Skip to content

Data & Analytics

Data Science

Analysis, statistics, experiments and models in service of one thing: a decision someone actually makes differently. We are sceptical by trade — about your data, about correlation dressed up as cause, and about models that look clever and change nothing.

Overview

Data science is often sold as the part where a model does something impressive. In practice the impressive model is the small end of the work, and frequently it is not needed at all. The job is to take a real business question — why are customers leaving, which change actually moved the number, what will demand look like next quarter — and answer it well enough, and clearly enough, that someone changes what they do. Everything in between is the craft: framing the question so it can be answered, getting the data into a state you can trust, choosing an analysis that fits the question rather than the fashion, and communicating the result so it lands with the people who decide.

Two things kill more data science projects than any modelling mistake. The first is a vague question. "Tell us what the data says" has no answer; a project that starts there produces a slide deck nobody uses. The second is dirty data — figures that mean different things in different systems, gaps that are not random, definitions that shifted halfway through the period. Most of the honest effort on any serious analysis goes into understanding and cleaning the data, and any consultancy that pretends otherwise is either inexperienced or selling you something. We spend real time on both, up front, because a fast answer to the wrong question or a confident answer built on broken data is worse than no answer at all.

We are also, deliberately, sceptics. Correlation is not causation, and a chart that suggests one thing is not evidence it is true. It is easy to overfit a model so it explains the past beautifully and predicts the future badly; easy to torture a dataset until some comparison crosses a significance threshold by chance. A rigorous, doubting approach — testing whether a finding holds up, asking what else could explain it, being honest about uncertainty — matters far more than the sophistication of the technique. The value is not the model. The value is a decision made better, and that is what we tie the work to.

Who it’s for — Organisations sitting on data they suspect holds an answer, who want a rigorous, honest read on a real decision rather than a dashboard or a model for its own sake.

What you get

  • A sharpened business question and a clear statement of what a useful answer would let you decide differently
  • Honest data assessment and cleaning — the state of the data named plainly, not glossed over
  • Exploratory analysis that surfaces what is actually in the data before any model is reached for
  • The right analysis for the question — descriptive, statistical, experimental or predictive — not a model applied by reflex
  • Findings communicated for the decision-maker: the answer, the confidence in it, and what it does not tell you
  • Where a model is warranted, a reproducible pipeline that can be re-run and, if it earns it, put into production
  • A candid account of the limitations — what we could not conclude and why — so the work is not over-trusted

What Data Science does for you

  • Better decisions, not more dashboards

    The point of the work is a choice made differently. We tie every piece of analysis to a decision on the table, so you get a recommendation you can act on rather than another report that gets admired and filed.

  • Findings you can trust under scrutiny

    When a result will be questioned in a board meeting or bet on with real money, it has to survive challenge. We test findings against alternative explanations and state our confidence plainly, so what you present holds up rather than falling apart on the first hard question.

  • Money not spent on the wrong model

    A great deal of data science spend goes on sophisticated models that a simple analysis would have beaten, or that answer a question nobody needed. Starting from the decision and preferring the simplest sufficient method keeps the effort proportionate to what is actually at stake.

Why teams choose us for Data Science

  • Senior data scientists do the work — the person interrogating your data and defending the conclusion is the person you met, not a junior handed the brief after the sale
  • We are honest about data quality: if the data cannot answer your question, we say so and tell you what it would take, rather than producing a confident answer on a broken foundation
  • We are sceptical by discipline — about causation, overfitting and cherry-picked significance — because a wrong conclusion held confidently is more dangerous than an open question
  • We tie the work to a decision and stop when it is answered; you are not sold an open-ended modelling programme when a two-week analysis settles the question

What Data Science includes

The concrete pieces of work this covers — scoped to what your problem actually needs.

  • Framing the question

    Turning a vague ask — "what does the data tell us" — into a specific, answerable question tied to a decision, and being willing to say when the question you brought is not the one worth answering.

  • Exploratory data analysis

    The disciplined first look: understanding distributions, spotting the gaps and errors and impossible values, seeing what is actually in the data before any hypothesis or model is reached for. Most of the real insight often surfaces here.

  • Statistics and hypothesis testing

    Applying statistical method properly — testing whether a difference is real or noise, quantifying uncertainty with intervals rather than false precision, and resisting the pull to keep slicing until something crosses a threshold by chance.

  • Experimentation and A/B testing

    Designing experiments that actually establish cause: sensible controls, enough power to detect an effect worth caring about, and analysis that is not undermined by peeking early or testing twenty things and reporting the one that worked.

  • Predictive and descriptive modelling

    Descriptive models that explain what is driving an outcome and predictive ones that forecast it — built with validation that reflects how they will be used, and with a hard eye on overfitting so the model works on data it has not seen.

  • Communicating findings

    Translating the analysis into something a decision-maker can act on — the answer, the confidence, the caveats, the recommendation — without hiding uncertainty behind jargon or overstating a result to make it land.

Where it fits

  • Understanding what drives an outcome

    Customers are churning, or conversion is falling, or a cost is creeping up, and you need to know why rather than that it is happening. We analyse the drivers, separate the real signal from the coincidences, and tell you which levers actually move the outcome and which only appear to.

  • Forecasting for a real decision

    Demand, capacity, revenue or risk that you need to plan against. We build a forecast grounded in the data’s actual behaviour — seasonality, trend, the irreducible uncertainty — and present it with honest error bars, so you plan against a range rather than a single number pretending to be exact.

  • Proving whether a change worked

    You changed a price, a page, a process, and the headline number moved — but was it the change, or something else that happened at the same time? We design or analyse the experiment properly, so you can attribute the effect with confidence instead of claiming a win the data does not actually support.

  • A one-off analytical question

    A specific question that needs answering once, rigorously — which segment behaves differently and why, whether two things are really related, what a pricing change would plausibly do. We answer it, document how, and hand you a conclusion you can rely on and move on from.

How we approach Data Science

We start from the decision, not the data. Before anyone opens a dataset we want to know what you would do differently depending on the answer — because a question whose answers all lead to the same action is not worth investigating, and a question you cannot act on is not worth answering. Framing this properly is often the most valuable hour of the engagement, and it is where we are most willing to push back on the brief you arrived with.

From there the work is deliberately unglamorous and deliberately sceptical. We interrogate the data before trusting it, we prefer the simplest analysis that answers the question over the most impressive one, and we hold our own findings at arm’s length — asking what else could explain the pattern, whether it survives a fair test, and how confident we can honestly be. You get the answer, the strength of the evidence behind it, and a straight account of what it does not settle.

How we work through a data science problem

We begin with the question and the decision behind it, not the data. We want to understand what you would do differently depending on the answer, what "good enough to act" looks like, and what the cost of being wrong is — because those determine how much rigour the problem earns and often reveal that the real question is not the one in the brief. This framing is short but it is the most leveraged part of the whole engagement.

Then comes the unglamorous majority of the work: getting the data and understanding it. We assess what you have, find the gaps and the errors and the shifting definitions, and clean it into a state we can reason about — reporting honestly on its condition as we go. Exploratory analysis follows, and frequently the answer, or most of it, appears here before any model is built. We resist reaching for a technique until the data has been understood on its own terms.

Only when the question warrants it do we build a model or a formal test — and we validate it the way it will actually be used, guard hard against overfitting and spurious significance, and check that a finding survives being challenged. The engagement ends not with a technical artefact but with a communicated conclusion: the answer, how confident we are, what it does not tell you, and what we would do next. If a model deserves to live on, we make the work reproducible so it can be re-run and, where justified, productionised.

From analysis to production, and reproducibility

A one-off analysis and a model that runs in production are different engineering problems, and conflating them causes trouble at both ends. The analysis needs to be reproducible — someone should be able to re-run it from the raw data and get the same result, which means the data pull, the cleaning steps and the analysis are captured as code rather than as a sequence of manual steps in a notebook that only worked once. A conclusion nobody can reproduce is a conclusion nobody should fully trust, including us.

When a model earns a place in production — scoring records, forecasting on a schedule, feeding a decision automatically — it stops being analysis and becomes software that has to run reliably. That is a genuine step: the ad hoc script that produced a good result in a notebook is not the same as a pipeline that ingests fresh data, applies the identical cleaning and transformation the model was trained on, produces a prediction, and is monitored for the day its inputs drift and its accuracy quietly decays. We are candid about when that step is worth taking and when a periodic re-run of a documented analysis is the more honest, cheaper answer.

Reproducibility is the thread that runs through both. Versioned code, a recorded account of which data and which choices produced a result, and an environment that can be recreated — these are what let a finding be checked, a model be retrained without mystery, and the work be handed over without walking out of the door with the only copy of how it was done. Where a model goes live, the same discipline is what makes retraining and monitoring possible rather than a periodic act of faith.

Data governance, PII and doing analysis responsibly

Data science works with your data at its most concentrated, which makes governance a first-order concern rather than a compliance footnote. We treat access on a need-to-use basis: we work with the minimum data the question requires, prefer anonymised or aggregated data where it will still answer the question, and are deliberate about where analytical copies of data live, who can reach them, and when they are deleted. Sensitive data does not get scattered across laptops and ad hoc exports because an analysis was in a hurry.

Personal data brings obligations that outlast the project. Under UK GDPR the purpose the data was collected for constrains what it can lawfully be used for, and an analysis is a use — so we are careful that what we do with personal data is compatible with why it was gathered. Where identifying individuals is not necessary to answer the question, and it usually is not, we pseudonymise or aggregate so that the analysis never depends on data more sensitive than the problem requires. This is both the law and simply the responsible way to work.

There is an ethical dimension that governance alone does not cover. A model trained on historical data learns the biases in that history, and a statistically valid finding can still be used to justify something unfair or simply wrong. We flag when a dataset is skewed in a way that would make a model discriminate, when a correlation is being read as a licence to act on individuals, and when a finding is being stretched past what it can honestly support. Being trusted with an organisation’s data means saying so — quietly and early — rather than delivering a technically correct result that causes harm.

Signs it’s time

  • You need to understand the drivers behind something — churn, conversion, cost, failure — not just observe that it is happening
  • A decision hinges on a forecast, and you want one built on evidence rather than a confident guess
  • You are about to change something and want to know, properly, whether the change actually works — an experiment, not a vibe
  • There is a specific analytical question you need answered once, rigorously, so you can move on and decide

Data science, data analytics and machine learning — and where the lines are

These three overlap enough to be used interchangeably in marketing, which helps nobody. Data analytics is largely about reporting and business intelligence: describing what happened, tracking it in dashboards, answering known questions with known metrics. It is indispensable and it is not what this service primarily is. Data science reaches further into the uncertain — asking why, testing whether, predicting what next — using statistics and modelling to answer questions that a dashboard cannot, and it leans on good analytics and good data engineering underneath it.

Machine learning is, in one honest framing, the subset of data science concerned with building models that learn patterns and make predictions, especially models intended to run in production and make or inform decisions at scale. A churn analysis that tells you why customers leave is data science; a model that scores every customer’s churn risk nightly and feeds a retention system is machine learning territory. The methods overlap heavily — the same statistical foundations, the same discipline about validation and overfitting — but the intent differs: insight for a human decision versus an automated prediction that runs on its own.

We work across all three and are straight about which one your problem actually needs, because the expensive mistakes come from mismatching them. Reaching for a production machine learning system when you needed a two-week analysis to answer a question once is money burnt; building yet another dashboard when the question is genuinely "why" and "what if" leaves the real question unanswered. We start from the decision and let it decide which discipline — often it is more analysis and better data than anyone expected, and less modelling.

Technologies we build it with

Chosen per problem, not per fashion — this is the stack we most often reach for on this work.

How we deliver

  1. 01

    Discover

    We map the system, the constraints and the business it serves — including the parts nobody documented.

    Architecture brief

  2. 02

    Architect

    Decisions get made, written down and defended before a line of production code exists.

    Decision records

  3. 03

    Build

    Short cycles against working software. You see progress in the product, not in a status deck.

    Shipping increments

  4. 04

    Operate

    Monitoring, incident response and iteration. The system is alive, so the engagement is too.

    Runbooks & SLOs

What changes

  • A decision you can defend

    Not a chart, but an answer to a real question — with the strength of the evidence stated, so you know how much weight it can bear.

  • Rigour over impressiveness

    Findings that have been tested against the obvious alternative explanations, so you are not acting on a coincidence or an artefact of dirty data.

  • Honest uncertainty

    A clear account of what the data does and does not support, so nothing is over-trusted — and you know where the guardrails are.

Industries we serve

Domain knowledge changes what gets built. A few of the sectors we know before the first meeting.

How pricing works

  • The largest and least predictable driver is the state of your data. A clean, well-understood, well-documented dataset lets us get to the actual question quickly; data scattered across systems with inconsistent definitions and unexplained gaps means the assessment and cleaning become a real part of the work. We would rather scope that honestly, and sometimes do a short data assessment first, than quote a fixed number against data neither of us has looked at yet.
  • The nature of the question matters as much as the data. A well-defined analytical question — is this real, what is driving this, what does the experiment say — is bounded work with a clear end. An open-ended predictive modelling problem, or one where the requirement is to get a model into production and keep it there, is a larger and more iterative undertaking, and it is priced as one.
  • We can work to a fixed scope for a clearly-defined analysis with an answer as its deliverable, or on a time-and-materials basis where the investigation is genuinely open and the next step depends on what the last one found — which is common in exploratory work, where you cannot honestly plan step five before you have seen the result of step two. We will tell you which shape fits your problem rather than force it into the one that is easier to quote.
  • Where the honest answer is that a smaller first piece — a data assessment, or a single focused analysis — should come before any larger commitment, we will say so. It is cheaper for you to find out early that the data cannot support the ambitious question than to fund a large programme that discovers it in month three.

Typical timeline

  1. 01

    Framing and data assessment

    Sharpening the question against the decision behind it, and taking an honest look at the data — what exists, what state it is in, and whether it can actually answer what you are asking. Short, and occasionally it changes the whole engagement.

  2. 02

    Cleaning and exploration

    Getting the data into a state we can trust and exploring it properly. Frequently the majority of the effort, and frequently where the answer — or most of it — first appears, before any model is built.

  3. 03

    Analysis, testing or modelling

    The core work chosen to fit the question — statistical testing, experiment analysis, or predictive and descriptive modelling — with validation that reflects real use and a hard guard against overfitting and spurious results.

  4. 04

    Communication and next steps

    The conclusion delivered for the decision-maker: the answer, the confidence in it, its limits, and a recommendation. Where a model earns it, reproducible work handed over ready to re-run or productionise.

Why teams choose us for Data Science

  • We are sceptics, and that protects you

    Our instinct is to doubt a finding until it has survived a fair test — to ask what else could explain it, whether the data supports it, how confident we can honestly be. That scepticism is exactly what stops you acting on a coincidence, an overfit model or a result that was really just dirty data.

  • We start from the decision, not the technique

    We care what you will do differently with the answer, which keeps the work tied to something that matters and keeps us from reaching for an impressive model when a simple analysis would settle it. An insight nobody acts on is worthless, so we do not produce them.

  • Honest about data and about limits

    If the data cannot answer your question, we tell you, and we tell you what it would take to change that. If a finding is weaker than you hoped, you hear that too. You get the real picture, including its uncertainty — not a confident story that falls apart under scrutiny.

  • Senior throughout, and we can operate it

    The person doing the analysis is the senior data scientist you met, and where a model goes into production we build it as software that has to run and be monitored — because we may be the ones running it, which changes every decision for the better.

How to engage us

Three ways to work with us on this — chosen to fit the problem, not our margin.

Related services

Part of Data Engineering. Other work we do alongside this.

Common questions

What is the difference between data science, data analytics and machine learning?

Data analytics is mainly reporting and business intelligence — describing what happened and tracking it. Data science reaches further, into why something is happening, whether an effect is real, and what will happen next, using statistics and modelling. Machine learning is best understood as the part of data science focused on models that learn patterns and make predictions, especially ones that run in production. They overlap heavily and share the same discipline; the difference is intent, and matching the right one to your problem is where a lot of wasted spend is avoided.

Do we need a machine learning model, or will an analysis do?

Far more often than people expect, an analysis is enough. If the goal is to understand something and make a decision once, a rigorous piece of analysis answers it faster and cheaper than building and maintaining a model. A model earns its keep when the same prediction has to be made repeatedly and automatically at scale. We start from your decision and tell you honestly which one it needs — and the honest answer is frequently "less modelling and better data than you thought".

Our data is messy and scattered. Can you still work with it?

Usually yes, but this is the part we are most honest about. Assessing and cleaning the data is often the largest share of a real project, and how messy the data is directly affects cost and timeline. Sometimes the honest finding is that the data cannot support the question you are asking — a gap that is not random, a definition that changed, a signal that simply is not there. We would rather establish that early, sometimes with a short data assessment first, than deliver a confident answer built on a broken foundation.

How do you make sure a finding is actually real and not a coincidence?

By being sceptical on purpose. Correlation is not causation, and it is easy to overfit a model to the past or to keep slicing data until something crosses a significance threshold by chance. We test whether a finding survives a fair test, ask what else could explain the pattern, validate models on data they have not seen, and quantify uncertainty honestly rather than reporting false precision. And we tell you the strength of the evidence, so you know how much weight a conclusion can actually bear.

What do we actually get at the end — a report, a model, or something else?

It depends on the question, and we agree that at the start. For most analytical questions the deliverable is a communicated conclusion built for a decision-maker: the answer, the confidence in it, what it does not tell you, and a recommendation — plus the reproducible work behind it so it can be checked and re-run. Where a model deserves to live on, you get a validated, documented, reproducible pipeline that can be put into production and monitored. What you will not get is an impressive artefact that changes nothing, because a result nobody acts on was not worth producing.

Let’s talk about Data Science.

Tell us what you’re building or fixing. A senior engineer reads every enquiry and replies within a business day.

Two fields required. We reply to real enquiries — no list, no sequence.