Skip to content

Data & Analytics

Data Management Services

We make an organisation’s data trustworthy, findable, secure and compliant. The unglamorous foundation that every analytics and AI initiative quietly depends on.

What Data Management means in practice

Who it’s for: Organisations whose data has quietly stopped being trustworthy (conflicting records, numbers nobody believes, no one able to find what exists), or who face GDPR and compliance pressure, or whose analytics and AI initiatives keep failing on bad data underneath.

Data management is the discipline that keeps the data an organisation runs on accurate, consistent, findable and safe to hold. It is the least glamorous work in the whole data field, and it is almost universally neglected until it causes a crisis: a regulator asking a question nobody can answer, a board decision made from a number that turned out to be wrong, or the slow realisation that nobody can find or trust the data they have. This service is about doing that foundational work deliberately, before the crisis, rather than scrambling to reconstruct it during one.

It sits underneath the analytical services rather than beside them. Data engineering builds the pipelines and the warehouse; data science and analytics build the models and the reports on top. But all of that rests on an assumption that is often false: that the underlying data is correct, that a customer is a single customer rather than four contradictory records, that a field means the same thing everywhere it appears, and that you are allowed to hold and use the data the way you are. When that assumption breaks, “garbage in, garbage out” takes over and every downstream initiative inherits the mess. Master data management, data quality, governance, cataloguing, lineage, retention and regulatory compliance are the disciplines that keep the assumption true.

We will be honest about the failure mode at the other end, too. Governance done badly becomes bureaucracy: a data council that meets monthly to approve nothing, a policy document nobody reads, a catalogue that goes stale the week after it launches. The point of this work is not process for its own sake; it is trustworthy data that people actually use and a compliance posture that survives scrutiny. We build the lightest governance that achieves that, wire it into the tools people already work in so it maintains itself, and stop there rather than building a bureaucracy that outweighs the problem it was meant to solve.

What you get

  • A data quality assessment, profiling the accuracy, completeness, consistency and uniqueness of your key datasets, with the concrete problems quantified rather than described in the abstract
  • Master data management for the entities that matter. A single authoritative record for customers, products or suppliers, with the matching, merging and survivorship rules that resolve conflicting duplicates
  • A pragmatic governance framework, data ownership and stewardship that names actual people, policies that are short enough to be read, and decision rights that are clear rather than ceremonial
  • A data catalogue and lineage. A searchable inventory of what data exists, what it means, where it came from and what depends on it, so people can find and trust data instead of asking around
  • Data classification and PII handling, sensitive and personal data identified, labelled, access-controlled and, where appropriate, masked or minimised, with retention and deletion rules that actually run
  • UK GDPR and data-protection alignment. A record of what personal data you hold and why, a lawful basis for it, and the machinery to answer data-subject access, correction and erasure requests without a fire drill
  • Documentation and handover so governance is something your own stewards and engineers operate day to day, not a consultant’s artefact that decays the moment we leave

What Data Management does for you

  • The foundation every analytics and AI project stands on

    Every dashboard, forecast and model inherits the quality of the data beneath it. Feed a model duplicated customers, inconsistent categories and missing fields and it will confidently produce wrong answers: “garbage in, garbage out” is not a slogan here, it is the mechanism by which expensive initiatives fail. Getting the data management right is the cheapest, highest-leverage thing you can do to make everything downstream work, and it is almost always the step that was skipped when a data project underdelivers.

  • Decisions made on numbers you can defend

    When the underlying records are trustworthy and a metric means one thing everywhere, the organisation stops wasting meetings arguing about whose figure is correct and starts acting on a shared version of reality. The value is not just faster reporting; it is that a decision made from the data can be defended, traced back to its source, and trusted by the people who have to live with it, instead of being quietly overridden by whoever has the most convincing private spreadsheet.

  • A compliance posture that survives scrutiny

    Holding personal data is a liability as much as an asset, and UK GDPR makes that concrete: you must know what you hold, why, on what lawful basis, and be able to honour a data subject’s rights over it. Doing this properly turns a looming audit or a data-subject request from a panicked reconstruction into a routine lookup, and it reduces the blast radius of any breach, because data you have classified, minimised and deleted on schedule is data that cannot leak.

Why teams choose us for Data Management

  • Your data has quietly stopped being trustworthy, conflicting records, metrics that disagree, a “single customer view” that does not exist, and you want it measured and fixed by engineers who treat data quality as the foundation it is, not a clean-up to do after the interesting work.
  • You face real data-protection pressure (an access request, an audit, a due-diligence questionnaire), and you want a genuine, evidenced UK GDPR posture, not a policy document written to look reassuring that falls apart the moment someone asks where a specific person’s data actually lives.
  • You have felt the other failure too (governance that became bureaucracy), and you want the lightest framework that actually keeps data trustworthy, wired into the tools people already use, rather than a data council and a stack of policies everyone routes around.
  • Your analytics or AI work keeps failing on bad data underneath, and you want someone to fix the foundation properly (master data, quality, lineage), rather than sell you another model or dashboard that will inherit the same mess.

What Data Management includes

The concrete pieces of work this covers, scoped to what your problem actually needs.

  • Data quality, profiling, measurement and remediation

    We start by measuring, not asserting: profiling your key datasets for accuracy, completeness, consistency, uniqueness, validity and timeliness, so the problems are quantified rather than felt. Then we remediate the ones that matter (standardising formats, filling or flagging gaps, resolving contradictions), and, crucially, put in place the ongoing checks and monitoring that catch quality problems as they arise rather than months later in a report. Quality is not a one-off clean-up; it is a property you maintain, so we build the rules and alerts that keep it from decaying back.

  • Master data management (MDM)

    The reason a customer appears four times with three spellings is that no system owns the authoritative record. We build that authority: matching and deduplication rules to identify when two records are the same real entity, survivorship rules to decide which values win when they conflict, and a golden record that the rest of the organisation can rely on for customers, products, suppliers or whatever your core entities are. This is the work that finally makes a genuine single view possible, and it is unglamorous, fiddly and enormously valuable.

  • Data governance, ownership and stewardship

    Governance that works names real people and gives them real, bounded authority: an owner accountable for each important dataset, stewards who curate it, and decision rights that are clear rather than ceremonial. We write policies short enough to be read and specific enough to be followed, and we resist the urge to build a heavyweight council that meets to approve nothing. The aim is the minimum structure that keeps data trustworthy and accountable, enough governance to matter, not so much that people route around it.

  • Data cataloguing and lineage

    People cannot use or trust data they cannot find or understand. A catalogue is a searchable inventory of what data exists, what each field means, who owns it and how sensitive it is; lineage traces where a number came from and what depends on it, so when a figure is questioned there is an answer and when a source changes you know what breaks. We stand these up with modern catalogue and lineage tooling, and (the part that actually determines success), wire them into your pipelines so they stay current instead of going stale the week after launch.

  • Classification, PII handling and access control

    Not all data carries the same risk, so we classify it (public, internal, confidential, and specifically personal and special-category data), and let that classification drive the controls. Personal and sensitive data is identified wherever it lives, access is restricted to those who genuinely need it, and where appropriate the data is masked, tokenised or minimised so a breach or a mistaken query exposes as little as possible. This is where data management and security meet: the classification is what makes least-privilege access and sensible retention enforceable rather than aspirational.

  • Retention, lifecycle and UK GDPR compliance

    Data has a lifecycle, and holding it forever is a liability, not prudence. We define retention schedules tied to purpose and legal basis, and (the step most organisations skip), make deletion actually happen on schedule rather than remain a policy nobody enforces. On the regulatory side we build the practical machinery of UK GDPR compliance: a record of processing, a defensible lawful basis for the personal data you hold, and repeatable processes to honour data-subject rights (access, rectification and erasure), so those requests are handled routinely instead of by hand each time.

Where it fits

  • Building a single customer view from conflicting records

    An organisation whose customers exist as duplicated, contradictory records scattered across a CRM, a billing system and a support tool, so nobody can answer “how many customers do we have” with a straight face. We profile the mess, build the matching and survivorship rules that collapse the duplicates into one authoritative record per real customer, and stand up the master-data process that keeps it that way. The result is a genuine single view that marketing, finance and support can all rely on, instead of three systems quietly disagreeing.

  • Getting ready for a GDPR audit or a wave of access requests

    A business that has realised it cannot readily say what personal data it holds, where it lives or on what basis, and now has a regulator, a large customer or a surge of data-subject requests forcing the issue. We map the personal data across systems, establish the lawful basis and a record of processing, classify and access-control the sensitive data, and build a repeatable process for answering access, correction and erasure requests. What was a fire drill becomes a lookup, and the compliance posture is one you can actually evidence.

  • Fixing the data foundation under a failing AI initiative

    A machine-learning or analytics project that keeps producing nonsense, where it has slowly become clear the model was never the problem. The training data is incomplete, inconsistent and full of duplicates. We do the unglamorous foundational work the project skipped: profile and remediate the quality problems, resolve the master-data conflicts, and document lineage so the team knows what they are actually feeding the model. Often the initiative that looked like it needed a better algorithm just needed data it could trust.

  • Making data findable and trusted across a growing organisation

    A company that has accumulated data faster than any understanding of it, where finding a dataset means asking around and trusting a number means knowing who to believe. We build a catalogue and lineage so people can discover what exists and see where it came from, assign ownership so each important dataset has someone accountable, and put lightweight governance around it. The change is cultural as much as technical: data stops being tribal knowledge and becomes something the organisation can find, understand and rely on.

How we approach Data Management

We start from the problem you can feel, not from a governance framework off a shelf. That usually means profiling the data first: measuring the actual completeness, consistency and duplication of the datasets that matter, so the conversation moves from “our data is a bit messy” to “thirty-one per cent of customer records have no valid postal address and there are four spellings of your largest account”. Concrete, measured problems are what justify and shape the work; an abstract maturity model is not. We fix the highest-leverage quality and master-data issues first, because those are the ones bleeding into every report and every model downstream.

Then we build governance to fit the organisation rather than an ideal of it. Ownership and stewardship name real people who already care about the data; policies are short enough that someone will actually read them; and the controls live in the catalogue, the warehouse and the tools people already use, so doing the right thing is the path of least resistance rather than a form to fill in. We deliberately build less governance than a textbook would prescribe, because the failure we see most often is not too little process but too much. A bureaucracy that people route around, which is worse than no governance at all because it creates the illusion of control.

How the engagement runs

We begin by measuring rather than theorising. The first work is profiling the datasets that matter (quantifying completeness, consistency, duplication and validity), and mapping where personal and sensitive data actually lives, so the conversation is grounded in specific, measured problems instead of a general sense that the data is messy. This is also where we learn how the organisation really works: who already cares about which data, where the trusted spreadsheets are and what they compensate for, and which conflicts and compliance gaps are doing the most damage. The output is a clear, prioritised picture of what is wrong and what it is costing, which is what shapes everything after.

From there we work in order of leverage, not in the order a framework prescribes. We fix the highest-impact quality and master-data problems first, because those bleed into everything downstream, and we stand up governance, cataloguing and compliance machinery incrementally around them rather than in a big-bang rollout that lands as bureaucracy. Each piece is wired into the tools people already use so it stays alive after we leave: catalogue entries maintained by the pipelines, quality checks running automatically, retention rules that actually execute. You see progress as measured improvements, duplication down, a metric reconciled, an access request answered from a process instead of by hand, and we hand over documentation and trained stewards so the discipline continues as an operating habit rather than a project that ended.

How we architect it

The architecture has two halves: a governance framework and the tooling that makes it real. The framework defines ownership, stewardship, classification and policy, who is accountable for what data, how sensitive each dataset is, and what the rules are for using, sharing and retaining it. We keep this deliberately lean, because a framework heavier than the organisation will carry is one it will quietly abandon. The tooling makes the framework operational rather than aspirational: a data catalogue as the searchable system of record for what exists and what it means, automated lineage so provenance and impact are visible without manual upkeep, and quality monitoring that runs as part of the pipelines rather than as a periodic manual audit.

For master data we design a clear model of the golden record and the matching, survivorship and stewardship rules around it, integrated with the source systems so the authoritative record stays authoritative as data changes. Classification is the connective tissue that ties governance to security: it drives access control, masking and retention, so a dataset labelled as containing personal data automatically inherits the right restrictions and lifecycle rather than depending on someone remembering. Throughout, we favour making the correct thing automatic (controls and metadata embedded in the flow of data), over relying on people to follow a policy manually, because governance that depends on constant human diligence is governance that decays.

Security and compliance

Security and compliance are not adjacent to this service; they are central to it, because the moment you hold personal or sensitive data you take on obligations and risk. It starts with classification: you cannot protect what you have not identified, so we find the personal and special-category data across your systems and label it, and that classification then drives everything else, access is restricted to those with a genuine need, sensitive fields are masked, tokenised or minimised so they are exposed as little as possible, and the data is encrypted in transit and at rest. The principle is least privilege and least data: hold the minimum personal data for the minimum time, so the blast radius of any breach is as small as it can be.

On the regulatory side we build the practical machinery of UK GDPR and data protection rather than a reassuring document. That means a record of what personal data you hold and why, a defensible lawful basis for processing it, retention schedules tied to purpose with deletion that actually runs, and repeatable processes to honour data-subject rights (access, rectification and erasure), so those requests are handled routinely instead of reconstructed by hand each time. Data-protection-by-design is the posture throughout: privacy considerations built into how data is collected, classified and retained, not reviewed as an afterthought before an audit. The result is a compliance stance you can evidence to a regulator, a customer or a partner, because it is grounded in controls that genuinely run rather than in intentions.

Signs it’s time

  • The same customer, product or account exists as several conflicting records across your systems, and no one can say which is the true one, so mailings duplicate, revenue is double-counted, and a “single view” is impossible
  • People no longer trust the numbers. The same metric means different things in different reports, decisions get relitigated over whose figure is right, and someone quietly keeps a private spreadsheet they believe instead
  • A regulator, customer or partner is asking data-protection questions (a data-subject access request, an audit, or a due-diligence questionnaire), and you cannot readily say what personal data you hold, where it is, or on what lawful basis
  • An analytics or AI initiative has stalled or produced nonsense because the data feeding it is incomplete, inconsistent or wrong, and it has become clear the model was never the problem. The foundation underneath it was

Our working method

Our organising belief is that the goal is trustworthy data people actually use, not process for its own sake, and that both neglect and over-governance defeat that goal. So we measure before we prescribe, fix the highest-leverage quality and master-data problems first, and build the lightest governance that keeps the data trustworthy. We are blunt about the trade-off: a heavyweight data council and a shelf of unread policies create the illusion of control while people route around them, which is worse than an honest absence of governance. We would rather ship a small, living framework that people follow than a comprehensive one they ignore.

The other half of the method is making the correct thing automatic. Governance that depends on constant human diligence decays the moment attention moves elsewhere, so we embed controls and metadata in the flow of data: catalogue entries maintained by pipelines, quality checks that run on every load, classification that drives access and retention without anyone remembering to apply it. And because this is foundational work that only pays off if it outlives us, we build for handover from the start: real stewards trained, documentation written, and the discipline established as an operating habit rather than a consultant’s artefact that quietly rots after the engagement ends.

Technologies we build it with

Chosen per problem, not per fashion. This is the stack we most often reach for on this work.

How we deliver

  1. 01

    Discover

    We map the system, the constraints and the business it serves, including the parts nobody documented.

    Architecture brief

  2. 02

    Architect

    Decisions get made, written down and defended before a line of production code exists.

    Decision records

  3. 03

    Build

    Short cycles against working software. You see progress in the product, not in a status deck.

    Shipping increments

  4. 04

    Operate

    Monitoring, incident response and iteration. The system is alive, so the engagement is too.

    Runbooks & SLOs

Want a straight answer on Data Management?

A short call with a senior engineer, before you write a brief. If Data Management is the wrong answer for your situation, we will say so and tell you what we think is right.

What changes

  • Data people trust again

    Quality problems measured and fixed, and a single authoritative record for the entities that matter, so the numbers stop being argued about and the private spreadsheets disappear because the shared data is finally believable.

  • Findable, understood data

    A catalogue and lineage that let anyone see what data exists, what it means and where it came from, so time is spent using data rather than hunting for it and asking around whether it can be trusted.

  • Compliance you can evidence

    A clear record of the personal data you hold and why, the classification and access controls around it, and the machinery to answer data-subject requests, so a regulator’s question is a lookup, not a fire drill.

Industries we serve

Domain knowledge changes what gets built. A few of the sectors we know before the first meeting.

How pricing works

  • A fixed-scope data quality and master-data assessment, profiling your key datasets, quantifying the duplication and inconsistency, and mapping where personal data lives, priced once the systems in scope are known, with a prioritised findings-and-fix report as the deliverable.
  • A focused compliance engagement, building the UK GDPR machinery of a record of processing, classification, retention and data-subject-request handling, scoped by the number of systems and the amount of personal data involved rather than by a fixed rate.
  • A delivery engagement to build the foundation, master data management, a catalogue and lineage, quality monitoring and a governance framework, quoted once the assessment has established which problems are worth fixing and in what order.
  • A monthly senior engagement for ongoing stewardship, where data quality, governance and compliance are maintained as the organisation and its data grow, rather than fixed once and left to decay back into the state that prompted the work.

Typical timeline

  1. 01

    Profiling and discovery

    One to two weeks measuring the actual quality of the key datasets, mapping where personal and sensitive data lives, and learning how the organisation really works, so the problems are quantified and prioritised rather than described in the abstract.

  2. 02

    Quality and master data

    Fixing the highest-leverage quality problems, building the matching and survivorship rules that collapse conflicting duplicates into authoritative golden records, and standing up the ongoing checks that keep quality from decaying back.

  3. 03

    Governance, catalogue and lineage

    Establishing lean ownership and stewardship, standing up a searchable catalogue and automated lineage wired into the pipelines, and classifying data so the framework drives access, masking and retention rather than sitting in a document.

  4. 04

    Compliance and handover

    Building the UK GDPR machinery (record of processing, lawful basis, retention that actually deletes, data-subject-request processes), and handing over to trained stewards with documentation so the discipline continues as an operating habit.

What working with us actually means

  • We treat the foundation as the priority, not the clean-up

    Most teams treat data management as the boring chore to do after the interesting analytics and AI work, which is exactly why those initiatives so often fail on the data underneath them. We treat the foundation as the priority it is, because we have watched clever models and expensive dashboards inherit the mess in the data and produce confident nonsense. Getting the quality, master data and lineage right is the highest-leverage work in the whole data stack, and we lead with it rather than bolt it on afterwards.

  • Honest about the over-governance trap

    We have seen governance become a bureaucracy that people route around, and we will tell you plainly that this is worse than no governance at all, because it manufactures the illusion of control. So we build the lightest framework that keeps data trustworthy, wire the controls into the tools people already use, and stop there. You are paying for judgement about how much governance is enough. A judgement that comes from having seen both the neglect and the over-correction fail.

  • We operate what we build

    Because we run the data foundations we design, the parts that only matter over time go in from the start: quality checks that keep running, catalogue entries that stay current because the pipelines maintain them, retention that actually deletes, and stewardship that survives handover. We optimise for a foundation that stays trustworthy as the data grows and the organisation forgets we were there, not for a tidy state on the day the project closes.

  • Senior engineers who understand the whole stack

    Data management sits beneath data engineering, analytics and AI, and getting it right requires understanding what those layers need from it. The people doing this work have built the pipelines, warehouses and models that sit on top, so the foundation we lay is shaped by what actually breaks downstream when it is wrong, not by a governance checklist written in isolation from the systems it is meant to serve.

How to engage us

Three ways to work with us on this, chosen to fit the problem, not our margin.

Related services

Part of Data Engineering. Other work we do alongside this.

Common questions

How is this different from your data engineering or analytics services?

Data engineering builds the pipelines and warehouse that move and store data; analytics and data science build the reports and models on top. This service is the foundation beneath all of them: making sure the data those layers rely on is actually accurate, consistent, deduplicated, findable and compliant to hold. It is the answer to “garbage in, garbage out”: you can build the finest pipeline and the cleverest model, but if a customer exists as four contradictory records and a field means three different things, everything downstream inherits that mess. We are frequently called in when an engineering or analytics project has stalled and it turns out the real problem was never the pipeline or the model, but the untrustworthy data underneath.

Isn’t data governance just bureaucracy that slows everyone down?

It becomes that when it is done badly, and we are the first to say so. A monthly data council that approves nothing, a stack of policies nobody reads, a catalogue that goes stale a week after launch. That kind of governance is genuinely worse than none, because it creates the illusion of control while people quietly route around it. Good governance is the opposite: the lightest structure that keeps data trustworthy and accountable, with the controls wired into the tools people already use so that doing the right thing is the path of least resistance rather than a form to fill in. The goal is trustworthy data people actually use, not process for its own sake, and we build only as much governance as it takes to get there.

What does master data management actually solve?

It solves the problem where the same real thing (a customer, a product, a supplier), exists as several conflicting records across your systems, so nobody can produce a single trustworthy view of it. Mailings duplicate, revenue gets double-counted, support cannot see a customer’s full history, and “how many customers do we have” has no honest answer. Master data management builds the authoritative record: matching rules to recognise when two records are the same entity, survivorship rules to decide which values win when they conflict, and a golden record the rest of the organisation can rely on. It is fiddly, unglamorous work, and it is what finally makes a genuine single view possible instead of three systems quietly disagreeing with each other.

We are worried about UK GDPR compliance, can you help us get it right?

Yes, and we build the practical machinery of compliance rather than a reassuring document that falls apart under scrutiny. That means mapping what personal data you actually hold and where it lives, establishing a defensible lawful basis and a record of processing, classifying and access-controlling the sensitive data, setting retention schedules that genuinely delete data when its purpose ends, and building repeatable processes to honour data-subject rights (access, correction and erasure), so those requests are handled routinely instead of reconstructed by hand each time one arrives. The test we hold the work to is whether you could answer a regulator’s or a large customer’s question with a lookup rather than a fire drill. We are engineers rather than a law firm, so where a genuinely legal judgement is needed we will say so, but the operational reality of protecting personal data is squarely our work.

How do you keep the quality and the catalogue from decaying after you leave?

By making the correct thing automatic and by handing over properly, because governance that depends on constant human diligence always decays the moment attention moves elsewhere. Quality checks run as part of the pipelines so problems are caught on every load rather than in a periodic manual audit; catalogue entries and lineage are maintained by the data flow itself rather than by someone remembering to update a wiki; classification drives access and retention automatically rather than waiting on a person to apply it. On top of that we train real stewards, document how everything works, and establish the discipline as an operating habit rather than a consultant’s artefact, because this is foundational work that only pays off if it outlives the engagement, and a foundation that rots the week after we leave was never worth building.

Thinking about Data Management?

Tell us the problem in your own words, not in requirements. A senior engineer reads it and comes back with a straight view on whether Data Management is the right answer here, or what would be.

  1. 01A senior engineer reads it. Not a form queue, and not an account manager.
  2. 02We reply either with questions or with a straight answer that we are not the right fit.
  3. 03If it looks like a fit, a technical call with the person who would actually run the delivery.
  4. 04Then scope, effort and risk in writing, before anyone signs anything.

Two fields required. We reply to real enquiries. No list, no sequence.