8 October 20267 min read

Who owns what? Running a shared data platform for consumer applications

A follow-up to my article on shared platforms: the architecture, using an insurer's customer data, and the split of responsibilities between the platform team, data producers, data governance and executive leadership.

By Srini Vankeepuram · Architecture, Engineering, Data & AI Leader

After I published my article on winning consumers for a shared platform, a product owner I've worked with called me. They'd read it and had sharp questions. Where does the platform team's job end and mine begin? If my marketing data is wrong, who fixes it? Who decides whose request goes first? And who do I call at 2 a.m.?

They're the right questions. A shared data platform rarely fails because of its technology. It fails when nobody is clear who owns what, so every problem lands in a queue between teams. This article is my answer, using an insurer's customer data as the example: the architecture first, then who is in charge of each part.

The architecture: customer data at an insurer

Picture a consumer insurer selling car and home cover. Policyholders buy and renew through a self-service portal, adding drivers, vehicles and addresses. The motor and home policy systems hold the cover itself. The marketing team runs promotions and records who responded and through which channel. Third-party providers enrich the picture: Dun & Bradstreet with company data for business and fleet customers, and LexisNexis Risk Solutions with identity, claims and no-claims history. Each of these is a data producer.

The shared data platform (SDP) sits in the middle. Producers publish into it through agreed data contracts. The platform matches records into one trusted customer and household view, validates the data and serves it to quote, policy, claims and service applications in real time. It also feeds Snowflake, the analytical warehouse, for pricing analysis, reporting and machine learning.

PRODUCERSCONTRACTSSHARED DATA PLATFORMCONSUMERSSelf-service portalPolicyholders, driversAPI callsreal timePolicy systemsMotor and home coverEventsstreamingMarketing teamPromotions, responsesETLnightly batchThird-party dataD&B, LexisNexisETL / APIfiles, lookupsDATA CONTRACTS · VALIDATIONQuarantineback to producerCustomer viewMatching and master dataCanonical modelCustomer, policy, claim, campaignQuality and lineageChecks and source-to-use trailAccess and privacyConsent, masking, auditServing layerAPIs and events, with published SLOsCustomer applicationsQuote and buyPolicy and renewalsClaims and serviceAPI calls and eventsSnowflake (OLAP)Full history, shaped for questionsELT / CDCscheduled, near real timeBI & reportsPricing & MLDATA GOVERNANCEOwners, definitions, privacy policy, enforced as pipeline checksEXECUTIVE LEADERSHIPFunding as a product, mandate, tie-breaks
Logical architecture: producers send data by API calls, events and ETL; contracts check it at the door; the platform serves applications through API calls and loads Snowflake through ELT.

Here is what each layer holds.

Data producers

Create the data and own its accuracy at source.

  • Self-service portalPolicyholder details, drivers, vehicles, consents and preferences.
  • Policy systemsMotor and home policies, renewals, cover changes and payments.
  • Marketing teamPromotions, campaign responses and channel preferences.
  • Third-party providersDun & Bradstreet company data; LexisNexis identity, claims and no-claims history.

Ingestion through contracts

How data enters: on the contract's terms, never around it.

  • Data contractsSchema, meaning, quality rules, owner and freshness, versioned.
  • PipelinesBatch ETL, streaming events and APIs, built by producers to the contract.
  • Validation at the doorRecords that break the contract are quarantined, not loaded.

Shared data platform

Built once, run well, owned by the platform team.

  • Customer and household viewMatching and master data: one trusted policyholder, linked to their household.
  • Quality and lineageChecks, metadata and a trail from source to use.
  • Access and privacyRole-based access, consent and masking of personal data.
  • Serving layerAPIs and events for applications, with published SLOs.

Consumers

Use the data to serve customers and decide better.

  • Customer applicationsQuote, buy, renew, claims and service, reading through APIs and events.
  • Snowflake (OLAP)Pricing analysis, reporting and machine learning on full history.
  • Analysts and data scientistsSelf-service queries on governed, documented data.
A shared data platform for an insurer's customer data: producers publish through contracts; the platform serves operational apps and feeds Snowflake.

The split between operational and analytical use matters. A quote or a claim needs current, accurate data in milliseconds; Snowflake needs complete history, shaped for questions. The platform serves the first through API calls and events, and feeds the second through ELT and change data capture (CDC), which copies only what changed, as it changes, rather than reloading whole tables overnight. Both come from the same governed model, so the number on the dashboard matches what the customer sees.

Customer data in a domain-driven world

In a domain-driven organisation, each domain owns its own data. Motor owns motor policies, claims owns claims, marketing owns campaigns. That's healthy: the people closest to the data own it. But every one of those domains also holds a copy of the customer, and they rarely agree. The motor system knows Jane Smith at one address; the home system knows J. Smith at another; marketing knows an email address that matches neither.

So the customer becomes a domain of its own. It owns identity, contact details, consent and the household: who lives with whom, and who is insured for what. Other domains keep the data they need for their own work, but they refer to the customer through one shared identifier and take changes from the customer domain rather than editing their own copy.

  • Domains own their facts: a policy belongs to the policy domain, a claim to claims, a campaign response to marketing.
  • The customer domain owns who the customer is: identity, contact details, consent and household links.
  • The platform publishes the trusted customer as a data product, with a contract, so every domain uses the same identity.
  • Third-party data enriches the customer record; it doesn't overwrite what the customer told you without a rule that says so.

What the platform team is in charge of

The platform team owns the platform as a product. It's led by a shared platform lead, who owns the roadmap and is the single point of accountability when something needs escalating. That remit is narrower than many people assume, and the narrowness is what makes it work.

  • Data modelling: the canonical model, and how it changes. Producers and consumers propose; the platform team decides, so the model stays coherent.
  • Data contracts: the template, the tooling and the checks. Every contract names an owner, a schema, quality rules and freshness, and every change is versioned and backward compatible.
  • Prioritisation: one public backlog for every producer and consumer request, ranked against agreed criteria, not by who shouts loudest.
  • SLAs and SLOs: published targets for availability, latency and data freshness, on dashboards everyone can see.
  • Support: a single front door, with clear routes for incidents, requests and questions, and an on-call rota for the platform itself.
  • Security and privacy controls: access, consent, masking and audit trails, built into the platform rather than bolted on by each consumer.
The platform team owns the road, the rules of the road and the signage. It doesn't drive everyone's car.

What the producers are in charge of

This is where my friend's first question lands. The producer owns the data, so the producer owns its accuracy. The platform can catch bad data at the door, but it can't know that a campaign code is wrong or that a supplier changed a field's meaning overnight.

  • Ingestion: building and running their pipelines (ETL, events or APIs) to the contract.
  • Quality at source: fixing data where it's created, not patching it downstream.
  • Meaning: keeping definitions accurate and telling the platform before anything changes.
  • Their own incidents: when their feed breaks or their data is wrong, they're first on the call.
  • Third parties: a producer team owns each external provider's feed and holds the provider to its contract.

Who's in charge of what, at a glance

AreaPlatform teamProducersData governanceExecutive leadership
Canonical data modelACCI
Data contractsA (standard and checks)A (their contracts)CI
Ingestion pipelines (ETL, events, APIs)CAII
Data quality at sourceCACI
Definitions and data ownershipCCAI
Prioritisation of the backlogACCC
SLAs and SLOsACII
Support and incidents (platform)ACII
Access, privacy and retention policyCCAI
Funding, mandate and conflictsCCCA
Responsibilities in a shared data platform. A = accountable (one owner); C = consulted; I = informed.

The role of data governance

Data governance sets the rules the platform enforces. It decides who owns each data domain, agrees business definitions, and sets policy on privacy, retention and access. It also resolves disputes about meaning: when marketing and finance define "active customer" differently, governance decides, and the contract records the answer.

Good governance is light and fast. It names owners and makes decisions; it doesn't become a committee every change must queue for. The platform team turns its policies into automated checks, so most governance happens in the pipeline, not in a meeting.

The role of executive leadership

Executives don't need to understand the data model. They need to do three things that nobody else can.

  • Fund the platform as a product, with a long-lived team, not as a project that ends at go-live.
  • Give it a mandate: the platform is the agreed route for shared data, and new duplicates need a reason.
  • Break ties: when two senior stakeholders both say their request comes first, the decision goes up, not round in circles.

As I wrote last time, a mandate is a backstop, not a strategy. But a shared data platform without visible executive backing becomes optional, and optional platforms get worked around.

Back to the product owner's questions

QuestionAnswer
Where does the platform's job end and mine begin?The platform owns the model, contracts, controls and service levels. You own your product's use of the data and the requests you raise.
If my marketing data is wrong, who fixes it?The producer, at source. The platform should have caught it at the door if it broke the contract, and will help find the cause.
Who decides whose request goes first?The platform team, in one public backlog, against agreed criteria. Ties go to executive leadership.
Who do I call at 2 a.m.?The platform's single front door. It routes to the platform on-call, or to the producer whose feed broke.
The questions I was asked, and the short answers.

The short version

  • Producers own their data and its quality; they publish through contracts.
  • The platform team owns the model, contracts, controls, priorities, SLOs and support.
  • Data governance owns definitions, ownership and policy, enforced in the pipeline.
  • Executive leadership funds the platform as a product, mandates it and breaks ties.
  • Serve applications and the warehouse from one governed model, so everyone sees the same number.
Clear ownership is the cheapest performance improvement a data platform will ever get.

Contact

Let's talk.

I'm based in London, UK. Whether it's a leadership role, an advisory engagement or a transformation that needs shaping — my inbox is open.

srini.vankee@gmail.com