Writing · Data Platform

One PR to Ingest: Building a Contract-Driven Data Platform

From ticket queue to self-service — Protobuf schemas and YAML contracts that generate the pipeline


TL;DR

  • We turned an org that received data-ingestion requests as tickets and hand-built pipelines into a platform where submitting one PR — a Protobuf schema plus a YAML contract file — auto-generates the pipeline.
  • From a single contract we auto-generate and apply the ETL pipeline, Kafka topics, schema evolution, permissions, retention, compaction, and lineage. Today it runs at hundreds of data models · dozens of pipeline contracts · real-time CDC from a dozen-plus sources · a dozen DaaS offerings feeding company-wide AI workloads.
  • Data ingestion became self-service — complete once the Data Owner submits a schema — and an AI skill automated writing the contract itself. This is the foundation that cut request lead time from 15 days to 3.
Data modelsHundredsProtobuf schemas
Pipeline contractsDozens+ a subscription-contract system
Real-time CDC sourcesA dozen-plusservice DB changes, live
DaaSA dozenfeeding company-wide AI

1. The limits of a request-processing org

Early-stage data platform orgs mostly look the same. A "please load this data into the lakehouse" request arrives as a ticket, an engineer studies the source, defines a schema, writes pipeline code, and configures permissions and retention. Every request gets human hands, end to end.

This works fine at a few requests a month. But as internal AI demand exploded and requests multiplied, something became obvious: building the pipelines was astonishingly repetitive. Whatever the source, the work was the same — define a schema, create topics, deploy an ingestion job, grant permissions, set retention and compaction, register lineage. Only the "values" differed; the "structure" was identical every time. An org that keeps taking repetitive requests as tickets can only grow linearly with headcount. We decided to turn the repetition into a product.

2. The design — separating contract (declaration) from implementation (generation)

The core idea: pipelines are declared as data contracts, not written as code. The party who wants their data in (the Data Owner) prepares exactly two things — a Protobuf schema (the data's structure: fields, types, keys) and a YAML contract file (the data's operational properties: source, ingestion mode, retention, compaction, access permissions). They open a PR with these two files, and after review and merge, everything else is automatic.

What the Data Owner submits (declaration) Protobuf schema structure — fields · types · keys YAML contract file operational properties — source · retention · permissions one PR → review · merge Generation layer (implementation — platform's responsibility) reads the contract, generates and applies the entire infrastructure ETL pipeline ingestion jobs · schema evolution Kafka topics streaming path setup Governance permissions · retention · compaction Lineage auto-registered in catalog A dozen DaaS offerings — abuse detection and more, feeding company-wide AI workloads hundreds of data models · dozens of pipeline contracts · real-time CDC from a dozen-plus sources
Two contract files are the only interface — the Data Owner declares the "what," the platform generates the "how"

The essence of this structure is separating declaration from implementation. When the implementation improves — swap the ingestion engine, change the compaction strategy — the contracts stay untouched and every pipeline improves together. And when the contract spec grows, the entire review surface collapses to a single YAML diff.

Governance lives in the schema, not in documents

Policies like retention, access permissions, and mutability live inside the contract file — right next to the schema — not in documents. Policy docs go stale; contracts can't, because that file is the literal source the real pipeline is generated from. Need an audit? Read the contract files in the repo and their git history.

3. Scale — what the contracts are carrying

What runs on this structure today is what the stat tiles up top show — hundreds of data models, dozens of pipeline contracts plus a separate subscription (access) contract system, real-time CDC from 13+ service sources, and a dozen DaaS offerings supplied to AI workloads across the company.

But the structure matters more than the numbers. While the model count passed 500, the headcount building pipelines barely moved. The proportionality between request volume and org size was severed — and that, I'd argue, is the definition of becoming a platform.

4. Self-service, and automating contract authoring itself

The most tangible effect of the platform shift is in the ingestion process. Before, our org sat across the entire "request → alignment → development → deployment" span. Now, ingestion is complete once the Data Owner submits a schema. We exist only as reviewers.

The last remaining friction was "I don't know how to write the contract file" — and we solved that with AI too. We built an AI skill that assists contract authoring: give it a source schema (DDL, ES mapping, etc.) and it drafts the Protobuf and YAML contract and even opens the PR. AI automation could be precise because it had a structured target — generating schema-validated YAML is far more verifiable than generating free-form code. This combination is what cut external request lead time from 15 days to 3.

5. Lessons

  1. Solve repetitive requests with a product, not tickets. When the same shape of request arrives three times, the fourth should be received by a system, not a person.
  2. Separate the contract (declaration) from the implementation (generation). Once the declaration is stable, you can swap implementations freely, and the user interface collapses to a single YAML diff.
  3. Plant governance in the schema, not in documents. A policy that sits on the execution path cannot go stale. Policy in a document is a recommendation; policy in a contract is code.

Becoming a platform org was as much an identity shift as a technical one. From "the team that processes requests" to "the team that makes requests get processed without us." The former looks productive the busier it gets — but only the latter can absorb demand that grows faster than the org.