Data contracts: how to stop downstream breakages
Write down what a shared dataset promises, and check it where changes are made. Writing the contract is the easy part. The milestone is a breaking change that fails in the producing team’s pipeline, before anyone downstream sees it.
By the Webair engineering team9 min read

A data contract is a written agreement between the team that produces a dataset and the teams that use it: what the data looks like, the quality it should meet, who owns it and where to ask about it.1 Keep it beside the code that produces the data and check every change against it in that team’s pipeline. Then a change that would break someone else’s report can fail there, before it ships.
What a data contract is
The Open Data Contract Standard (ODCS) puts it in one line: “A data contract defines the agreement between a data producer and consumers.”1 The producer is the team whose pipeline writes a table, a file or a stream of events. The consumers are everyone who reads it: dashboards, finance reports, forecasting models, other services and AI features. During a migration, an old and a new system can depend on the same data at once, one of the harder parts of replacing a legacy system piece by piece.
dbt Labs, maker of the dbt transformation tool, puts it in software terms: a dbt model, the code that builds a table, shared with other teams or systems “is operating like an API”.2 Change an API without warning and its clients break. A shared table is the same, except its owner may not know who all its clients are.
dbt’s guidance says a breaking change can be “as obvious as removing or renaming a column, or more subtle, like changing its data type or nullability”.2 Its reference guide gives an example of the subtle kind:3
Even a subtle change in data type, such as from boolean (true/false) to integer (0/1), could cause queries to fail in surprising ways.
None of these sources measures how often this happens, so this article won’t guess.
One word, two meanings. A dbt model contract covers only the shape of the data: columns, types and constraints. dbt leaves data tests out on purpose: by its analogy with an API, the structure of the response is the contract, and quality and reliability are “not part of the contract per se”.4 An ODCS contract also covers quality, service levels, ownership and support.1 When a tool or a vendor says “contract”, check which kind they mean.
What goes in one
The standard’s sections follow the questions a consumer would ask. These matter most:1
- Fundamentals. What the data is for, the contract’s version and its status, such as draft, active or retired.
- Schema. The tables and fields, which the standard calls objects and properties, with their types, their keys and whether a field must always have a value.5
- Data quality. Rules the data must pass, in a form a tool can run.
- Service levels. What the producer commits to, such as how up to date the data is, how often it’s refreshed and how quickly problems are fixed.
- Team and support. Who looks after the data, including its owner, and the channels where consumers get help and announcements go.
- Servers. Where the data lives.
There are more, such as pricing and tags.
Here’s a short one for a shared orders table, written by us. Note the last two fields: one is on its way out, and names its replacement.
# Our example: a contract for a shared orders table, in ODCS v3.2.0.
apiVersion: v3.2.0
kind: DataContract
id: checkout-orders
name: orders
version: 2.1.0
status: active
description:
purpose: One row per completed order, for finance reporting and forecasting.
team:
name: checkout
members:
- username: orders-owner
role: owner
support:
- channel: "#orders-data"
tool: slack
scope: announcements
schema:
- name: orders
logicalType: object
physicalType: table
properties:
- name: order_id
logicalType: string
primaryKey: true
required: true
- name: amount
logicalType: number
deprecated: true
description: "DEPRECATED: use amount_incl_tax. To be removed in 3.0.0."
- name: amount_incl_tax
logicalType: number
required: true
Keep it in the repository that builds the table, so both change in the same pull request.
An open standard to start from
You could write your own template, but a standard has one advantage: tools already read it. ODCS is maintained by Bitol, an LF AI & Data Foundation project, under the Apache 2.0 licence.1 Its JSON Schema lets an editor check a contract as you write it,1 and open-source tools such as the Data Contract CLI test data against it.6
Version 3.2.0, released on 8 September 2026, adds, among other things:7
- Deprecation, for retiring fields without breaking the teams that use them.
- Variables, so one contract can serve development, test and production without keeping secrets in Git.
- Enums, vectors and maps, for fixed lists of values, embeddings and key-value data.
- Context, a block of instructions for AI agents: how to use the data, which statements about it are verified, and what they must never do.
That last one matters once agents use your data. An agent querying a table through an MCP server depends on what the data means as much as a dashboard does, and the contract is one place to write that down.
Tools are still catching up: the Data Contract CLI says support for the new fields is arriving “step by step”,8 so check that yours acts on a field before you rely on it.
The standard gives you a format, not enforcement. On its own, a contract checks nothing, and one that’s never checked goes out of date like any other document.
Check it where changes happen
The place to catch a breaking change is the producing team’s pipeline, before it ships. Once a consumer finds it, it’s already in someone’s report. Four checks can help. The middle column is our advice, not either tool’s rule:
| Check | Where to run it | What it catches |
|---|---|---|
| dbt enforced contract | Every build of the table | Columns whose names or types don’t match the contract. The table isn’t built |
| dbt comparison with the previous state | Every pull request | A removed column, a changed type or constraint, or a deleted or renamed table |
datacontract breaking | Every change to the contract | Changes that could make existing data fail the contract, plus warnings for changes to review |
datacontract test | Every pipeline run, against the real data | Data that doesn’t match the schema, the quality rules or the service levels |
dbt’s checks cover the shape. With a contract enforced, dbt stops before building the table if the columns or types don’t match.3 Comparing with the previous state, it raises a contract error for a removed column or a changed type or constraint. Deleting or renaming a model is an error only if the model is versioned, and otherwise a warning.4 And most platforms only enforce not-null constraints, so a primary key in a contract may be recorded rather than enforced.4
datacontract test connects to the data named in the contract.6 Its dry run works on a pull request with no warehouse access, but it reads no data and always passes: the documentation calls it “a plan, not a verdict”.6 datacontract breaking treats a change that could make existing data fail the contract as an error and fails, so the pipeline can stop on it.9
None of these checks sees a change in meaning. A column recalculated under the same name and type passes the shape checks, yet dbt’s guidance warns this can “substantially change the results seen by downstream queriers”, and calls whether to release it as a new version a judgement call.2 That takes people: the owner tells the consumers before it ships. An AI feature that retrieves the data is a consumer too, and a change in meaning is a reason to rerun its test set.
Writing the YAML is the easy part. The hard part is the agreement behind it: which datasets count as shared, who owns each one, what counts as a breaking change, and who decides when a check fails the day before a release.
Start small, where it counts. dbt recommends contracts for the models other teams rely on, and warns that adding them while models are still changing can make rollbacks harder and add maintenance.4 So pick one dataset several teams depend on, give it a named owner, keep its contract with the code that builds it, and run the checks in that team’s pipeline. Then track one measure: how many breaking changes are caught there, and how many still reach a consumer.
Changing a contract without breaking anyone
A contract doesn’t freeze the data. It makes change deliberate. dbt’s advice is to prefer non-breaking changes, such as adding a column, wherever you can.2 When a breaking change is needed, both tools give consumers time to move.
In ODCS, you flag the old field as deprecated. It stays documented and validated, so consumers can keep using it for now. Tools may warn when it’s used, and its description should point to the replacement.5 That’s what amount does in the example above.
In dbt, you publish a new version of the model beside the old one: build it, make it the latest, set a deprecation date for the old one, update the references downstream, then remove it.2 dbt suggests keeping two or three versions live, not more, and clearing out deprecated columns on a predictable schedule, such as once or twice a year, announced well in advance.2
Versions have a cost too. dbt says making every consumer handle a breaking change at once is the right answer at many smaller organisations, and while models are young and changing fast, but it doesn’t scale well beyond that.2 If a handful of people build every table and read every report, a message may be enough. Once teams can’t see each other’s work, it isn’t.
In practice
Retire a field in three steps. First, add the replacement and mark the old field deprecated, pointing to it. Second, set a removal date, announce it where consumers will see it, and keep both fields filled until then. Third, remove the old field in a new major version of the contract, once you’ve checked who still reads it. The date matters: dbt notes that a clear window puts a known limit on what the migration costs.2
Writing a contract isn’t the milestone. A breaking change that fails in the producing team’s pipeline, before it reaches anyone’s report, is.
What to ask your team
Contracts are judged by where breakages surface: in the producing team’s pipeline, or in someone else’s report. Four answers show where yours surface today:
- which datasets other teams and AI features depend on, and who owns each one
- which of those have a contract, kept beside the code that builds them
- which checks run in the producing team’s pipeline, and who decides when one fails
- how a breaking change is announced, and how long consumers get to move
Then the question to take to them: if the team behind your most-used dataset renamed a column today, who would find out first, that team or the people whose reports broke?
Sources
- Bitol, Open Data Contract Standard v3.2.0: Home, with its Fundamentals, Team, Support and Service-level agreement pages. Published under the LF AI & Data Foundation; Apache 2.0. The pages carry no date.
- dbt Labs, Model versions, dbt Developer Hub, last updated 10 September 2026. dbt Labs sells the dbt platform; model versions are part of dbt itself.
- dbt Labs, contract, resource configuration reference, dbt Developer Hub, last updated 8 September 2026.
- dbt Labs, Model contracts, dbt Developer Hub, last updated 10 September 2026.
- Bitol, Schema, Open Data Contract Standard v3.2.0. The deprecated flag is new in this version.
- Entropy Data, Test your data, Data Contract CLI documentation. The CLI is open source; Entropy Data, which maintains it, sells a data contract platform.
- Bitol, Open Data Contract Standard v3.2.0 release notes, GitHub, 8 September 2026.
- Entropy Data, Open Data Contract Standard, Data Contract CLI documentation.
- Entropy Data, Compare contract versions, Data Contract CLI documentation.


