Data contracts: how to stop downstream breakages

Write down what a shared dataset promises, and check it where changes are made. Writing the contract is the easy part. The milestone is a breaking change that fails in the producing team’s pipeline, before anyone downstream sees it.

By the Webair engineering team9 min read

A contract with the first stroke of a coral signature on its signature line, about to be signed

A data contract is a written agreement between the team that produces a dataset and the teams that use it: what the data looks like, the quality it should meet, who owns it and where to ask about it.1 Keep it beside the code that produces the data and check every change against it in that team’s pipeline. Then a change that would break someone else’s report can fail there, before it ships.

What a data contract is

The Open Data Contract Standard (ODCS) puts it in one line: “A data contract defines the agreement between a data producer and consumers.”1 The producer is the team whose pipeline writes a table, a file or a stream of events. The consumers are everyone who reads it: dashboards, finance reports, forecasting models, other services and AI features. During a migration, an old and a new system can depend on the same data at once, one of the harder parts of replacing a legacy system piece by piece.

dbt Labs, maker of the dbt transformation tool, puts it in software terms: a dbt model, the code that builds a table, shared with other teams or systems “is operating like an API”.2 Change an API without warning and its clients break. A shared table is the same, except its owner may not know who all its clients are.

dbt’s guidance says a breaking change can be “as obvious as removing or renaming a column, or more subtle, like changing its data type or nullability”.2 Its reference guide gives an example of the subtle kind:3

Even a subtle change in data type, such as from boolean (true/false) to integer (0/1), could cause queries to fail in surprising ways.

dbt Labs, contract configuration reference

None of these sources measures how often this happens, so this article won’t guess.

One word, two meanings. A dbt model contract covers only the shape of the data: columns, types and constraints. dbt leaves data tests out on purpose: by its analogy with an API, the structure of the response is the contract, and quality and reliability are “not part of the contract per se”.4 An ODCS contract also covers quality, service levels, ownership and support.1 When a tool or a vendor says “contract”, check which kind they mean.

What goes in one

The standard’s sections follow the questions a consumer would ask. These matter most:1

  • Fundamentals. What the data is for, the contract’s version and its status, such as draft, active or retired.
  • Schema. The tables and fields, which the standard calls objects and properties, with their types, their keys and whether a field must always have a value.5
  • Data quality. Rules the data must pass, in a form a tool can run.
  • Service levels. What the producer commits to, such as how up to date the data is, how often it’s refreshed and how quickly problems are fixed.
  • Team and support. Who looks after the data, including its owner, and the channels where consumers get help and announcements go.
  • Servers. Where the data lives.

There are more, such as pricing and tags.

Here’s a short one for a shared orders table, written by us. Note the last two fields: one is on its way out, and names its replacement.

# Our example: a contract for a shared orders table, in ODCS v3.2.0.
apiVersion: v3.2.0
kind: DataContract
id: checkout-orders
name: orders
version: 2.1.0
status: active
description:
  purpose: One row per completed order, for finance reporting and forecasting.
team:
  name: checkout
  members:
    - username: orders-owner
      role: owner
support:
  - channel: "#orders-data"
    tool: slack
    scope: announcements
schema:
  - name: orders
    logicalType: object
    physicalType: table
    properties:
      - name: order_id
        logicalType: string
        primaryKey: true
        required: true
      - name: amount
        logicalType: number
        deprecated: true
        description: "DEPRECATED: use amount_incl_tax. To be removed in 3.0.0."
      - name: amount_incl_tax
        logicalType: number
        required: true

Keep it in the repository that builds the table, so both change in the same pull request.

An open standard to start from

You could write your own template, but a standard has one advantage: tools already read it. ODCS is maintained by Bitol, an LF AI & Data Foundation project, under the Apache 2.0 licence.1 Its JSON Schema lets an editor check a contract as you write it,1 and open-source tools such as the Data Contract CLI test data against it.6

Version 3.2.0, released on 8 September 2026, adds, among other things:7

  • Deprecation, for retiring fields without breaking the teams that use them.
  • Variables, so one contract can serve development, test and production without keeping secrets in Git.
  • Enums, vectors and maps, for fixed lists of values, embeddings and key-value data.
  • Context, a block of instructions for AI agents: how to use the data, which statements about it are verified, and what they must never do.

That last one matters once agents use your data. An agent querying a table through an MCP server depends on what the data means as much as a dashboard does, and the contract is one place to write that down.

Tools are still catching up: the Data Contract CLI says support for the new fields is arriving “step by step”,8 so check that yours acts on a field before you rely on it.

The standard gives you a format, not enforcement. On its own, a contract checks nothing, and one that’s never checked goes out of date like any other document.

Check it where changes happen

The place to catch a breaking change is the producing team’s pipeline, before it ships. Once a consumer finds it, it’s already in someone’s report. Four checks can help. The middle column is our advice, not either tool’s rule:

Four checks, where we’d run them, and what each one catches
CheckWhere to run itWhat it catches
dbt enforced contractEvery build of the tableColumns whose names or types don’t match the contract. The table isn’t built
dbt comparison with the previous stateEvery pull requestA removed column, a changed type or constraint, or a deleted or renamed table
datacontract breakingEvery change to the contractChanges that could make existing data fail the contract, plus warnings for changes to review
datacontract testEvery pipeline run, against the real dataData that doesn’t match the schema, the quality rules or the service levels

dbt’s checks cover the shape. With a contract enforced, dbt stops before building the table if the columns or types don’t match.3 Comparing with the previous state, it raises a contract error for a removed column or a changed type or constraint. Deleting or renaming a model is an error only if the model is versioned, and otherwise a warning.4 And most platforms only enforce not-null constraints, so a primary key in a contract may be recorded rather than enforced.4

datacontract test connects to the data named in the contract.6 Its dry run works on a pull request with no warehouse access, but it reads no data and always passes: the documentation calls it “a plan, not a verdict”.6 datacontract breaking treats a change that could make existing data fail the contract as an error and fails, so the pipeline can stop on it.9

None of these checks sees a change in meaning. A column recalculated under the same name and type passes the shape checks, yet dbt’s guidance warns this can “substantially change the results seen by downstream queriers”, and calls whether to release it as a new version a judgement call.2 That takes people: the owner tells the consumers before it ships. An AI feature that retrieves the data is a consumer too, and a change in meaning is a reason to rerun its test set.

Writing the YAML is the easy part. The hard part is the agreement behind it: which datasets count as shared, who owns each one, what counts as a breaking change, and who decides when a check fails the day before a release.

Start small, where it counts. dbt recommends contracts for the models other teams rely on, and warns that adding them while models are still changing can make rollbacks harder and add maintenance.4 So pick one dataset several teams depend on, give it a named owner, keep its contract with the code that builds it, and run the checks in that team’s pipeline. Then track one measure: how many breaking changes are caught there, and how many still reach a consumer.

Changing a contract without breaking anyone

A contract doesn’t freeze the data. It makes change deliberate. dbt’s advice is to prefer non-breaking changes, such as adding a column, wherever you can.2 When a breaking change is needed, both tools give consumers time to move.

In ODCS, you flag the old field as deprecated. It stays documented and validated, so consumers can keep using it for now. Tools may warn when it’s used, and its description should point to the replacement.5 That’s what amount does in the example above.

In dbt, you publish a new version of the model beside the old one: build it, make it the latest, set a deprecation date for the old one, update the references downstream, then remove it.2 dbt suggests keeping two or three versions live, not more, and clearing out deprecated columns on a predictable schedule, such as once or twice a year, announced well in advance.2

Versions have a cost too. dbt says making every consumer handle a breaking change at once is the right answer at many smaller organisations, and while models are young and changing fast, but it doesn’t scale well beyond that.2 If a handful of people build every table and read every report, a message may be enough. Once teams can’t see each other’s work, it isn’t.

In practice

Retire a field in three steps. First, add the replacement and mark the old field deprecated, pointing to it. Second, set a removal date, announce it where consumers will see it, and keep both fields filled until then. Third, remove the old field in a new major version of the contract, once you’ve checked who still reads it. The date matters: dbt notes that a clear window puts a known limit on what the migration costs.2

Writing a contract isn’t the milestone. A breaking change that fails in the producing team’s pipeline, before it reaches anyone’s report, is.

What to ask your team

Contracts are judged by where breakages surface: in the producing team’s pipeline, or in someone else’s report. Four answers show where yours surface today:

  • which datasets other teams and AI features depend on, and who owns each one
  • which of those have a contract, kept beside the code that builds them
  • which checks run in the producing team’s pipeline, and who decides when one fails
  • how a breaking change is announced, and how long consumers get to move

Then the question to take to them: if the team behind your most-used dataset renamed a column today, who would find out first, that team or the people whose reports broke?

Sources

  1. Bitol, Open Data Contract Standard v3.2.0: Home, with its Fundamentals, Team, Support and Service-level agreement pages. Published under the LF AI & Data Foundation; Apache 2.0. The pages carry no date.
  2. dbt Labs, Model versions, dbt Developer Hub, last updated 10 September 2026. dbt Labs sells the dbt platform; model versions are part of dbt itself.
  3. dbt Labs, contract, resource configuration reference, dbt Developer Hub, last updated 8 September 2026.
  4. dbt Labs, Model contracts, dbt Developer Hub, last updated 10 September 2026.
  5. Bitol, Schema, Open Data Contract Standard v3.2.0. The deprecated flag is new in this version.
  6. Entropy Data, Test your data, Data Contract CLI documentation. The CLI is open source; Entropy Data, which maintains it, sells a data contract platform.
  7. Bitol, Open Data Contract Standard v3.2.0 release notes, GitHub, 8 September 2026.
  8. Entropy Data, Open Data Contract Standard, Data Contract CLI documentation.
  9. Entropy Data, Compare contract versions, Data Contract CLI documentation.
  1. Legacy modernisation

    The strangler fig pattern: modernise without a rewrite

    Replace an old system one piece at a time while it keeps running. Moving the first piece is the easy part. Switching the old one off is the milestone.

    9 min read

    Railway points with the coral switch rail standing open, about to move across and send the track to the new line
  2. AI agents

    MCP explained: connecting AI agents to your systems

    The Model Context Protocol gives AI agents one standard way to reach your tools and data. What an agent may do once it’s connected is still your decision.

    8 min read

    A plug lined up with a socket, about to slide in
  3. AI product engineering

    How to test an AI feature before it launches

    A demo shows that an AI feature can work. Before launch, you need evidence of how often it does, on the cases your users will bring.

    8 min read

    A balance scale with its coral pan raised, about to settle level

Working on something similar?

Tell us what you’re building or fixing, and we’ll share how we’d approach it.