Data Contracts in Practice
Progress
0%
ende
Getting Started
  • Welcome15′
  • Setup20′
Part A · The Source Data Product
  • 1.Put Your Data Under Contract60′
  • 2.Data Contract Evolution30′
  • 3.Describe Your Data Product20′
Part B · The Consumer-Aligned Data Product
  • 4.Design Contract-First30′
  • 5.Implement Your Data Product25′
  • 6.Consumer-Driven Contracts30′
Part C · Automate
  • 7.CI/CD with GitHub Actions45′
Part D · Data Platformoptional
  • 8.Publish to Entropy Data40′
  • 9.Semantics25′
Wrap-up
  • Wrap-up10′
Part C · Automate

Exercise 7 · CI/CD with GitHub Actions

Test every change automatically and block breaking changes in pull requests.

~45 min0 of 10 steps done
Previous
Consumer-Driven Contracts
Next
Publish to Entropy Data
Maintained byEntropy Data

Your contracts are only tested when someone runs datacontract test. Someone forgets, and a broken contract slips through. In this chapter, GitHub Actions does it for you: every push lints and tests all contracts, and every pull request is checked for breaking changes before it is merged.

You will learn
  • Version contracts, data products, and SQL in your fork
  • Lint ODCS and ODPS files on every push
  • Test all contracts against a database in the pipeline
  • Block breaking changes in pull requests with datacontract breaking
Contracts as code

A contract is a YAML file in Git, so it gets the same treatment as code: reviews, pull requests, and automated checks. The pipeline checks three things:

  1. Syntax: Is every file valid ODCS or ODPS? (lint)
  2. Reality: Does the data match the contract? (ci)
  3. Compatibility: Does a change break consumers? (breaking)

The earlier a problem is found, the cheaper it is to fix. This is called shift left.

Push your work

Commit your work to your fork

The workshop repository ignores the files you create in the exercises. In your fork, they are your source code. Open .gitignore and delete the last block: the comment # files created during the exercises and the five lines below it.

Your views should already be in sql/sku_sales_input.sql and sql/sku_sales_per_year.sql (from Consumer-Driven Contracts). The pipeline applies them in this order.

Commit and push:

git add .gitignore sql/ *.odcs.yaml *.odps.yaml
git commit -m "Add data contracts, data products, and views"
git push
[main 6ee37e4] Add data contracts, data products, and views
 9 files changed, 601 insertions(+)
 create mode 100644 orders.odps.yaml
 create mode 100644 orders_v1.odcs.yaml
 …
 create mode 100644 sql/sku_sales_per_year.sql
Enumerating objects: 14, done.
Counting objects: 100% (14/14), done.
Delta compression using up to 18 threads
Compressing objects: 100% (11/11), done.
Writing objects: 100% (12/12), 5.32 KiB | 5.32 MiB/s, done.
Total 12 (delta 1), reused 0 (delta 0), pack-reused 0 (from 0)
To https://github.com/<your-username>/learn.datacontract.com.git
   18f0f52..6ee37e4  main -> main
Warning

Your fork is public. Never commit an API key. Check with git diff .env that .env only contains the workshop credentials.

Enable GitHub Actions in your fork

GitHub disables workflows in forks by default. Open the Actions tab of your fork on GitHub and click I understand my workflows, go ahead and enable them.

Lint on every push

Create the workflow

Create the file .github/workflows/datacontract.yml:

.github/workflows/datacontract.yml
name: Data Contracts

on:
  push:
    branches: [main]
  pull_request:

env:
  COLUMNS: 200 # wider tables in the logs

jobs:
  test:
    name: Lint and test
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7

      - uses: astral-sh/[email protected]

      - name: Install the CLIs
        run: |
          uv tool install --python 3.11 'datacontract-cli[postgres]==1.2.2'
          uv tool install --python 3.11 'dataproduct-cli==0.2.0'

      - name: Lint data contracts
        run: |
          for file in *.odcs.yaml; do
            datacontract lint "$file"
          done

      - name: Lint data products
        run: |
          for file in *.odps.yaml; do
            dataproduct lint "$file"
          done

Commit and push it, then open the Actions tab. The run turns green after about a minute.

Break the lint

In orders_v1.odcs.yaml, change the logicalType of order_id to text. That is not an ODCS logical type. Push, and open the failing run. The log names the invalid field and lists the allowed values.

Revert the change and push again.

Test against the database

Linting checks the syntax. Now check reality: the pipeline starts the workshop database, creates your views, and tests every contract.

Add the database tests

Add these steps at the end of the test job:

YAML
      - name: Start the database
        run: docker compose up -d --wait

      - name: Create the views
        run: |
          docker compose exec -T postgres psql -U workshop -d workshop -v ON_ERROR_STOP=1 < sql/sku_sales_input.sql
          docker compose exec -T postgres psql -U workshop -d workshop -v ON_ERROR_STOP=1 < sql/sku_sales_per_year.sql

      - name: Test data contracts
        run: datacontract ci *.odcs.yaml

datacontract ci is datacontract test for pipelines. It tests several contracts at once, writes a summary to the run page, and annotates failed checks on the contract file. The database credentials come from the .env file in your repository.

Push, open the run, and scroll down to the summary.

The workflow run in the Actions tab: both jobs, triggered by a push to main
The step summary of datacontract ci: all four contracts passed, with every check listed below

Make a test fail

In sku_sales_per_year.odcs.yaml, change the physicalType of total_quantity to integer and push. The run fails. Find the annotation: it says the column is bigint, not integer.

Revert and push again.

A stand-in for staging and production

The database in the pipeline is a stand-in. In a real setup, the pipeline tests a staging environment before each deployment. A scheduled workflow (on: schedule: with a cron expression) tests production regularly, because data can break without any code change.

Stop breaking changes

A breaking change in a contract breaks consumers. The pipeline should catch it before the merge. The idea: compare each contract in a pull request with its version on the base branch.

Try datacontract breaking locally

Compare two contracts on your machine:

datacontract breaking orders_v1.odcs.yaml orders_v2.odcs.yaml
Summary
[ 11 Warning ]  [ 6 Info ]
╭──────────┬─────────┬───────────────────────────────────────────────────────────╮
│ Severity │ Change  │ Field                                                     │
├──────────┼─────────┼───────────────────────────────────────────────────────────┤
│ INFO     │ Updated │ id                                                        │
│ WARNING  │ Updated │ schema.line_items.properties.lines_item_id.quality.[1]    │
│ INFO     │ Added   │ schema.line_items.properties.quantity                     │
│ WARNING  │ Updated │ schema.line_items.properties.sku.quality.[1]              │
│ …        │         │                                                           │
│ WARNING  │ Updated │ schema.orders.quality.[2]                                 │
│ INFO     │ Updated │ servers.postgres                                          │
│ INFO     │ Updated │ version                                                   │
╰──────────┴─────────┴───────────────────────────────────────────────────────────╯

Details
…

Adding quantity is INFO. The changed SQL quality queries are WARNINGs: worth a look, but not breaking. The command exits with code 0. Now compare in the other direction, as if you removed quantity again:

datacontract breaking orders_v2.odcs.yaml orders_v1.odcs.yaml
echo $?
Summary
[ 1 Error ]  [ 11 Warning ]  [ 5 Info ]
╭──────────┬─────────┬───────────────────────────────────────────────────────────╮
│ Severity │ Change  │ Field                                                     │
├──────────┼─────────┼───────────────────────────────────────────────────────────┤
│ INFO     │ Updated │ id                                                        │
│ WARNING  │ Updated │ schema.line_items.properties.lines_item_id.quality.[1]    │
│ ERROR    │ Removed │ schema.line_items.properties.quantity                     │
│ …        │         │                                                           │
│ INFO     │ Updated │ version                                                   │
╰──────────┴─────────┴───────────────────────────────────────────────────────────╯

Details
…
1

The removed property is an ERROR, and the exit code is 1. That is what fails a pipeline.

Add the breaking change check

Add a second job to the workflow. It only runs for pull requests:

YAML
  breaking-changes:
    name: Breaking changes
    if: github.event_name == 'pull_request'
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7
        with:
          fetch-depth: 0 # the base branch is needed for the comparison

      - uses: astral-sh/[email protected]

      - name: Install the CLI
        run: uv tool install --python 3.11 'datacontract-cli[postgres]==1.2.2'

      - name: Compare with the base branch
        env:
          BASE: origin/${{ github.base_ref }}
        run: |
          status=0
          # every contract on the base branch must stay compatible (new contracts have nothing to compare)
          for file in $(git ls-tree --name-only "$BASE" | grep '\.odcs\.yaml
#x27;); do
if [ ! -f "$file" ]; then echo "::error file=$file::Data contract $file was removed" status=1 continue fi git show "$BASE:$file" > "$RUNNER_TEMP/$file" if ! datacontract breaking "$RUNNER_TEMP/$file" "$file"; then echo "::error file=$file::Breaking change in $file. Release a new major version instead." status=1 fi done exit $status

The ::error lines are GitHub annotations. They show up on the pull request. Commit and push to main.

Open a pull request with a breaking change

Now you are the orders team, and you want to "clean up" a column. Create a branch and remove the customer_id property from the orders schema in orders_v2.odcs.yaml:

git switch -c remove-customer-id
# edit orders_v2.odcs.yaml: remove customer_id
git commit -am "Remove customer_id"
git push -u origin remove-customer-id
Switched to a new branch 'remove-customer-id'
[remove-customer-id 86465a7] Remove customer_id
 1 file changed, 10 deletions(-)
Enumerating objects: 5, done.
Counting objects: 100% (5/5), done.
Delta compression using up to 18 threads
Compressing objects: 100% (3/3), done.
Writing objects: 100% (3/3), 305 bytes | 305.00 KiB/s, done.
Total 3 (delta 2), reused 0 (delta 0), pack-reused 0 (from 0)
remote:
remote: Create a pull request for 'remove-customer-id' on GitHub by visiting:
remote:      https://github.com/<your-username>/learn.datacontract.com/pull/new/remove-customer-id
remote:
To https://github.com/<your-username>/learn.datacontract.com.git
 * [new branch]      remove-customer-id -> remove-customer-id
branch 'remove-customer-id' set up to track 'origin/remove-customer-id'.

Open a pull request on GitHub.

Choose your fork as base

For a fork, GitHub proposes the original workshop repository as the base. Change base repository to <your-username>/learn.datacontract.com and base main.

The Breaking changes check fails, with an annotation on orders_v2.odcs.yaml.

The pull request: the Breaking changes check fails, Lint and test still passes
The job log: datacontract breaking reports the removed customer_id as an ERROR

Make a compatible change instead

Restore customer_id. Instead, add a description to it, or a new tag. Push to the same branch. The check turns green, and you can merge.

A change that consumers must adapt to needs a new major version: a new contract, like orders_v2 for orders_v1 in Data Contract Evolution.

Quick check
A pull request removes customer_id from orders_v2, and the column still exists in the database. Which check makes the pipeline fail?

Bonus

  • Make the checks mandatory: in your fork, go to Settings → Rules → Rulesets and require the checks Lint and test and Breaking changes for main. Now a breaking change can't be merged.
  • Consumer-driven contracts in the producer's pipeline: your consumer contract orders_v2.consumer_sku_sales.odcs.yaml runs in the same pipeline. In a real setup, the orders team runs the contracts of all its consumers. Then a failing check shows exactly whom a change would break.
  • Read the changelog: datacontract changelog orders_v1.odcs.yaml orders_v2.odcs.yaml lists all changes, not just the breaking ones.
  • Publish test results: in Part D, you can publish the results to Entropy Data from the pipeline. The reference workflow in the solution contains the step as a comment.

Bonus: Schedule the tests with Airflow

CI tests a contract when it changes. The data changes every day, so production tests also need a schedule. If your pipelines run in Apache Airflow, the Data Contract provider for Airflow adds a DataContractTestOperator: it runs datacontract test as a quality gate in your DAG and fails the task when the contract is violated, so bad data stops before it reaches downstream tasks.

dags/orders_contract.py
from datetime import datetime
from airflow.sdk import dag
from datacontract_provider.operators.datacontract import DataContractTestOperator


@dag(schedule="0 2 * * *", start_date=datetime(2026, 1, 1), catchup=False)
def orders_contract():
    DataContractTestOperator(
        task_id="test_orders_v2",
        data_contract_file="orders_v2.odcs.yaml",
        server="Orders",
        server_conn_id="workshop_postgres",  # Airflow connection with the database credentials
    )


orders_contract()

Install it with pip install "airflow-provider-datacontract[postgres]". See the scheduling docs for Airflow of the Data Contract CLI for connections, publishing results, and the results view. No Airflow? A schedule: trigger in GitHub Actions does the same, see the scheduling docs for GitHub Actions.