# CI/CD with GitHub Actions

Test every change automatically and block breaking changes in pull requests.

Source: https://learn.datacontract.com/en/ci-cd/

Your contracts are only tested when someone runs `datacontract test`. Someone forgets, and a broken contract slips through.
In this chapter, GitHub Actions does it for you: every push lints and tests all contracts, and every pull request is checked for **breaking changes** before it is merged.

**You will learn:**
- Version contracts, data products, and SQL in your fork
- Lint ODCS and ODPS files on every push
- Test all contracts against a database in the pipeline
- Block breaking changes in pull requests with `datacontract breaking`

> **Contracts as code**
>
> A contract is a YAML file in Git, so it gets the same treatment as code: reviews, pull requests, and automated checks.
> The pipeline checks three things:
>
> 1. **Syntax:** Is every file valid ODCS or ODPS? (`lint`)
> 2. **Reality:** Does the data match the contract? (`ci`)
> 3. **Compatibility:** Does a change break consumers? (`breaking`)
>
> The earlier a problem is found, the cheaper it is to fix. This is called *shift left*.

## Push your work

### Commit your work to your fork

The workshop repository ignores the files you create in the exercises. In your fork, they are your source code.
Open `.gitignore` and **delete the last block**: the comment `# files created during the exercises` and the five lines below it.

Your views should already be in `sql/sku_sales_input.sql` and `sql/sku_sales_per_year.sql` (from [Consumer-Driven Contracts](https://learn.datacontract.com/en/consumer-driven/)). The pipeline applies them in this order.

Commit and push:

macOS / Linux:

```bash
git add .gitignore sql/ *.odcs.yaml *.odps.yaml
git commit -m "Add data contracts, data products, and views"
git push
```

Windows (PowerShell):

```powershell
git add .gitignore sql/ *.odcs.yaml *.odps.yaml
git commit -m "Add data contracts, data products, and views"
git push
```

Output:

```text
[main 6ee37e4] Add data contracts, data products, and views
 9 files changed, 601 insertions(+)
 create mode 100644 orders.odps.yaml
 create mode 100644 orders_v1.odcs.yaml
 …
 create mode 100644 sql/sku_sales_per_year.sql
Enumerating objects: 14, done.
Counting objects: 100% (14/14), done.
Delta compression using up to 18 threads
Compressing objects: 100% (11/11), done.
Writing objects: 100% (12/12), 5.32 KiB | 5.32 MiB/s, done.
Total 12 (delta 1), reused 0 (delta 0), pack-reused 0 (from 0)
To https://github.com/<your-username>/learn.datacontract.com.git
   18f0f52..6ee37e4  main -> main
```

> **Warning**
>
> Your fork is public. Never commit an API key. Check with `git diff .env` that `.env` only contains the workshop credentials.

### Enable GitHub Actions in your fork

GitHub disables workflows in forks by default.
Open the **Actions** tab of your fork on GitHub and click **I understand my workflows, go ahead and enable them**.

## Lint on every push

### Create the workflow

Create the file `.github/workflows/datacontract.yml`:

```yaml title=.github/workflows/datacontract.yml
name: Data Contracts

on:
  push:
    branches: [main]
  pull_request:

env:
  COLUMNS: 200 # wider tables in the logs

jobs:
  test:
    name: Lint and test
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7

      - uses: astral-sh/setup-uv@v10.2.0

      - name: Install the CLIs
        run: |
          uv tool install --python 3.11 'datacontract-cli[postgres]==1.2.2'
          uv tool install --python 3.11 'dataproduct-cli==0.2.0'

      - name: Lint data contracts
        run: |
          for file in *.odcs.yaml; do
            datacontract lint "$file"
          done

      - name: Lint data products
        run: |
          for file in *.odps.yaml; do
            dataproduct lint "$file"
          done
```

Commit and push it, then open the **Actions** tab. The run turns green after about a minute.

### Break the lint

In `orders_v1.odcs.yaml`, change the `logicalType` of `order_id` to `text`. That is not an ODCS logical type.
Push, and open the failing run. The log names the invalid field and lists the allowed values.

Revert the change and push again.

## Test against the database

Linting checks the syntax. Now check reality: the pipeline starts the workshop database, creates your views, and tests every contract.

### Add the database tests

Add these steps at the end of the `test` job:

```yaml
      - name: Start the database
        run: docker compose up -d --wait

      - name: Create the views
        run: |
          docker compose exec -T postgres psql -U workshop -d workshop -v ON_ERROR_STOP=1 < sql/sku_sales_input.sql
          docker compose exec -T postgres psql -U workshop -d workshop -v ON_ERROR_STOP=1 < sql/sku_sales_per_year.sql

      - name: Test data contracts
        run: datacontract ci *.odcs.yaml
```

`datacontract ci` is `datacontract test` for pipelines. It tests several contracts at once, writes a summary to the run page, and annotates failed checks on the contract file.
The database credentials come from the `.env` file in your repository.

Push, open the run, and scroll down to the **summary**.

![The workflow run in the Actions tab: both jobs, triggered by a push to main](https://learn.datacontract.com/screenshots/ci-run-overview.webp)
![The step summary of datacontract ci: all four contracts passed, with every check listed below](https://learn.datacontract.com/screenshots/ci-step-summary.webp)

### Make a test fail

In `sku_sales_per_year.odcs.yaml`, change the `physicalType` of `total_quantity` to `integer` and push.
The run fails. Find the annotation: it says the column is `bigint`, not `integer`.

Revert and push again.

> **A stand-in for staging and production**
>
> The database in the pipeline is a stand-in. In a real setup, the pipeline tests a staging environment before each deployment.
> A scheduled workflow (`on: schedule:` with a cron expression) tests production regularly, because data can break without any code change.

## Stop breaking changes

A breaking change in a contract breaks consumers. The pipeline should catch it before the merge.
The idea: compare each contract in a pull request with its version on the base branch.

### Try datacontract breaking locally

Compare two contracts on your machine:

macOS / Linux:

```bash
datacontract breaking orders_v1.odcs.yaml orders_v2.odcs.yaml
```

Windows (PowerShell):

```powershell
datacontract breaking orders_v1.odcs.yaml orders_v2.odcs.yaml
```

Output:

```text
Summary
[ 11 Warning ]  [ 6 Info ]
╭──────────┬─────────┬───────────────────────────────────────────────────────────╮
│ Severity │ Change  │ Field                                                     │
├──────────┼─────────┼───────────────────────────────────────────────────────────┤
│ INFO     │ Updated │ id                                                        │
│ WARNING  │ Updated │ schema.line_items.properties.lines_item_id.quality.[1]    │
│ INFO     │ Added   │ schema.line_items.properties.quantity                     │
│ WARNING  │ Updated │ schema.line_items.properties.sku.quality.[1]              │
│ …        │         │                                                           │
│ WARNING  │ Updated │ schema.orders.quality.[2]                                 │
│ INFO     │ Updated │ servers.postgres                                          │
│ INFO     │ Updated │ version                                                   │
╰──────────┴─────────┴───────────────────────────────────────────────────────────╯

Details
…
```

Adding `quantity` is `INFO`. The changed SQL quality queries are `WARNING`s: worth a look, but not breaking. The command exits with code `0`.
Now compare in the other direction, as if you removed `quantity` again:

macOS / Linux:

```bash
datacontract breaking orders_v2.odcs.yaml orders_v1.odcs.yaml
echo $?
```

Windows (PowerShell):

```powershell
datacontract breaking orders_v2.odcs.yaml orders_v1.odcs.yaml
$LASTEXITCODE
```

Output:

```text
Summary
[ 1 Error ]  [ 11 Warning ]  [ 5 Info ]
╭──────────┬─────────┬───────────────────────────────────────────────────────────╮
│ Severity │ Change  │ Field                                                     │
├──────────┼─────────┼───────────────────────────────────────────────────────────┤
│ INFO     │ Updated │ id                                                        │
│ WARNING  │ Updated │ schema.line_items.properties.lines_item_id.quality.[1]    │
│ ERROR    │ Removed │ schema.line_items.properties.quantity                     │
│ …        │         │                                                           │
│ INFO     │ Updated │ version                                                   │
╰──────────┴─────────┴───────────────────────────────────────────────────────────╯

Details
…
1
```

The removed property is an `ERROR`, and the exit code is `1`. That is what fails a pipeline.

### Add the breaking change check

Add a second job to the workflow. It only runs for pull requests:

```yaml
  breaking-changes:
    name: Breaking changes
    if: github.event_name == 'pull_request'
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7
        with:
          fetch-depth: 0 # the base branch is needed for the comparison

      - uses: astral-sh/setup-uv@v10.2.0

      - name: Install the CLI
        run: uv tool install --python 3.11 'datacontract-cli[postgres]==1.2.2'

      - name: Compare with the base branch
        env:
          BASE: origin/${{ github.base_ref }}
        run: |
          status=0
          # every contract on the base branch must stay compatible (new contracts have nothing to compare)
          for file in $(git ls-tree --name-only "$BASE" | grep '\.odcs\.yaml$'); do
            if [ ! -f "$file" ]; then
              echo "::error file=$file::Data contract $file was removed"
              status=1
              continue
            fi
            git show "$BASE:$file" > "$RUNNER_TEMP/$file"
            if ! datacontract breaking "$RUNNER_TEMP/$file" "$file"; then
              echo "::error file=$file::Breaking change in $file. Release a new major version instead."
              status=1
            fi
          done
          exit $status
```

The `::error` lines are GitHub annotations. They show up on the pull request.
Commit and push to `main`.

### Open a pull request with a breaking change

Now you are the orders team, and you want to "clean up" a column. Create a branch and remove the `customer_id` property from the `orders` schema in `orders_v2.odcs.yaml`:

macOS / Linux:

```bash
git switch -c remove-customer-id
# edit orders_v2.odcs.yaml: remove customer_id
git commit -am "Remove customer_id"
git push -u origin remove-customer-id
```

Windows (PowerShell):

```powershell
git switch -c remove-customer-id
# edit orders_v2.odcs.yaml: remove customer_id
git commit -am "Remove customer_id"
git push -u origin remove-customer-id
```

Output:

```text
Switched to a new branch 'remove-customer-id'
[remove-customer-id 86465a7] Remove customer_id
 1 file changed, 10 deletions(-)
Enumerating objects: 5, done.
Counting objects: 100% (5/5), done.
Delta compression using up to 18 threads
Compressing objects: 100% (3/3), done.
Writing objects: 100% (3/3), 305 bytes | 305.00 KiB/s, done.
Total 3 (delta 2), reused 0 (delta 0), pack-reused 0 (from 0)
remote:
remote: Create a pull request for 'remove-customer-id' on GitHub by visiting:
remote:      https://github.com/<your-username>/learn.datacontract.com/pull/new/remove-customer-id
remote:
To https://github.com/<your-username>/learn.datacontract.com.git
 * [new branch]      remove-customer-id -> remove-customer-id
branch 'remove-customer-id' set up to track 'origin/remove-customer-id'.
```

Open a pull request on GitHub.

> **Choose your fork as base**
>
> For a fork, GitHub proposes the original workshop repository as the base. Change **base repository** to `<your-username>/learn.datacontract.com` and base `main`.

The **Breaking changes** check fails, with an annotation on `orders_v2.odcs.yaml`.

![The pull request: the Breaking changes check fails, Lint and test still passes](https://learn.datacontract.com/screenshots/ci-pr-checks.webp)
![The job log: datacontract breaking reports the removed customer_id as an ERROR](https://learn.datacontract.com/screenshots/ci-breaking-log.webp)

### Make a compatible change instead

Restore `customer_id`. Instead, add a `description` to it, or a new tag. Push to the same branch.
The check turns green, and you can merge.

A change that consumers must adapt to needs a new major version: a new contract, like `orders_v2` for `orders_v1` in [Data Contract Evolution](https://learn.datacontract.com/en/evolution/).

**Solution: .github/workflows/datacontract.yml**

```yaml title=datacontract.yml
# Put this file at .github/workflows/datacontract.yml in your fork.
name: Data Contracts

on:
  push:
    branches: [main]
  pull_request:

env:
  # wider tables in the logs of the Data Contract CLI
  COLUMNS: 200

jobs:
  test:
    name: Lint and test
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7

      - uses: astral-sh/setup-uv@v10.2.0

      - name: Install the CLIs
        run: |
          uv tool install --python 3.11 'datacontract-cli[postgres]==1.2.2'
          uv tool install --python 3.11 'dataproduct-cli==0.2.0'

      - name: Lint data contracts
        run: |
          for file in *.odcs.yaml; do
            datacontract lint "$file"
          done

      - name: Lint data products
        run: |
          for file in *.odps.yaml; do
            dataproduct lint "$file"
          done

      - name: Start the database
        run: docker compose up -d --wait

      - name: Create the views
        run: |
          docker compose exec -T postgres psql -U workshop -d workshop -v ON_ERROR_STOP=1 < sql/sku_sales_input.sql
          docker compose exec -T postgres psql -U workshop -d workshop -v ON_ERROR_STOP=1 < sql/sku_sales_per_year.sql

      - name: Test data contracts
        # the database credentials come from the .env file in the repository
        run: datacontract ci *.odcs.yaml

      # Part D: publish the test results to Entropy Data.
      # Add your API key as repository secret ENTROPY_DATA_API_KEY, then replace the step above with:
      #
      # - name: Test data contracts
      #   run: datacontract ci *.odcs.yaml --publish https://api.entropy-data.com/api/test-results
      #   env:
      #     ENTROPY_DATA_API_KEY: ${{ secrets.ENTROPY_DATA_API_KEY }}
      #
      # Contracts linked to semantic concepts (Part D) are resolved on the Entropy Data host:
      # add --no-inline-references if the pipeline can't reach it.

  breaking-changes:
    name: Breaking changes
    if: github.event_name == 'pull_request'
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7
        with:
          fetch-depth: 0 # the base branch is needed for the comparison

      - uses: astral-sh/setup-uv@v10.2.0

      - name: Install the CLI
        run: uv tool install --python 3.11 'datacontract-cli[postgres]==1.2.2'

      - name: Compare with the base branch
        env:
          BASE: origin/${{ github.base_ref }}
        run: |
          status=0
          # every contract on the base branch must stay compatible (new contracts have nothing to compare)
          for file in $(git ls-tree --name-only "$BASE" | grep '\.odcs\.yaml$'); do
            if [ ! -f "$file" ]; then
              echo "::error file=$file::Data contract $file was removed"
              status=1
              continue
            fi
            git show "$BASE:$file" > "$RUNNER_TEMP/$file"
            if ! datacontract breaking "$RUNNER_TEMP/$file" "$file"; then
              echo "::error file=$file::Breaking change in $file. Release a new major version instead."
              status=1
            fi
          done
          exit $status
```

**Quick check:** A pull request removes customer_id from orders_v2, and the column still exists in the database. Which check makes the pipeline fail?

- datacontract lint
- datacontract ci
- datacontract breaking against the base branch (correct)
- None, the contract is still valid

`lint` checks the syntax, and `ci` checks the contract against the data. Only `breaking` compares the new version with the base branch and fails on the removed column (ERROR).

## Bonus

- **Make the checks mandatory:** in your fork, go to **Settings → Rules → Rulesets** and require the checks **Lint and test** and **Breaking changes** for `main`. Now a breaking change can't be merged.
- **Consumer-driven contracts in the producer's pipeline:** your consumer contract `orders_v2.consumer_sku_sales.odcs.yaml` runs in the same pipeline. In a real setup, the orders team runs the contracts of all its consumers. Then a failing check shows exactly whom a change would break.
- **Read the changelog:** `datacontract changelog orders_v1.odcs.yaml orders_v2.odcs.yaml` lists all changes, not just the breaking ones.
- **Publish test results:** in [Part D](https://learn.datacontract.com/en/publish/), you can publish the results to Entropy Data from the pipeline. The reference workflow in the solution contains the step as a comment.

### Bonus: Schedule the tests with Airflow

CI tests a contract when it changes. The data changes every day, so production tests also need a schedule. If your pipelines run in Apache Airflow, the [Data Contract provider for Airflow](https://github.com/datacontract/airflow-provider-datacontract) adds a `DataContractTestOperator`: it runs `datacontract test` as a quality gate in your DAG and fails the task when the contract is violated, so bad data stops before it reaches downstream tasks.

```python title=dags/orders_contract.py
from datetime import datetime
from airflow.sdk import dag
from datacontract_provider.operators.datacontract import DataContractTestOperator

@dag(schedule="0 2 * * *", start_date=datetime(2026, 1, 1), catchup=False)
def orders_contract():
    DataContractTestOperator(
        task_id="test_orders_v2",
        data_contract_file="orders_v2.odcs.yaml",
        server="Orders",
        server_conn_id="workshop_postgres",  # Airflow connection with the database credentials
    )

orders_contract()
```

Install it with `pip install "airflow-provider-datacontract[postgres]"`. See the [scheduling docs for Airflow](https://docs.datacontract.com/scheduling/airflow) of the Data Contract CLI for connections, publishing results, and the results view. No Airflow? A `schedule:` trigger in GitHub Actions does the same, see the [scheduling docs for GitHub Actions](https://docs.datacontract.com/scheduling/github-actions).
