Simor
Data pipeline testing strategy: unit, integration, and contract tests

Data pipeline testing strategy: unit, integration, and contract tests

Simor Consulting | 27 Sep, 2026 | 05 Mins read

Data pipelines break in production more often than they should, and the breakage is expensive. A pipeline that silently produces wrong data for three days before anyone notices has corrupted downstream reports, dashboards, and decisions. The fix is not better monitoring alone. The fix is testing: the same discipline that software engineering applies to application code, applied to data transformations.

Most data teams have no tests. Some have a few ad hoc checks. Very few have a testing strategy. This post gives you a three-layer strategy: unit tests for transformations, integration tests for data flows, and contract tests for schema compliance. Each layer catches different failure modes and has different cost and speed characteristics.

Why Data Pipelines Need Different Testing

Application code has well-defined inputs and outputs. A function takes arguments and returns a value. Testing is straightforward: supply inputs, check outputs. Data pipelines have characteristics that make testing harder.

First, the inputs are entire datasets, not individual arguments. Testing a transformation requires representative data, which means managing test fixtures that are large enough to exercise edge cases but small enough to run quickly. A test that takes 20 minutes is a test nobody runs before committing code.

Second, the outputs are not always deterministic. Aggregations, joins, and window functions can produce different results depending on data ordering, partition boundaries, and late-arriving data. Tests need to account for this non-determinism without becoming so loose that they pass when they should fail.

Third, data pipelines are chains of transformations where a failure in one stage propagates downstream. A schema change in the source system breaks the ingestion step, which breaks the transformation, which breaks the output table, which breaks the dashboard. Testing each stage in isolation is necessary but not sufficient. You also need tests that verify the chain works end to end.

Layer 1: Unit Tests for Transformations

Unit tests validate individual transformation logic. Given this input dataset, does the transformation produce the expected output? The input dataset is a test fixture: a small, curated dataset designed to exercise the transformation’s logic.

What to Test

Every transformation has invariants. A deduplication step should reduce row count or leave it unchanged. A join should not produce more rows than the cartesian product. A filter should produce fewer rows than its input. A type cast should not produce nulls from non-null inputs. These invariants are your unit tests.

Test the happy path: representative data that exercises the main logic. Test edge cases: empty inputs, null values, duplicate keys, boundary dates, maximum-length strings. Test the failure mode: what happens when the input violates an assumption the transformation makes?

How to Build Test Fixtures

Create fixtures by sampling production data and anonymising it. A fixture of 100-1,000 rows is enough to exercise most transformation logic. The fixture should include edge cases deliberately, not by random sampling alone. Random sampling misses rare edge cases. Deliberate inclusion covers them.

Store fixtures alongside the transformation code. When the transformation changes, update the fixture if the expected output changes. If the fixture is stored separately from the code, they drift apart and the tests become unreliable.

Speed Requirement

Unit tests should run in seconds, not minutes. If a unit test takes more than 30 seconds, the fixture is too large or the transformation is doing too much. Split the transformation and test each piece separately.

Layer 2: Integration Tests for Data Flows

Integration tests validate that stages of the pipeline work together. They test data flow from source to destination through the transformation chain. The input is a test source system or a replay of production data. The output is checked at the destination.

What to Test

Row counts between stages. If stage one produces 10,000 rows and stage two filters to 8,000 rows, the integration test should verify both counts. Unexpected changes in row counts are the most common signal of pipeline breakage.

Data freshness. If the pipeline should produce daily outputs, the integration test checks that today’s output exists by the expected time. Latency violations are integration failures.

End-to-end transformations. Take a known input record, run it through the full pipeline, and verify the output record matches expectations. This catches issues that unit tests miss: interaction effects between transformations, partition boundary problems, and serialisation issues.

How to Run Integration Tests

Integration tests need a test environment that mirrors production configuration. This does not mean a full production replica. It means the same orchestration tool, the same compute configuration, and the same connection patterns. If your production pipeline runs on Airflow with Spark, your integration tests should run on Airflow with Spark, even if the Spark cluster is smaller.

Run integration tests on a schedule: nightly at minimum, after every deployment at best. Integration tests are slower than unit tests. Accept that and do not try to make them fast enough to run on every commit. Run unit tests on every commit. Run integration tests nightly and on deployment.

Layer 3: Contract Tests for Schema Compliance

Contract tests validate that data producers and consumers agree on the schema. The producer promises that the data will have these columns, these types, and these constraints. The consumer promises that it expects exactly that. The contract test verifies both sides.

What to Test

Schema presence and types. Every column the consumer expects must exist with the expected type. A column that changes from integer to string breaks downstream aggregations. A column that disappears breaks queries.

Nullability constraints. If the consumer assumes a column is non-null and the producer allows nulls, the consumer will encounter nulls it cannot handle. The contract should specify which columns are nullable and which are not.

Value range constraints. If a status column should contain only “active,” “inactive,” and “pending,” the contract should enforce this. Values outside the expected range indicate either a producer change or data corruption.

Freshness contracts. The producer promises that data will be available by a certain time. The consumer promises to read it within a certain window. Violations of either side are contract failures.

How to Implement

Define contracts as data schemas with constraints. Store them in a shared location accessible to both producers and consumers. Run contract tests at the boundary between producer and consumer, in the ingestion step, before the data enters the consumer’s pipeline.

When a producer changes their schema, the contract test fails. The failure forces a conversation: update the contract (and the consumer), or fix the producer. Without the contract test, the schema change propagates silently and breaks the consumer in production.

The Testing Pyramid for Data

This diagram requires JavaScript.

Enable JavaScript in your browser to use this feature.

You should have many unit tests, fewer integration tests, and enough contract tests to cover every producer-consumer boundary. Unit tests catch logic errors. Integration tests catch flow errors. Contract tests catch interface errors. Together, they cover the failure modes that cause silent data corruption.

Next Step

Pick your most critical pipeline. Write unit tests for its three most complex transformations. Write one integration test that verifies row counts end to end. Write one contract test for the output schema. Run them. Fix what breaks. You now have a testing foundation that you can extend to the rest of your pipelines.

Shipping a production AI system?

Find where your AI spend leaks and where quality slips. Take the AI Production Scorecard for a fast baseline across the seven layers, or book a free AI cost review and we will turn it into a plan.

Similar Articles

5 AI Workflows Professional Services Firms Can Deploy This Quarter
5 AI Workflows Professional Services Firms Can Deploy This Quarter
10 Jul, 2026 | 12 Mins read

Professional services firms sell judgment, billed by the hour or by the matter. That makes them both the biggest winners and the most cautious adopters of AI. The upside is real: every firm carries ho

Legacy Data Pipeline Modernisation Without Rewriting Everything
Legacy Data Pipeline Modernisation Without Rewriting Everything
10 Jul, 2026 | 10 Mins read

The pipeline runs every night at 2 a.m. Nobody fully understands it. The original author left in 2019. It is part SAS, part shell, part stored procedures, and part a spreadsheet someone emails in. It

Lightweight MLOps for Mid-Market Teams: Ship Models Without a Platform Engineering Org
Lightweight MLOps for Mid-Market Teams: Ship Models Without a Platform Engineering Org
10 Jul, 2026 | 11 Mins read

A head of ML at a 120-person company told us recently that his team had spent nine months trying to stand up a "proper MLOps platform." They had evaluated three orchestration tools, designed a feature

Building AI-Ready Data Pipelines: Key Architecture Considerations
Building AI-Ready Data Pipelines: Key Architecture Considerations
04 Mar, 2025 | 02 Mins read

Data pipelines built for business intelligence often fail when supporting AI workloads. The root cause is usually architectural: BI pipelines assume bounded, relatively static datasets, while AI syste

Anatomy of an AI Incident: Post-Mortem of a Model Provider Outage
Anatomy of an AI Incident: Post-Mortem of a Model Provider Outage
19 Jun, 2026 | 09 Mins read

On a Tuesday at 2:14 PM, a major model provider began returning elevated error rates for a specific model endpoint. By 2:31 PM, a customer support platform that depended on that endpoint was producing

AI Rollback Patterns: When to Roll Back a Prompt, a Model, or the Whole Release
AI Rollback Patterns: When to Roll Back a Prompt, a Model, or the Whole Release
27 Jun, 2026 | 11 Mins read

Software rollbacks are well-understood. You deploy a new version, detect an issue, and roll back to the previous version. The rollback is atomic: the entire application reverts to the previous state.

The 7-step vector database selection checklist
The 7-step vector database selection checklist
26 Apr, 2026 | 06 Mins read

Most vector database selection failures come down to one mistake: picking the technology before mapping the workload. Teams benchmark embedding search speed on a curated dataset, pick the fastest opti

Build vs buy: a decision tree for AI infrastructure
Build vs buy: a decision tree for AI infrastructure
03 May, 2026 | 06 Mins read

Every AI infrastructure team eventually faces the same argument. One faction wants to build a custom solution because the commercial options do not handle their specific requirements. The other factio

How to design a prompt ops pipeline from scratch
How to design a prompt ops pipeline from scratch
10 May, 2026 | 06 Mins read

Prompt management in most AI teams starts the same way. One engineer writes a prompt, it works well enough, and the prompt gets committed to a config file. Three months later, there are forty prompts

The data quality scorecard: metrics that actually matter
The data quality scorecard: metrics that actually matter
17 May, 2026 | 06 Mins read

Most data quality initiatives fail not because teams lack tools, but because they measure the wrong things. Teams track hundreds of data quality metrics, generate dashboards full of green indicators,

Conference report: key takeaways from Data Council 2026
Conference report: key takeaways from Data Council 2026
23 May, 2026 | 04 Mins read

Data Council 2026 wrapped in Austin last week, and the signal-to-noise ratio was higher than in recent years. The conference has historically been the venue where data infrastructure practitioners (no

A cost optimisation framework for LLM inference
A cost optimisation framework for LLM inference
24 May, 2026 | 06 Mins read

LLM inference costs follow a pattern that catches teams off guard. The first prototype costs almost nothing: a few hundred dollars a month during development. The pilot scales to a few thousand. Produ

Migration playbook: batch to streaming in 5 phases
Migration playbook: batch to streaming in 5 phases
31 May, 2026 | 06 Mins read

The case for streaming is straightforward: data that arrives in minutes instead of hours enables decisions that were previously impossible. Fraud detection catches transactions before they clear. Pers

How to audit your AI pipeline for bias: step by step
How to audit your AI pipeline for bias: step by step
07 Jun, 2026 | 06 Mins read

Bias in AI systems is not a theoretical risk. It is a measurable property that can be detected, quantified, and mitigated at every stage of the pipeline. The teams that treat bias as an audit problem

The 30-day AI readiness assessment
The 30-day AI readiness assessment
14 Jun, 2026 | 07 Mins read

Organisations that skip readiness assessment before investing in AI tend to discover their gaps expensively. A financial services firm spent four months building a customer churn prediction model only

Data Pipelines for Time Series Forecasting
Data Pipelines for Time Series Forecasting
21 Mar, 2024 | 02 Mins read

Time series forecasting requires specialised pipeline architecture. Unlike standard batch processing, time series work demands strict chronological ordering, historical context, time-based feature eng

The death of the dashboard: what replaces BI?
The death of the dashboard: what replaces BI?
20 Jun, 2026 | 03 Mins read

The traditional BI dashboard, a grid of charts that a business user opens every morning to check KPIs, is losing its grip on how organisations consume data. The decline is not dramatic. No one declare

Your first 90 days as a Head of AI Engineering
Your first 90 days as a Head of AI Engineering
28 Jun, 2026 | 07 Mins read

The first Head of AI Engineering at a company inherits one of three situations. Situation one: there is no AI team, no AI infrastructure, and the mandate is to build from scratch. Situation two: there

The RAG evaluation framework you'll actually use
The RAG evaluation framework you'll actually use
08 Jul, 2026 | 06 Mins read

Most RAG systems are evaluated with vibes. An engineer runs ten queries, eyeballs the results, and declares the system "working." Three months later, a customer reports that the system confidently ret

Why your AI strategy needs a data strategy (not the other way around)
Why your AI strategy needs a data strategy (not the other way around)
11 Jul, 2026 | 03 Mins read

The majority of enterprise AI strategies are built on an implicit assumption: that the organisation's data is ready to support AI workloads. The assumption is almost always wrong. Data that is adequat

How to write an AI incident response plan
How to write an AI incident response plan
12 Jul, 2026 | 07 Mins read

AI systems fail differently than traditional software. A traditional software bug produces incorrect output deterministically. The same input always produces the same wrong output, and a fix eliminate

Data Contracts: Building Trust Between Teams
Data Contracts: Building Trust Between Teams
29 Jan, 2024 | 03 Mins read

Data contracts are formal agreements that define the structure, semantics, quality standards, and delivery expectations for data exchanged between teams. They specify schema definitions, SLAs, ownersh

Capacity planning for vector databases
Capacity planning for vector databases
19 Jul, 2026 | 07 Mins read

Vector database capacity planning fails in predictable ways. Teams estimate storage based on vector count alone and discover at 60% capacity that memory consumption is growing faster than disk because

The procurement checklist for AI vendors
The procurement checklist for AI vendors
26 Jul, 2026 | 07 Mins read

AI vendor procurement is where organisations make binding commitments that are expensive to unwind. A three-year contract with a model provider locks you into their pricing, their rate limits, their m

Setting up a model registry: the minimal viable approach
Setting up a model registry: the minimal viable approach
02 Aug, 2026 | 06 Mins read

A model registry is the version control system for your trained models. Without one, teams track model versions by filename, store artifacts in ad-hoc cloud storage locations, and discover which model

Data contract template and negotiation guide
Data contract template and negotiation guide
09 Aug, 2026 | 07 Mins read

Data pipelines break because data producers and data consumers have different assumptions. The producer assumes the consumer can handle null values in a column. The consumer assumes the column is neve

How to run an AI architecture review
How to run an AI architecture review
12 Aug, 2026 | 07 Mins read

An architecture review for an AI system catches design flaws at the cheapest possible stage: before implementation. A data pipeline that cannot handle the expected volume, a model serving architecture

The observability maturity model for AI systems
The observability maturity model for AI systems
16 Aug, 2026 | 07 Mins read

Most AI systems in production operate with observability that was designed for traditional software. Teams monitor CPU, memory, network, and error rates. These metrics tell you whether the server is r

Building an internal AI platform team: org chart and responsibilities
Building an internal AI platform team: org chart and responsibilities
23 Aug, 2026 | 07 Mins read

The decision to centralise AI infrastructure into a platform team usually comes after a period of decentralised pain. Three product teams independently built model serving pipelines. None of them shar

LLM cost calculator: estimating spend before you deploy
LLM cost calculator: estimating spend before you deploy
30 Aug, 2026 | 05 Mins read

Teams approve LLM projects based on per-query cost estimates, then get blindsided by the actual invoice. The gap between estimate and reality is not a rounding error. It is a structural problem: the e

Designing a data mesh operating model: roles, responsibilities, and boundaries
Designing a data mesh operating model: roles, responsibilities, and boundaries
06 Sep, 2026 | 04 Mins read

Most data mesh initiatives fail not because the architecture is wrong, but because nobody can answer the question: who owns this data product? When ownership is ambiguous, quality drops, SLAs go unmet

The rise of vertical AI: industry-specific models outperform generalists
The rise of vertical AI: industry-specific models outperform generalists
12 Sep, 2026 | 04 Mins read

The benchmark results from the past quarter are hard to ignore. On tasks spanning legal document analysis, medical coding, financial risk assessment, and manufacturing quality inspection, vertical AI

The LLM cost optimisation playbook: 12 techniques that actually save money
The LLM cost optimisation playbook: 12 techniques that actually save money
13 Sep, 2026 | 04 Mins read

LLM costs are easy to start and hard to control. A team ships a feature that calls GPT-4, the feature works, users like it, and the invoice climbs 15 percent month over month. The cost is not a proble

How to run a pre-mortem on your AI project
How to run a pre-mortem on your AI project
23 Sep, 2026 | 04 Mins read

Post-mortems are useful. Pre-mortems are cheaper. A post-mortem tells you why a project failed after the money is gone. A pre-mortem tells you why a project might fail while you can still change cours

Building Synthetic Data Pipelines for ML Testing
Building Synthetic Data Pipelines for ML Testing
24 May, 2024 | 04 Mins read

# Building Synthetic Data Pipelines for ML Testing Synthetic data addresses real ML development problems: privacy restrictions on real data, class imbalance, and edge case coverage. It does not repla

The AI project scoping template: right-size before you build
The AI project scoping template: right-size before you build
04 Oct, 2026 | 04 Mins read

AI projects have a scoping problem. Teams either scope too loosely, "use AI to improve customer experience", or too tightly: "build a transformer model with 12 attention layers for intent classificati

Zero-trust architecture for data pipelines: a practical guide
Zero-trust architecture for data pipelines: a practical guide
11 Oct, 2026 | 04 Mins read

Data pipelines are soft targets. They move sensitive data across network boundaries, authenticate with service accounts that have broad permissions, and log enough information to reconstruct entire da

Feature Store Architectures: Building the Foundation for Enterprise ML
Feature Store Architectures: Building the Foundation for Enterprise ML
18 Jan, 2024 | 03 Mins read

Organisations scaling ML efforts encounter a predictable problem: feature engineering work duplicates across teams, training-serving skew causes model failures in production, and point-in-time correct

Time-Travel Queries: Implementing Temporal Data Access
Time-Travel Queries: Implementing Temporal Data Access
02 Oct, 2024 | 03 Mins read

Time-travel queries (the ability to access data as it existed at any point in the past) have become essential in modern data platforms. This capability transforms how organisations approach data gover

Choosing a Vector Database for Production AI Applications
Choosing a Vector Database for Production AI Applications
10 Jul, 2026 | 12 Mins read

You have a retrieval-augmented generation proof of concept that works on a laptop. The embeddings are in a CSV file, the search is brute force, and the demo impresses the steering committee. Now someo