Debugging a Flaky API Test That Only Fails in CI
There's a particular kind of frustration reserved for the test that passes every single time on your laptop and then fails — sometimes — in CI. Not always. Not on a schedule. Just often enough to block a merge and make you doubt everything. I've seen teams spend days chasing this pattern, and the fix is almost never "retry the test." The fix is understanding why the two environments are different in the first place.
The core problem is that CI is not your laptop. The network latency is different, the clock might be slightly off, the environment variables are injected differently, the filesystem is ephemeral, and the test runner might be executing things in a different order or in parallel. Any one of those differences can turn a perfectly reasonable test into a non-deterministic mess. The good news is that each of those causes leaves a fingerprint, and once you know what to look for, you can usually narrow it down in one or two focused debugging sessions.
In this article I'll walk through the most common root causes I've encountered for CI-only flakiness in API tests, show you how to instrument your tests to surface the real signal, and give you concrete fixes you can apply today — not just "add a sleep and hope for the best."
Learn Python, Behave, GitHub Copilot, APIs, and CI/CD by building a real framework you can finish in a weekend.
The Five Root Causes Behind CI-Only API Test Failures
Before you can fix a flaky CI failure, you need a mental model of what's actually different between your local environment and the pipeline. In practice, almost every CI-only flaky API test traces back to one of these five categories:
1. Timing and latency assumptions
The most common culprit. Your test calls an endpoint that triggers an async operation — a job queue, a webhook, a cache invalidation — and then immediately asserts on the result. Locally, everything is fast enough that the assertion wins the race. In CI, the container is cold, the network hop is real, and the assertion fires before the operation completes. A hardcoded time.sleep(1) that "works locally" is not a fix; it's a liability. Replace it with a proper polling loop:
import time
import requests
def wait_for_status(url, expected_status, timeout=10, interval=0.5, headers=None):
deadline = time.time() + timeout
while time.time() < deadline:
response = requests.get(url, headers=headers)
if response.status_code == expected_status:
return response
time.sleep(interval)
raise TimeoutError(f"Status {expected_status} not reached at {url} within {timeout}s")
This gives CI the time it needs without burning arbitrary seconds on every run.
2. Environment variable differences
Your local .env file has a value that CI doesn't. The test silently falls back to a default — maybe an empty string, maybe None — and the API call either hits the wrong host or sends a malformed auth header. Add an explicit assertion at the top of your test module or fixture:
import os
import pytest
@pytest.fixture(scope="session", autouse=True)
def require_env_vars():
required = ["API_BASE_URL", "API_KEY"]
missing = [v for v in required if not os.getenv(v)]
if missing:
pytest.fail(f"Missing required environment variables: {missing}")
This turns a mysterious 401 or a connection error into an immediate, readable failure message.
3. Test ordering and shared state
CI pipelines often run tests in a different order than your local session — especially if you've recently added tests or changed a file. If Test B relies on data that Test A created, and the order changes, Test B fails. This is also the dominant cause of flakiness when you start running tests in parallel. Every API test should own its setup and teardown. Use fixtures to create the data it needs and clean it up afterward — never depend on a prior test having run.
4. Clock skew and token expiry
JWT tokens have expiry times. If CI's system clock is even slightly behind the token-issuing server's clock, a freshly minted token can appear expired before the first request fires. Check your CI runner's NTP configuration, and build a small buffer into any token-expiry logic in your test helpers.
5. Rate limiting and throttling
CI often runs multiple jobs concurrently against the same staging API. If your pipeline doesn't account for rate limits, tests will start receiving 429 responses that never appear during a single local run. Treat a 429 as a retryable condition in your HTTP client wrapper, and consider staggering test jobs in your pipeline configuration.
Instrumenting the Test to See What CI Actually Experienced
The biggest mistake teams make when chasing CI flakiness is trying to fix the test without first capturing what actually happened. CI logs are often thin by default — you see a failed assertion, but not the response body, the headers, the latency, or the sequence of calls. Before you change a single line of test logic, add the instrumentation that makes the failure self-describing.
Good logging at the HTTP layer is the fastest path to a diagnosis. I've written before about structuring your test logging to surface the right signal, and the same principles apply here. The key is capturing request and response details in a way that shows up in CI output without drowning you in noise on a passing run. A pytest fixture using requests hooks is a clean way to do this:
import logging
import pytest
import requests
logger = logging.getLogger(__name__)
@pytest.fixture
def logged_session():
session = requests.Session()
def log_response(response, *args, **kwargs):
elapsed = response.elapsed.total_seconds()
logger.info(
"HTTP %s %s -> %d (%.3fs)",
response.request.method,
response.request.url,
response.status_code,
elapsed,
)
if not response.ok:
logger.debug("Response body: %s", response.text[:500])
session.hooks["response"].append(log_response)
return session
Now every test that uses logged_session automatically emits timing and status information. When a test fails in CI, you'll see exactly which call was slow, which returned an unexpected status, and what the body looked like — without having to reproduce the failure locally.
Capturing timing distributions, not just pass/fail
A single slow response isn't always a bug. A response that is consistently slow only in CI points to a real infrastructure difference worth investigating. Add elapsed-time assertions as soft warnings rather than hard failures during your investigation phase:
def test_get_user_profile(logged_session):
response = logged_session.get(f"{BASE_URL}/users/42")
assert response.status_code == 200
elapsed = response.elapsed.total_seconds()
if elapsed > 2.0:
logger.warning("Slow response: %.3fs — investigate CI network latency", elapsed)
data = response.json()
assert data["id"] == 42
This keeps the test from failing on latency alone while still flagging the condition in your CI logs. Once you've confirmed the threshold, you can promote it to a hard assertion.
Reproducing CI conditions locally
If you still can't reproduce the failure, try to close the gap between environments. Set only the environment variables your CI pipeline sets (use a .env.ci file), run your tests with pytest -p no:randomly disabled to match CI's default ordering, and if your CI runs tests in parallel, run them locally with pytest-xdist at the same concurrency level. Narrowing the environmental delta is often enough to make the failure reproducible.
Fixing the Flake for Good — and Keeping It Fixed
Once you've identified the root cause, the fix needs to be structural, not cosmetic. Increasing a timeout or adding a retry decorator might make the failure less frequent, but it doesn't eliminate the underlying fragility. Here's how to close each category of flakiness permanently.
Replace timing assumptions with explicit wait strategies
Every place in your test suite where you're waiting for an async side effect, use the polling helper shown in section one — or, if the API supports it, use a webhook or a polling endpoint that signals completion. Design your tests to wait for a condition, not for a duration. This is especially important as your suite grows; a test suite built on well-defined layers keeps async concerns at the API layer where they belong, rather than leaking timing hacks into every test.
Enforce environment contract at startup
The require_env_vars fixture pattern from section one should be a permanent part of your conftest, not a temporary debugging aid. Make it scope="session" and autouse=True so it runs once per CI job and fails loudly if the pipeline is misconfigured. This prevents a whole class of mysterious 401s and connection errors from ever masquerading as test failures.
Isolate test data ownership
Audit every test that reads from shared state — a shared user account, a shared record ID, a shared database row. Each of those is a time bomb. Replace them with fixtures that create a unique resource for each test run and delete it on teardown. Use a naming convention like test-{uuid4()} for created resources so you can identify and clean up orphaned test data in staging environments.
import uuid
import pytest
import requests
@pytest.fixture
def test_user(api_session):
unique_name = f"test-{uuid.uuid4().hex[:8]}"
response = api_session.post("/users", json={"username": unique_name})
response.raise_for_status()
user_id = response.json()["id"]
yield user_id
# Teardown — always runs, even on test failure
api_session.delete(f"/users/{user_id}")
Add a flakiness gate to your CI pipeline
Once your suite is stable, protect it. Configure your CI to re-run only failed tests once (most pipelines support this natively) and treat a test that fails on the first run but passes on the retry as a flakiness signal worth investigating — not a green light to merge. Log these retry-pass events explicitly. A test that "passes eventually" is still a flaky test; it's just failing quietly.
The goal isn't a test suite that never fails. It's a test suite where every failure means something real. Flaky CI failures erode trust in the entire suite — developers start ignoring red builds, and the suite stops doing its job. Fixing flakiness isn't housekeeping; it's what keeps your test investment from depreciating.