REAL-WORLD SCENARIOS

Handling Third-Party API Downtime in Your Test Suite

Third-party APIs go down. Payment gateways have maintenance windows, weather services hit rate limits, and authentication providers throw 503s at the worst possible moment — usually right in the middle of a CI run. I've seen teams lose hours chasing what looks like a broken test, only to discover the real culprit is an upstream service they don't control. The test wasn't wrong. The world was.

The problem is that most test suites treat third-party downtime as an edge case rather than a design constraint. If your integration tests make live calls to external services, you're not just testing your code — you're testing whether those services happen to be healthy right now. That's a brittle dependency, and it turns real failures into noise and noise into missed real failures. The fix isn't to stop testing integrations; it's to be deliberate about when and how you lean on live external calls.

In this article I'll walk through the practical patterns I rely on: how to isolate third-party calls so downtime doesn't cascade, how to write tests that distinguish between "our code is broken" and "their service is down," and how to make your CI pipeline resilient without sacrificing confidence in your coverage. Every technique here is something you can wire into an existing Python test suite today.

Build an API Automation Framework With Node.js

Learn Node.js, Cucumber, GitHub Copilot, APIs, CI/CD, and modern automation by building a complete framework.

Learn more

Isolating Third-Party Calls So One Outage Can't Collapse Your Whole Suite

The first thing to get right is structural isolation. If a third-party call lives buried inside a function that also does business logic, a 503 from that service will fail the test and tell you nothing useful. You need a seam — a clear boundary between your code and the external world — so you can substitute a fake when the real thing isn't available.

In Python this usually means a thin client wrapper around the external service:

# services/payment_client.py
import requests

class PaymentClient:
    def __init__(self, base_url, api_key, session=None):
        self.base_url = base_url
        self.api_key = api_key
        self._session = session or requests.Session()

    def charge(self, amount_cents, token):
        response = self._session.post(
            f"{self.base_url}/charges",
            json={"amount": amount_cents, "token": token},
            headers={"Authorization": f"Bearer {self.api_key}"},
            timeout=5,
        )
        response.raise_for_status()
        return response.json()

Now your business logic depends on PaymentClient, not on requests directly. In tests you inject a fake or a mock session. The real HTTP call is confined to one place, and when the payment gateway goes down, only the tests that explicitly exercise that client are affected — not every test that touches checkout logic.

This maps directly to the principle behind layering your test suite: unit tests mock the client, integration tests hit a sandbox or a recorded response, and only a small slice of end-to-end tests ever make live calls. When you've got that layering in place, a third-party outage degrades one layer — it doesn't bring down the whole pyramid.

Tag your tests to make the boundary explicit. With pytest, a simple marker does the job:

# conftest.py
import pytest

def pytest_configure(config):
    config.addinivalue_line(
        "markers", "live: marks tests that call real external services"
    )
# test_payment_live.py
import pytest

@pytest.mark.live
def test_real_charge_returns_transaction_id(payment_client):
    result = payment_client.charge(100, "tok_visa")
    assert "transaction_id" in result

Now CI can run pytest -m "not live" on every push and reserve pytest -m live for a scheduled nightly job or a manual gate. Downtime in the payment gateway breaks the nightly run — which is expected and visible — but it never blocks a developer's pull request.

Writing Tests That Tell You "Their Service Is Down," Not Just "Something Failed"

Isolation buys you structural protection, but you still need the tests that do call live services to fail informatively. A bare requests.exceptions.ConnectionError in a CI log is almost useless. You want your test output to say: "The external service returned 503 — this is likely a provider outage, not a code defect."

The pattern I use is a small availability check that runs before the live test suite and short-circuits with a clear skip message if the service is unreachable:

# conftest.py
import pytest
import requests

def is_service_available(url):
    try:
        r = requests.get(url, timeout=3)
        return r.status_code < 500
    except requests.exceptions.RequestException:
        return False

@pytest.fixture(scope="session", autouse=False)
def require_payment_service():
    if not is_service_available("https://api.payments.example.com/health"):
        pytest.skip("Payment service unavailable — skipping live tests")
# test_payment_live.py
@pytest.mark.live
def test_real_charge_returns_transaction_id(payment_client, require_payment_service):
    result = payment_client.charge(100, "tok_visa")
    assert "transaction_id" in result

Now a downed service produces SKIPPED in your report instead of ERROR or a misleading FAILED. That distinction matters enormously when you're triaging CI results at speed. A skip is a signal that says "external dependency unavailable"; a failure says "your code is broken." Conflating them is how teams start ignoring red builds.

For tests that must run even during downtime — say, you're verifying your own retry and fallback logic — use responses or pytest-httpserver to simulate the outage conditions your code is supposed to handle:

import responses as rsps

@rsps.activate
def test_charge_retries_on_503():
    rsps.add(rsps.POST, "https://api.payments.example.com/charges",
             status=503, json={"error": "service_unavailable"})
    rsps.add(rsps.POST, "https://api.payments.example.com/charges",
             status=200, json={"transaction_id": "txn_abc123"})

    client = PaymentClient("https://api.payments.example.com", "key_test")
    result = client.charge_with_retry(100, "tok_visa", max_retries=2)
    assert result["transaction_id"] == "txn_abc123"

This test runs in full isolation and proves your retry logic works — without needing the real service to actually be down. That's the kind of coverage that protects you when a third-party outage hits production. For more on how this intersects with CI-specific flakiness, the patterns in debugging a flaky API test that only fails in CI are worth keeping in mind — a lot of the same environmental instability applies here.

Keeping Your CI Pipeline Green Through Third-Party Outages Without Losing Real Coverage

The end goal is a CI pipeline that stays green during an outage and still catches real regressions. Those two things feel like they're in tension, but they're not — you just need to be deliberate about what "green" means at each stage of the pipeline.

Here's the pipeline split I recommend for suites that depend on third-party services:

  • On every push / PR: Run unit tests and mocked integration tests only (-m "not live"). These must always pass. They test your code, not the world.
  • On merge to main: Run the full integration suite including sandbox/staging calls, but wrap live-service tests in availability checks so they skip gracefully instead of failing hard.
  • On a nightly schedule: Run live end-to-end tests against real external services. Alert on failures, but don't block deployments on this result alone.

The nightly job is where you catch genuine integration drift — things like a payment provider quietly changing a response field, which is exactly the kind of subtle breakage that API contract testing is designed to surface. A contract test can tell you "the field transaction_id disappeared from the response schema" even when you're running against a recorded fixture, giving you early warning before the nightly live run confirms it.

One more thing worth wiring in: record and replay. The vcrpy library lets you record real HTTP interactions on a healthy day and replay them in CI when the service is down:

import vcr

@vcr.use_cassette("fixtures/cassettes/charge_success.yaml")
def test_charge_success_recorded():
    client = PaymentClient("https://api.payments.example.com", "key_test")
    result = client.charge(100, "tok_visa")
    assert result["transaction_id"].startswith("txn_")

The cassette file lives in version control. When the real service is healthy, you regenerate it periodically to keep the fixture fresh. When the service is down, the test runs against the recording and your pipeline stays green. The tradeoff is that stale cassettes can mask drift — so pair this with a scheduled live run that regenerates cassettes and alerts if the real response no longer matches the recorded shape.

Managing the test data that feeds these cassettes and fixtures can get complicated across a large suite. The strategies for sharing test data across a Behave suite safely translate well here — the same concerns about fixture ownership, mutation, and scope apply whether you're using Behave or pytest.

The bottom line: third-party downtime is a fact of life, not an excuse for a broken pipeline. Isolate the seams, tag the live tests, skip gracefully, simulate the failure modes you care about, and record the happy paths. Do those four things and an upstream outage becomes a known, managed event instead of a fire drill.