Skip to content

Running Benchmarks in CI Without Flaky Results

A benchmark in CI is supposed to fail when code gets slower. On shared runners it fails when the machine gets busier, and after a few false alarms nobody trusts it. Measured on a shared 24-core Linux machine with other workloads running (load average 5–8), using a benchmark that creates and awaits 20,000 asyncio tasks: a single run's coefficient of variation was 22–28%. Comparing two runs of identical code against a 5% threshold flagged a regression in 5–17 of 30 comparisons; taking the best of five runs on each side still flagged 3–7 of 30. Collecting garbage and disabling the collector around the timed region cut the variation to 7–10%. Comparing 15 interleaved pairs and requiring a bootstrap confidence interval to exclude zero cut false alarms to 0–1 of 30 — but on this machine it caught a real 17% slowdown only 12 of 30 times, because paired ratios ranged from 0.74 to 1.87. Counting executed bytecode instructions with sys.monitoring gave the same count, 460,438, on every warm run of identical code, and 628,438 for the slower variant. This guide builds a benchmark gate that fails for code, not for noise.

Prerequisites

1. Measure the noise before choosing a threshold

A regression threshold below the benchmark's run-to-run variation flags noise. Measure that variation on the machines CI actually uses, with the benchmark exactly as CI runs it:

import statistics, time

def noise(bench, runs=40):
    xs = [bench() for _ in range(runs)]
    mean = statistics.mean(xs)
    return statistics.stdev(xs) / mean, min(xs), max(xs)

cv, lo, hi = noise(lambda: timed(lambda: asyncio.run(create_and_await(20_000))))
print(f"CV {cv:.1%}, range {lo * 1e3:.1f}..{hi * 1e3:.1f} ms")

Measured: CV 22–28% with the garbage collector left on, and a range of roughly 2× between the fastest and slowest of 40 runs. With that much variation, a 5% threshold is not a test. Comparing two runs of identical code against it flagged a regression 5–17 times in 30 across four sessions — a false alarm in up to half of builds. Pinning the benchmark to one CPU core made it worse on this machine, because other processes were also scheduled on that core; pinning helps only on cores reserved for the benchmark.

Verify: you know the CV of your benchmark on CI hardware, and your threshold is several times larger — or you use a method from the steps below.

Variation of one benchmark run, 60 runs each A grid of 4 rows by 4 columns. Variation of one benchmark run, 60 runs each clock GC CV (unpinned) CV (pinned to a shared core) wall, perf_counter on 22.2% 27.6% wall, perf_counter collect + disable 9.6% 21.2% CPU, process_time collect + disable 8.3% 14.8% CPU, thread_time collect + disable 7.4% 18.6% Shared 24-core machine, load average 5-8 from other work.

2. Remove the noise your own process creates

Some variation comes from the benchmark itself. Asyncio benchmarks allocate many objects — tasks, futures, coroutines — and the cyclic garbage collector runs whenever allocation counters cross thresholds, which depends on everything allocated before the timed region:

import gc

def timed(fn, clock=time.perf_counter):
    gc.collect()                      # start every run from the same heap state
    gc.disable()
    try:
        start = clock()
        fn()
        return clock() - start
    finally:
        gc.enable()

Measured: the mean run time fell from 50–72 ms to 35–47 ms and the CV from 22–28% to 7–10% (unpinned), because collections no longer landed in some runs and not others. Disabling the collector means the benchmark no longer measures GC cost, which is the right trade for a regression gate — measure GC separately if it matters. Run a few untimed warm-up iterations first, and make each timed run long enough — tens of milliseconds at least — that timer resolution and scheduler ticks are small relative to it.

Verify: after these changes, the CV measured in step 1 has fallen, and its new value sets your threshold.

3. Compare interleaved pairs, and decide with a confidence interval

Machine load drifts over seconds and minutes, so running all baseline iterations and then all candidate iterations compares two different machines. Alternate them in random order, compute the ratio within each pair, and decide on the distribution of ratios:

import random, statistics

def regression(bench_base, bench_new, pairs=15, min_effect=1.03):
    ratios = []
    for _ in range(pairs):
        if random.random() < 0.5:
            a = bench_base(); b = bench_new()
        else:
            b = bench_new(); a = bench_base()
        ratios.append(b / a)
    medians = sorted(statistics.median(random.choices(ratios, k=len(ratios))) for _ in range(2000))
    low = medians[int(0.025 * len(medians))]               # 95% bootstrap lower bound
    return low > 1.0 and statistics.median(ratios) > min_effect

Measured over 30 comparisons of identical code: 0–1 false alarms with GC controlled, against 5–17 for a single run and 3–7 for best-of-five. The cost was power: a real 17% slowdown was flagged in only 12 of 30 comparisons, and a 4% one in 2 of 30, because on this machine individual paired ratios still ranged from 0.74 to 1.87. More pairs narrow the interval — its width shrinks with the square root of the count — so the number of pairs is a budget decision: quadrupling them halves the effect you can detect. On a quiet, dedicated runner the same method needs far fewer.

Verify: run the gate on identical code 20 times; it should almost never fail. Then run it on a deliberately slowed copy and check that it fails.

How often each rule flagged identical code and a real slowdown A grid of 3 rows by 4 columns. How often each rule flagged identical code and a real slowdown rule identical code flagged 4.4% slowdown caught 17% slowdown caught 1 run each, > 5% 5-17 / 30 13 / 30 25 / 30 best of 5 each, > 5% 3-7 / 30 11 / 30 28 / 30 15 interleaved pairs, CI > 0 and > 3% 0-1 / 30 2 / 30 12 / 30 The rules that catch more also cry wolf more; noise sets the trade-off.

4. Gate on something deterministic

Timing on shared hardware will always be noisy. A count of work done does not depend on the machine. For pure-Python code paths, sys.monitoring can count executed bytecode instructions:

import sys

def count_instructions(fn) -> int:
    mon, tool = sys.monitoring, sys.monitoring.PROFILER_ID
    n = 0

    def on_instruction(code, offset):
        nonlocal n
        n += 1

    mon.use_tool_id(tool, "icount")
    mon.register_callback(tool, mon.events.INSTRUCTION, on_instruction)
    mon.set_events(tool, mon.events.INSTRUCTION)
    try:
        fn()
    finally:
        mon.set_events(tool, 0)
        mon.register_callback(tool, mon.events.INSTRUCTION, None)
        mon.free_tool_id(tool)
    return n

Measured on a 2,000-task version of the benchmark: 460,581 instructions on the first run and exactly 460,438 on every run after it — the first run includes one-time specialisation — and 628,438 for the variant with extra work per task. A gate of "no more than 2% more instructions than the stored baseline" therefore never fails by chance. The count covers Python bytecode only: work done inside C — the event loop's internals, json, a database driver — is invisible to it, and an instruction is not a fixed amount of time. Use it to catch accidental extra Python work in hot paths, alongside other deterministic counts such as allocations from tracemalloc, queries per request, or round trips per operation.

Verify: the deterministic metric is identical across repeated CI runs of the same commit.

5. Split fast gates from slow trend tracking

Put the two kinds of measurement where each works. In every pull request, gate on deterministic counts and only on large timing regressions — a threshold several times the measured CV. Track precise timings on a schedule, on a dedicated or at least consistent machine, and look at trends rather than single comparisons:

# pull requests: fast, deterministic
- run: python -m benchmarks.count_gate --max-increase 0.02

# nightly, dedicated runner: timings with pyperf, stored for trend analysis
- run: python -m pyperf system tune || true
- run: python -m benchmarks.suite -o results/$(date +%F).json
- run: python -m pyperf compare_to results/baseline.json results/$(date +%F).json --table

pyperf compare_to reports whether differences are significant, and a nightly series makes a slow drift visible that no single comparison would flag. Keep the benchmark's environment fixed — Python version, dependency lock, CPU model — and record it with every result, since an interpreter upgrade alone can move timings by more than the regressions you are hunting, as measured in measuring asyncio speedups between Python versions.

Verify: pull-request benchmarks fail only on deterministic regressions or large timing jumps, and a nightly series exists for everything else.

Which benchmark check belongs where? A decision on What change must be caught with 4 outcomes. Which benchmark check belongs where? What change must be caught? extra Python work in a hot path instruction / allocation counts per PR identical every run a large slowdown timing gate at several x CV rare false alarms a few percent slower nightly pyperf on a fixed runner trend, not one run 5% on a shared runner do not gate on it 5-17 / 30 false alarms Gate on what is deterministic; trend what is not.

Verification

CI benchmarks are trustworthy when:

  • Run-to-run noise is measured on the CI hardware, and thresholds sit well above it.
  • GC and warm-up are controlled inside the timed region.
  • Comparisons use interleaved pairs and a confidence interval, not two single numbers.
  • Pull requests gate on deterministic counts, with precise timings tracked nightly.

Diagnostic Hook: record the CV of every benchmark run alongside its result. When a timing gate fails, check whether that run's CV was unusually high; a failure during a noisy run is a symptom of the runner, and a gate that refuses to decide when noise is high is better than one that decides wrongly.

Pitfalls & edge cases

  • A threshold below the noise. Measured: identical code flagged in up to 17 of 30 runs.
  • Best-of-N as a cure. It still flagged identical code in 3–7 of 30.
  • Pinning to a shared core. It increased the variation here.
  • Treating instruction counts as time. They miss C-level work entirely.

Frequently Asked Questions

Why are my CI benchmarks flaky?

Run-to-run variation on shared runners is often larger than the threshold: one run varied by 22 to 28% here, so a 5% threshold flagged identical code in 5 to 17 of 30 comparisons.

How do I reduce benchmark noise in Python?

Collect and disable the garbage collector around the timed region, warm up, and make runs long enough; this cut variation from 22 to 28% to 7 to 10% for an asyncio benchmark. Then compare interleaved pairs rather than separate batches.

How can a benchmark gate be deterministic?

Count work instead of timing it: bytecode instructions via sys.monitoring were identical on every run of the same code (460,438) and rose to 628,438 with extra work. Counts miss C-level work, so pair them with nightly timings.

How many benchmark runs do I need to detect a 5% regression?

It depends on the noise: with paired ratios ranging from 0.74 to 1.87, 15 pairs caught a 17% slowdown only 12 times in 30. Reduce noise first; quadrupling pairs halves the detectable effect.