Running Benchmarks in CI Without Flaky Results¶
A benchmark in CI is supposed to fail when code gets slower. On shared runners it fails when the machine gets busier, and after a few false alarms nobody trusts it. Measured on a shared 24-core Linux machine with other workloads running (load average 5–8), using a benchmark that creates and awaits 20,000 asyncio tasks: a single run's coefficient of variation was 22–28%. Comparing two runs of identical code against a 5% threshold flagged a regression in 5–17 of 30 comparisons; taking the best of five runs on each side still flagged 3–7 of 30. Collecting garbage and disabling the collector around the timed region cut the variation to 7–10%. Comparing 15 interleaved pairs and requiring a bootstrap confidence interval to exclude zero cut false alarms to 0–1 of 30 — but on this machine it caught a real 17% slowdown only 12 of 30 times, because paired ratios ranged from 0.74 to 1.87. Counting executed bytecode instructions with sys.monitoring gave the same count, 460,438, on every warm run of identical code, and 628,438 for the slower variant. This guide builds a benchmark gate that fails for code, not for noise.
Prerequisites¶
- Python 3.12+ for
sys.monitoring; pytest or a plain script in CI. - Microbenchmarking basics, from microbenchmarking coroutines with pyperf and timeit.
- The topic overview, Load Testing & Benchmarking.
1. Measure the noise before choosing a threshold¶
A regression threshold below the benchmark's run-to-run variation flags noise. Measure that variation on the machines CI actually uses, with the benchmark exactly as CI runs it:
import statistics, time
def noise(bench, runs=40):
xs = [bench() for _ in range(runs)]
mean = statistics.mean(xs)
return statistics.stdev(xs) / mean, min(xs), max(xs)
cv, lo, hi = noise(lambda: timed(lambda: asyncio.run(create_and_await(20_000))))
print(f"CV {cv:.1%}, range {lo * 1e3:.1f}..{hi * 1e3:.1f} ms")
Measured: CV 22–28% with the garbage collector left on, and a range of roughly 2× between the fastest and slowest of 40 runs. With that much variation, a 5% threshold is not a test. Comparing two runs of identical code against it flagged a regression 5–17 times in 30 across four sessions — a false alarm in up to half of builds. Pinning the benchmark to one CPU core made it worse on this machine, because other processes were also scheduled on that core; pinning helps only on cores reserved for the benchmark.
Verify: you know the CV of your benchmark on CI hardware, and your threshold is several times larger — or you use a method from the steps below.
2. Remove the noise your own process creates¶
Some variation comes from the benchmark itself. Asyncio benchmarks allocate many objects — tasks, futures, coroutines — and the cyclic garbage collector runs whenever allocation counters cross thresholds, which depends on everything allocated before the timed region:
import gc
def timed(fn, clock=time.perf_counter):
gc.collect() # start every run from the same heap state
gc.disable()
try:
start = clock()
fn()
return clock() - start
finally:
gc.enable()
Measured: the mean run time fell from 50–72 ms to 35–47 ms and the CV from 22–28% to 7–10% (unpinned), because collections no longer landed in some runs and not others. Disabling the collector means the benchmark no longer measures GC cost, which is the right trade for a regression gate — measure GC separately if it matters. Run a few untimed warm-up iterations first, and make each timed run long enough — tens of milliseconds at least — that timer resolution and scheduler ticks are small relative to it.
Verify: after these changes, the CV measured in step 1 has fallen, and its new value sets your threshold.
3. Compare interleaved pairs, and decide with a confidence interval¶
Machine load drifts over seconds and minutes, so running all baseline iterations and then all candidate iterations compares two different machines. Alternate them in random order, compute the ratio within each pair, and decide on the distribution of ratios:
import random, statistics
def regression(bench_base, bench_new, pairs=15, min_effect=1.03):
ratios = []
for _ in range(pairs):
if random.random() < 0.5:
a = bench_base(); b = bench_new()
else:
b = bench_new(); a = bench_base()
ratios.append(b / a)
medians = sorted(statistics.median(random.choices(ratios, k=len(ratios))) for _ in range(2000))
low = medians[int(0.025 * len(medians))] # 95% bootstrap lower bound
return low > 1.0 and statistics.median(ratios) > min_effect
Measured over 30 comparisons of identical code: 0–1 false alarms with GC controlled, against 5–17 for a single run and 3–7 for best-of-five. The cost was power: a real 17% slowdown was flagged in only 12 of 30 comparisons, and a 4% one in 2 of 30, because on this machine individual paired ratios still ranged from 0.74 to 1.87. More pairs narrow the interval — its width shrinks with the square root of the count — so the number of pairs is a budget decision: quadrupling them halves the effect you can detect. On a quiet, dedicated runner the same method needs far fewer.
Verify: run the gate on identical code 20 times; it should almost never fail. Then run it on a deliberately slowed copy and check that it fails.
4. Gate on something deterministic¶
Timing on shared hardware will always be noisy. A count of work done does not depend on the machine. For pure-Python code paths, sys.monitoring can count executed bytecode instructions:
import sys
def count_instructions(fn) -> int:
mon, tool = sys.monitoring, sys.monitoring.PROFILER_ID
n = 0
def on_instruction(code, offset):
nonlocal n
n += 1
mon.use_tool_id(tool, "icount")
mon.register_callback(tool, mon.events.INSTRUCTION, on_instruction)
mon.set_events(tool, mon.events.INSTRUCTION)
try:
fn()
finally:
mon.set_events(tool, 0)
mon.register_callback(tool, mon.events.INSTRUCTION, None)
mon.free_tool_id(tool)
return n
Measured on a 2,000-task version of the benchmark: 460,581 instructions on the first run and exactly 460,438 on every run after it — the first run includes one-time specialisation — and 628,438 for the variant with extra work per task. A gate of "no more than 2% more instructions than the stored baseline" therefore never fails by chance. The count covers Python bytecode only: work done inside C — the event loop's internals, json, a database driver — is invisible to it, and an instruction is not a fixed amount of time. Use it to catch accidental extra Python work in hot paths, alongside other deterministic counts such as allocations from tracemalloc, queries per request, or round trips per operation.
Verify: the deterministic metric is identical across repeated CI runs of the same commit.
5. Split fast gates from slow trend tracking¶
Put the two kinds of measurement where each works. In every pull request, gate on deterministic counts and only on large timing regressions — a threshold several times the measured CV. Track precise timings on a schedule, on a dedicated or at least consistent machine, and look at trends rather than single comparisons:
# pull requests: fast, deterministic
- run: python -m benchmarks.count_gate --max-increase 0.02
# nightly, dedicated runner: timings with pyperf, stored for trend analysis
- run: python -m pyperf system tune || true
- run: python -m benchmarks.suite -o results/$(date +%F).json
- run: python -m pyperf compare_to results/baseline.json results/$(date +%F).json --table
pyperf compare_to reports whether differences are significant, and a nightly series makes a slow drift visible that no single comparison would flag. Keep the benchmark's environment fixed — Python version, dependency lock, CPU model — and record it with every result, since an interpreter upgrade alone can move timings by more than the regressions you are hunting, as measured in measuring asyncio speedups between Python versions.
Verify: pull-request benchmarks fail only on deterministic regressions or large timing jumps, and a nightly series exists for everything else.
Verification¶
CI benchmarks are trustworthy when:
- Run-to-run noise is measured on the CI hardware, and thresholds sit well above it.
- GC and warm-up are controlled inside the timed region.
- Comparisons use interleaved pairs and a confidence interval, not two single numbers.
- Pull requests gate on deterministic counts, with precise timings tracked nightly.
Diagnostic Hook: record the CV of every benchmark run alongside its result. When a timing gate fails, check whether that run's CV was unusually high; a failure during a noisy run is a symptom of the runner, and a gate that refuses to decide when noise is high is better than one that decides wrongly.
Pitfalls & edge cases¶
- A threshold below the noise. Measured: identical code flagged in up to 17 of 30 runs.
- Best-of-N as a cure. It still flagged identical code in 3–7 of 30.
- Pinning to a shared core. It increased the variation here.
- Treating instruction counts as time. They miss C-level work entirely.
Frequently Asked Questions¶
Why are my CI benchmarks flaky?
Run-to-run variation on shared runners is often larger than the threshold: one run varied by 22 to 28% here, so a 5% threshold flagged identical code in 5 to 17 of 30 comparisons.
How do I reduce benchmark noise in Python?
Collect and disable the garbage collector around the timed region, warm up, and make runs long enough; this cut variation from 22 to 28% to 7 to 10% for an asyncio benchmark. Then compare interleaved pairs rather than separate batches.
How can a benchmark gate be deterministic?
Count work instead of timing it: bytecode instructions via sys.monitoring were identical on every run of the same code (460,438) and rose to 628,438 with extra work. Counts miss C-level work, so pair them with nightly timings.
How many benchmark runs do I need to detect a 5% regression?
It depends on the noise: with paired ratios ranging from 0.74 to 1.87, 15 pairs caught a 17% slowdown only 12 times in 30. Reduce noise first; quadrupling pairs halves the detectable effect.
Related¶
- Load Testing & Benchmarking — up to the topic overview.
- Capacity planning with Little's law — from benchmark numbers to capacity.
- Resilience, Cancellation & Error Handling — the section overview.