trajectory
py-flaky-test-01core-12pythonDifficulty tier 3/5

Fix a test that fails one run in four without deleting it

flakinessdeterminismtestingrandomness

Task parameters

Reference steps
11
Step ceiling
40
Runs
15
Solved
12 of 15
Models
5

What is broken, and what fixed means

A sampling library for a request auditor draws from the global `random` stream, so the behaviour a test observes depends on whatever entropy the process started with. One test asserts that a sample of 8 requests out of 40 covers all four tenants, which is true about seven runs in ten, so the suite fails roughly one run in four and passes on rerun. Fixing it means making the randomness injectable and pinning the draw in the test, not seeding a global inside the library, not loosening the assertion, and not marking the test away. Production sampling has to stay random, and the hidden tests check that it still is.

Results by model

One group per model. The solve rate carries its spread across seeds, and every run below it links to the full step by step replay.

stub:hasty

0.0%+/- 0.0% over 3 seeds

3 runs, seeds 0, 1, 2

stub:methodical

100.0%+/- 0.0% over 3 seeds

3 runs, seeds 0, 1, 2

stub:reckless

100.0%+/- 0.0% over 3 seeds

3 runs, seeds 0, 1, 2

stub:sloppy

100.0%+/- 0.0% over 3 seeds

3 runs, seeds 0, 1, 2

stub:thrasher

100.0%+/- 0.0% over 3 seeds

3 runs, seeds 0, 1, 2