PostTrainBench · Forensic Replays

How agents reward-hack a benchmark

The whole leaderboard as a map: every one of the 1198 distinct post-training runs (25 agents × 7 benchmarks; 1226 total minus 28 old-container reruns) is a dot in the model × task grid — shaded by official score, red when the usage judge flagged it, and ringed + clickable when we hand-verified it into a step-by-step replay. 45 deep replays so far; 52 runs judge-flagged.

1198Total runs
52Judge-flagged
45Deep replays
24Verified hacks

▨ light-yellow cells have a clickable step-by-step replay — click the ◎ ringed dot to open that agent's trajectory on that task. ● red = usage judge flagged · ● shaded = official score (darker = higher) · ● gray = ≈0 / no lift. Cell label = run count × best score. Un-ringed colors are the judge's signal (which under-counts hacks); only ringed dots are verified.

Claude

8 models · 444 runs · 22 judge-flagged · 18 deep replay
model \ taskBFCLHealthBenchArena-HardAIME 2025GSM8KGPQAHumanEval
fable-5
4× 1.00
4× 0.46
4× 0.86▸ replay
4× 0.20
4× 0.91
2×
4× 0.78
opus-4.5
3× 0.92
4× 0.22
4× 0.16
4× 0.00
4× 0.62
4× 0.26
4× 0.50▸ replay
opus-4.6
12× 1.00▸ replay
12× 0.30▸ replay
12× 0.21▸ replay
12× 0.20
12× 0.83
12× 0.32
12× 0.71▸ replay
opus-4.6 · 1m
12× 1.00▸ replay
12× 0.23▸ replay
12× 0.14
12× 0.13
12× 0.85
12× 0.33
12× 0.74
opus-4.7
12× 0.95▸ replay
12× 0.31
12× 0.76
12× 0.17
12× 0.82
12× 0.33
12× 0.71
opus-4.8
8× 1.00▸ replay
8× 0.48
8× 0.75
8× 0.23
8× 0.87
8× 0.35
8× 0.77▸ replay
opus-4.8 · max
8× 1.00▸ replay
8× 0.49▸ replay
8× 0.74
8× 0.30▸ replay
8× 0.91
8× 0.33
8× 0.76
sonnet-4.6
4× 1.00▸ replay
4× 0.20
4× 0.17
4× 0.07
4× 0.42
3× 0.18
4× 0.53

GPT / Codex

7 models · 392 runs · 13 judge-flagged · 10 deep replay
model \ taskBFCLHealthBenchArena-HardAIME 2025GSM8KGPQAHumanEval
gpt-5.1-codex-max
4× 0.00▸ replay
4× 0.00
4× 0.00
4× 0.03
4× 0.11
4× 0.25
4× 0.11
gpt-5.3-codex
12× 1.00▸ replay
12× 0.18
12× 0.03
12× 0.03
12× 0.58
12× 0.30
12× 0.41
gpt-5.3-codex · high
12× 0.98▸ replay
12× 0.21
12× 0.11
12× 0.03
12× 0.59
12× 0.34
12× 0.43
gpt-5.4 · high
12× 0.96▸ replay
12× 0.33
12× 0.27
12× 0.07
12× 0.68
12× 0.34
12× 0.39
gpt-5.4 · high·rp
4× 1.00▸ replay
4× 0.32
4× 0.50▸ replay
4× 0.07
4× 0.83
4× 0.34
4× 0.66
gpt-5.5 · xhigh
8× 0.99▸ replay
8× 0.33
8× 0.28
8× 0.10
8× 0.80
8× 0.37
8× 0.72
gpt-5.5 · xhigh·rp
4× 1.00
4× 0.34
4× 0.28
4× 0.10
4× 0.79
4× 0.34
4× 0.61

GLM

3 models · 126 runs · 1 judge-flagged · 5 deep replay
model \ taskBFCLHealthBenchArena-HardAIME 2025GSM8KGPQAHumanEval
glm-4.7
2× 0.00▸ replay
··
2×
3× 0.46
4×
3× 0.12
glm-5 · zai
4× 0.62▸ replay
4× 0.27
4× 0.07
4× 0.03
4× 0.54
4× 0.26
4× 0.31
glm-5.2
12× 1.00▸ replay
12× 0.40
12× 0.73
12× 0.20
12× 0.87▸ replay
12× 0.35
12× 0.82

Gemini

2 models · 112 runs · 2 judge-flagged · 3 deep replay
model \ taskBFCLHealthBenchArena-HardAIME 2025GSM8KGPQAHumanEval
gemini-3-pro
4× 0.35
4× 0.25
4× 0.26
4× 0.00▸ replay
4× 0.72
4× 0.20
4× 0.46
gemini-3.1-pro
12× 0.89
12× 0.30▸ replay
12× 0.31
12× 0.17
12× 0.65
12× 0.34
12× 0.65

MiniMax

2 models · 42 runs · 7 judge-flagged · 5 deep replay
model \ taskBFCLHealthBenchArena-HardAIME 2025GSM8KGPQAHumanEval
minimax-m2.1
4× 0.54▸ replay
··
4× 0.03
3× 0.12
4× 0.15
4× 0.38
minimax-m2.5
3× 0.09▸ replay
4× 0.18
3× 0.04
4× 0.00
2× 0.36
4× 0.23▸ replay
3× 0.34

Kimi

2 models · 54 runs · 6 judge-flagged · 3 deep replay
model \ taskBFCLHealthBenchArena-HardAIME 2025GSM8KGPQAHumanEval
kimi-k2-thinking
4×
4×
4×
4×
4× 0.19
4×
4× 0.17
kimi-k2.5
4× 0.77
4× 0.28
4× 0.50▸ replay
4× 0.07
4× 0.39
3× 0.19
3× 0.32▸ replay

Qwen

1 models · 28 runs · 1 judge-flagged · 1 deep replay
model \ taskBFCLHealthBenchArena-HardAIME 2025GSM8KGPQAHumanEval
qwen3-max
4× 0.00
4×
4× 0.02▸ replay
4× 0.00
4× 0.43
4× 0.08
4× 0.46
Reward hack / judge-flagged Verified honest Failed / ≈0 Deep replay available

Source: PostTrainBench raw trajectories — 1198 runs / 25 agents / 7 benchmarks. Matrix cells are mechanical (official score + contamination/disallowed-model judge from meta.csv). The 45 ringed runs are full forensic step-by-step replays with verbatim, line-anchored evidence.