Benchmark prefix maximum vs elapsed time
Fable and GPT-5.6-Sol research trajectories · raw scores are faint marks · curves are elapsed-time prefix maxima.
GRPO baseline 说明:这里的 current GRPO-style baseline 不是 naive GRPO,而是包含 decoupled clip 和其他组件的当前参考配置。它仅作为对照,不代表模型只有 GRPO 时代的水平。
显示的 baselines
X 轴 · elapsed hours
Y 轴范围 · 每个 benchmark
claude-fable-5gpt-5.6-sol xhigh
Average is the unweighted mean of AIME 2026, HMMT Feb 2026, and MATH500. Evaluation time is measured when the full evaluation finishes. Baseline reproductions are excluded from prefix max. Click any raw evaluation point for its algorithm change.
| Algorithm / checkpoint | Average | AIME 2026 | HMMT Feb 2026 | MATH500 |
|---|---|---|---|---|
| Current GRPO-style baseline · reproduced mean | 53.05 | 42.00 | 25.15 | 92.00 |
| CISPO token-loss · final33 | 61.36 | 56.67 | 31.82 | 95.60 |
| Dr.GRPO denom8192 · final35 | 55.51 | 46.67 | 27.27 | 92.60 |
| GSPO clip3e3 · final53 | 51.97 | 39.17 | 25.76 | 91.00 |
| DAPO overlong · final31 | 49.46 | 30.83 | 25.76 | 91.80 |
Core change
Versus current GRPO-style baseline