Benchmark prefix maximum vs elapsed time

Fable and GPT-5.6-Sol research trajectories · raw scores are faint marks · curves are elapsed-time prefix maxima.

GRPO baseline 说明:这里的 current GRPO-style baseline 不是 naive GRPO,而是包含 decoupled clip 和其他组件的当前参考配置。它仅作为对照,不代表模型只有 GRPO 时代的水平。

显示的 baselines
X 轴 · elapsed hours
Y 轴范围 · 每个 benchmark
claude-fable-5gpt-5.6-sol xhigh

Average is the unweighted mean of AIME 2026, HMMT Feb 2026, and MATH500. Evaluation time is measured when the full evaluation finishes. Baseline reproductions are excluded from prefix max. Click any raw evaluation point for its algorithm change.

Algorithm / checkpointAverageAIME 2026HMMT Feb 2026MATH500
Current GRPO-style baseline · reproduced mean53.0542.0025.1592.00
CISPO token-loss · final3361.3656.6731.8295.60
Dr.GRPO denom8192 · final3555.5146.6727.2792.60
GSPO clip3e3 · final5351.9739.1725.7691.00
DAPO overlong · final3149.4630.8325.7691.80