Related Models
Benchmarks
MMLU-Pro
40.6%
GPQA Diamond
32.7%
HLE
5.0%
LiveCodeBench
9.8%
SciCodeNot evaluated
TerminalBench HardNot evaluated
MATH-500
32.3%
AIME
0.0%
AIME 2025Not evaluated
IFBenchNot evaluated
Long Context RecallNot evaluated
Tau2Not evaluated
Market AverageTop Score