Mathematical Reasoning on AIME 25 (Pass@1 Accuracy)
89.1Pass@1 AccuracyNemotron 3 Nano 30B-A3B
Evaluation Results
| Method | Links | |
|---|---|---|
| Nemotron 3 Nano 30B-A3BOpen-Source=✓, Size=30B-A3B, no tools=true2026.04 | 89.1 | |
| Nemotron 3 Nano OmniOpen-Source=✓, Size=30B-A3B, no tools=true2026.04 | 82.1 | |
| Qwen3-OmniOpen-Source=✓, Size=30B-A3B, no tools=true2026.04 | 73.7 | |
| ROPDThinking Mode=Thinking2026.05 | 68.75 | |
| FP16Model=Qwen3-8B, Quantization Bits=FP16, Speedup=1.0×2026.05 | 67.66 | |
| GPT-5.2-chat (teacher)2026.05 | 67.08 | |
| OVDThinking Mode=Thinking2026.05 | 65.83 | |
| T-JudgeThinking Mode=Thinking2026.05 | 65.48 | |
| KnowRL-Nemotron-1.5BHint Setting=CSS, Evaluation Protocol=mean@322026.04 | 65.21 | |
| KnowRL-Nemotron-1.5BHint Setting=CBRS, Evaluation Protocol=mean@322026.04 | 65 | |
| QuestAHint Setting=CSS, Evaluation Protocol=mean@322026.04 | 64.99 | |
| KnowRL-Nemotron-1.5BHint Setting=w/o KP, Evaluation Protocol=mean@322026.04 | 64.69 | |
| FP16Model=Qwen3-4B, Quantization Bits=FP16, Speedup=1.0×2026.05 | 63.75 | |
| JustRLHint Setting=w/o KP, Evaluation Protocol=mean@322026.04 | 62.92 | |
| JustRLHint Setting=CBRS, Evaluation Protocol=mean@322026.04 | 62.36 | |
| QuestAHint Setting=w/o KP, Evaluation Protocol=mean@322026.04 | 62.08 | |
| QuestAHint Setting=CBRS, Evaluation Protocol=mean@322026.04 | 62 | |
| JustRLHint Setting=CSS, Evaluation Protocol=mean@322026.04 | 61.43 | |
| LAQuantModel=Qwen3-8B, Quantization Bits=3-bit, Speedup=5.15×2026.05 | 61.3 | |
| GADThinking Mode=Thinking2026.05 | 61.28 | |
| Qwen3-30B-A3B (teacher)Access=–2026.05 | 61.25 | |
| Qwen3-4B (student)Thinking Mode=Thinking2026.05 | 59.58 | |
| ROPDThinking Mode=Non-Thinking2026.05 | 58.75 | |
| ParoQ++Model=Qwen3-8B, Quantization Bits=3-bit, Speedup=4.63×2026.05 | 58.39 | |
| T-JudgeThinking Mode=Non-Thinking2026.05 | 56.64 | |
| ROPDAccess=text2026.05 | 55.93 | |
| OVDThinking Mode=Non-Thinking2026.05 | 55.71 | |
| LAQuantModel=Qwen3-4B, Quantization Bits=3-bit, Speedup=3.42×2026.05 | 54.69 | |
| ESampModel=GPT-OSS-20B2026.04 | 54.2 | |
| Min-pModel=GPT-OSS-20B2026.04 | 53.8 | |
| FIREModel=GPT-OSS-20B2026.04 | 53.5 | |
| OverRIDEModel=GPT-OSS-20B2026.04 | 53.3 | |
| ParoQModel=Qwen3-8B, Quantization Bits=3-bit, Speedup=4.63×2026.05 | 52.92 | |
| ParoQ++Model=Qwen3-4B, Quantization Bits=3-bit, Speedup=3.01×2026.05 | 52.76 | |
| Nemotron-1.5BHint Setting=CSS, Evaluation Protocol=mean@322026.04 | 50.1 | |
| Nemotron-1.5BHint Setting=CBRS, Evaluation Protocol=mean@322026.04 | 49 | |
| VanillaModel=GPT-OSS-20B2026.04 | 49 | |
| Nemotron-1.5BHint Setting=w/o KP, Evaluation Protocol=mean@322026.04 | 48.33 | |
| R-QATModel=Qwen3-8B, Quantization Bits=3-bit, Speedup=5.15×2026.05 | 43.91 | |
| FIREModel=Qwen3-8B2026.04 | 42.4 | |
| ExOPDAccess=logit2026.05 | 41.25 | |
| GPTQModel=Qwen3-8B, Quantization Bits=3-bit, Speedup=5.15×2026.05 | 40.68 | |
| ParoQModel=Qwen3-4B, Quantization Bits=3-bit, Speedup=3.01×2026.05 | 39.58 | |
| ESampModel=Qwen3-8B2026.04 | 39.1 | |
| LOPDAccess=logit2026.05 | 38.75 | |
| Min-pModel=Qwen3-8B2026.04 | 38.6 | |
| VanillaModel=Qwen3-8B2026.04 | 38 | |
| OverRIDEModel=Qwen3-8B2026.04 | 37.8 | |
| FP16Model=Qwen3-1.7B, Quantization Bits=FP16, Speedup=1.0×2026.05 | 35.94 | |
| R-QATModel=Qwen3-4B, Quantization Bits=3-bit, Speedup=3.42×2026.05 | 35.31 | |
| ASLEC-CASLBackbone=Llama3-3B2026.04 | 33.33 | |
| AMS-ExpectedKV Cache Length (T_keep)=10242026.05 | 33.33 | |
| CASPOBase Model=Qwen3-8B-Base, RM Data=02026.05 | 33.3 | |
| FP16Model=R1-Distill-Llama-8B, Quantization Bits=FP16, Speedup=1.0×2026.05 | 31.2 | |
| EEPOBase Model=Qwen3-14B-Base2025.10 | 30 | |
| Full KVKV Cache Length (T_keep)=Full2026.05 | 30 | |
| LAQuantModel=Qwen3-1.7B, Quantization Bits=3-bit, Speedup=2.29×2026.05 | 29.33 | |
| LAQuantModel=R1-Distill-Llama-8B, Quantization Bits=3-bit, Speedup=5.56×2026.05 | 28.54 | |
| ParoQ++Model=Qwen3-1.7B, Quantization Bits=3-bit, Speedup=2.05×2026.05 | 28.23 | |
| R-QATModel=R1-Distill-Llama-8B, Quantization Bits=3-bit, Speedup=5.56×2026.05 | 27.5 | |
| SCRLCandidate responses=32, Training samples=16, Backbone=Qwen2.5-Math-7B2026.03 | 26.9 | |
| Re-ScheduleBackbone=Qwen3-4B-Base2025.10 | 26.9 | |
| GRPOBase Model=Qwen3-14B-Base2025.10 | 26.7 | |
| SatoriBase Model=Qwen3-8B-Base, RM Data=240K2026.05 | 26.7 | |
| SA + GRPOTraining Paradigm=Single Agent2026.05 | 26.67 | |
| MetaAgent-X RLTraining Paradigm=RL-based Auto MAS2026.05 | 26.67 | |
| AdaKV-ExpE2KV Cache Length (T_keep)=10242026.05 | 26.67 | |
| GRACEBackbone=Llama3-3B2026.04 | 26.66 | |
| RL-PLUSBase Model=Qwen2.5-Math-7B, Training Method=RL-PLUS2025.07 | 25.9 | |
| ParoQ++Model=R1-Distill-Llama-8B, Quantization Bits=3-bit, Speedup=4.90×2026.05 | 25.57 | |
| MaASTraining Paradigm=RL-based Auto MAS2026.05 | 25 | |
| R-QATModel=Qwen3-1.7B, Quantization Bits=3-bit, Speedup=2.29×2026.05 | 23.7 | |
| GADThinking Mode=Non-Thinking2026.05 | 23.34 | |
| AMS-ExpectedKV Cache Length (T_keep)=5122026.05 | 23.33 | |
| R-KVKV Cache Length (T_keep)=10242026.05 | 23.33 | |
| RPCKV Cache Length (T_keep)=10242026.05 | 23.33 | |
| ACCBackbone=Qwen3-4B-Base2025.10 | 23.3 | |
| rStar-MathBase Model=Qwen3-8B-Base, RM Data=3.64M2026.05 | 23.3 | |
| ADASTraining Paradigm=Search-based Auto MAS2026.05 | 23 | |
| SCRLCandidate responses=64, Training samples=32, Backbone=Qwen2.5-Math-7B2026.03 | 22.8 | |
| SFTAccess=text2026.05 | 22.5 | |
| SFTBase Model=Qwen2.5-Math-7B, Training Method=SFT2025.07 | 22.3 | |
| GRPOBackbone=Qwen3-4B-Base2025.10 | 21.8 | |
| Qwen3-4B (student)Thinking Mode=Non-Thinking2026.05 | 20.83 | |
| Qwen3-4B (student)Access=–2026.05 | 20.83 | |
| Qwen3-8B-BaseRM Data=–2026.05 | 20 | |
| AMS-TOVAKV Cache Length (T_keep)=10242026.05 | 20 | |
| GPTQModel=R1-Distill-Llama-8B, Quantization Bits=3-bit, Speedup=5.56×2026.05 | 19.64 | |
| SATraining Paradigm=Single Agent2026.05 | 19.1 | |
| TTRLCandidate responses=64, Training samples=32, Backbone=Qwen2.5-Math-7B2026.03 | 19 | |
| ParoQModel=Qwen3-1.7B, Quantization Bits=3-bit, Speedup=2.05×2026.05 | 16.88 | |
| TTRLCandidate responses=32, Training samples=16, Backbone=Qwen2.5-Math-7B2026.03 | 16.8 | |
| ScoreFlowTraining Paradigm=RL-based Auto MAS2026.05 | 16.7 | |
| MetaAgent-X SFTTraining Paradigm=RL-based Auto MAS2026.05 | 16.7 | |
| ChunkKV-ExpectedKV Cache Length (T_keep)=10242026.05 | 16.67 | |
| GRPOBase Model=Qwen2.5-Math-7B, Training Method=GRPO2025.07 | 15.3 | |
| Re-Schedule_sigmoidBackbone=Qwen2.5-7B, Weight Function=sigmoid2025.10 | 14 | |
| RL-PLUSBase Model=Qwen2.5-Math-1.5B, Training Method=RL-PLUS2025.07 | 13.6 | |
| AFlowTraining Paradigm=Search-based Auto MAS2026.05 | 13.33 | |
| StreamingLLMKV Cache Length (T_keep)=10242026.05 | 13.33 |