Mathematical Reasoning on Math 500 (Acc, Fmt)
72.03AccuracyQwen3-235B-A22b-Instruct
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen3-235B-A22b-InstructModel=Qwen3-235B-A22b-Instruct, Sampling Temperature=1.0, Max Output Length=2048, Samples per problem=162026.05 | 72.03 | 87.19 | |
| EP-GRPOBackbone=Qwen2.5-7B, Sampling Temperature=1.0, Max Output Length=2048, Samples per problem=162026.05 | 64.53 | 99.53 | |
| GRPOBackbone=Qwen2.5-7B, Sampling Temperature=1.0, Max Output Length=2048, Samples per problem=162026.05 | 61.25 | 99.22 | |
| EP-GRPOBackbone=Qwen2.5-3B, Sampling Temperature=1.0, Max Output Length=2048, Samples per problem=162026.05 | 56.88 | 99.06 | |
| GRPO (More Rollouts)Backbone=Qwen2.5-3B, Sampling Temperature=1.0, Rollouts (G)=10, Max Output Length=2048, Samples per problem=162026.05 | 55.16 | 99.22 | |
| GRPO (Higher Temp)Backbone=Qwen2.5-3B, Sampling Temperature=1.2, Rollouts (G)=8, Max Output Length=2048, Samples per problem=162026.05 | 51.41 | 97.5 | |
| GRPOBackbone=Qwen2.5-3B, Sampling Temperature=1.0, Rollouts (G)=8, Max Output Length=2048, Samples per problem=162026.05 | 50.94 | 98.28 | |
| DeepSeek-R1-671B-0528Model=DeepSeek-R1-671B-0528, Sampling Temperature=1.0, Max Output Length=2048, Samples per problem=162026.05 | 47.03 | 58.28 | |
| Qwen2.5-3B BaseBackbone=Qwen2.5-3B, Sampling Temperature=1.0, Max Output Length=2048, Samples per problem=162026.05 | 31.56 | 73.12 | |
| Qwen2.5-7B BaseBackbone=Qwen2.5-7B, Sampling Temperature=1.0, Max Output Length=2048, Samples per problem=162026.05 | 29.06 | 79.53 |