Causal Reasoning on CaLM Mathematical
93.5AccuracyGRPO
Evaluation Results
| Method | Links | |
|---|---|---|
| GRPOTraining=GRPO2026.02 | 93.5 | |
| BaseTraining=Base2026.02 | 27.2 | |
| GPT-3.5-TurboContext=Best Performance Reported in the Original Paper2026.02 | 27 |
| Method | Links | |
|---|---|---|
| GRPOTraining=GRPO2026.02 | 93.5 | |
| BaseTraining=Base2026.02 | 27.2 | |
| GPT-3.5-TurboContext=Best Performance Reported in the Original Paper2026.02 | 27 |