Scientific Reasoning on MMLU-Redux
87.3Mean@1 AccuracyAgentic Proposing
Evaluation Results
| Method | Links | |
|---|---|---|
| Agentic ProposingTraining Data=Agentic-Proposer-4B Generated Problems, Training Budget=10,000 trajectories, Optimizer=GRPO2026.02 | 87.3 | |
| Qwen3-4B-Instruct-2507Protocol=zero-shot2026.02 | 84.1 |