Multi-turn Collaboration Reasoning on MATH-Chat
91.5AccuracyCalibrated Interactive RL
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Calibrated Interactive RLParadigm=Calibrated RL, Backbone=Gemma-3-4B-IT2026.05 | 91.5 | 1.86 | |
| Oracle (Proxy Human)Paradigm=–2026.05 | 89.7 | 1.27 | |
| Naive Interactive RLParadigm=Interactive RL, Backbone=Gemma-3-4B-IT2026.05 | 89.3 | 1.97 | |
| Static ContextParadigm=Static RL, Backbone=Gemma-3-4B-IT2026.05 | 85 | 1.63 | |
| Gemma-3-4B-ITParadigm=Base Model, Backbone=Gemma-3-4B-IT2026.05 | 82.3 | 1.76 | |
| COLLABLLMParadigm=Offline DPO, Backbone=Gemma-3-4B-IT2026.05 | 82.3 | 1.63 |