Symbolic Reasoning on MMLU-Hard abstract_algebra 100 items
41.7Baseline AccuracyUncoded single call
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Uncoded single callChannel models=Nemotron-Nano-3 + Devstral-Small-2, Judge model=GLM-5.1, n=100, Method category=Baseline2026.05 | 41.7 | — | — | |
| FountainChannel models=Nemotron-Nano-3 + Devstral-Small-2, Judge model=GLM-5.1, n=100, Method category=Best fixed technique, Agentic Framework=AGENTCODEC2026.05 | — | 66.2 | 24.4 |