Code Synthesis on HumanEval-Hard (100-item)
75.7Baseline ScoreUncoded single call
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Uncoded single callChannel models=Nemotron-Nano-3 + Devstral-Small-2, Judge model=GLM-5.1, n=100, Method category=Baseline2026.05 | 75.7 | — | — | |
| HARQ-IRChannel models=Nemotron-Nano-3 + Devstral-Small-2, Judge model=GLM-5.1, n=100, Method category=Best fixed technique, Agentic Framework=AGENTCODEC2026.05 | — | 86 | 10.3 |