Knowledge-intensive Reasoning on 2Wiki (F1, EM, TC)
52F1 ScoreEAPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| EAPOBackbone=Llama3.1-8B-Instruct2026.06 | 52 | 44.8 | 1.69 | |
| GRPOBackbone=Llama3.1-8B-Instruct2026.06 | 48 | 40.7 | 3.27 | |
| Reinforce++Backbone=Llama3.1-8B-Instruct2026.06 | 47.1 | 38.4 | 2.13 | |
| TIRBackbone=Llama3.1-8B-Instruct2026.06 | 33.7 | 28.6 | 2.98 | |
| BaseBackbone=Llama3.1-8B-Instruct2026.06 | 23.7 | 10.2 | — |