Question Answering on ARC-C (T, U, F, Rely metrics)
88T ScorePrompting
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| PromptingBase Model=Qwen2.5-7B-Inst2026.04 | 88 | 0.9 | 11.1 | 88.9 | |
| BinaryBase Model=Qwen2.5-7B-Inst2026.04 | 87.6 | 0 | 12.4 | 87.6 | |
| KARLBase Model=Qwen2.5-7B-Inst2026.04 | 87.2 | 2.3 | 10.5 | 89.4 | |
| SFTBase Model=Qwen2.5-7B-Inst2026.04 | 87 | 0 | 13 | 87 | |
| RLKFBase Model=Qwen2.5-7B-Inst2026.04 | 85.1 | 4.2 | 10.7 | 89.1 | |
| Ternary (TruthRL)Base Model=Qwen2.5-7B-Inst2026.04 | 83.6 | 8.2 | 8.2 | 91.1 | |
| BinaryBase Model=Llama3.1-8B-Inst2026.04 | 82.4 | 0 | 17.6 | 82.4 | |
| PromptingBase Model=Llama3.1-8B-Inst2026.04 | 81.7 | 0.3 | 18 | 82 | |
| KARLBase Model=Llama3.1-8B-Inst2026.04 | 80 | 9.1 | 10.9 | 88.3 | |
| SFTBase Model=Llama3.1-8B-Inst2026.04 | 77.6 | 0 | 22.4 | 77.6 | |
| Ternary (TruthRL)Base Model=Llama3.1-8B-Inst2026.04 | 71.9 | 18.1 | 10 | 86.7 | |
| RLKFBase Model=Llama3.1-8B-Inst2026.04 | 54.9 | 5.2 | 39.9 | 59.8 | |
| R-TuningBase Model=Qwen2.5-7B-Inst2026.04 | 52.6 | 45 | 2.4 | 77.4 | |
| R-TuningBase Model=Llama3.1-8B-Inst2026.04 | 10.3 | 87.8 | 1.9 | 21 |