Language Understanding on MMLU Redux (test)
66.9AccuracyOurs (theory-guided context selection strategy)
Evaluation Results
| Method | Links | |
|---|---|---|
| Ours (theory-guided context selection strategy)Model=Qwen3-8B2026.02 | 66.9 | |
| ReMemModel=Qwen3-8B2026.02 | 66.8 | |
| ExpRAGModel=Qwen3-8B2026.02 | 66.6 | |
| DCModel=Qwen3-8B2026.02 | 66.5 | |
| BM25Model=Qwen3-8B2026.02 | 66 | |
| ZeroModel=Qwen3-8B2026.02 | 65.8 | |
| Ours (theory-guided context selection strategy)Model=Llama-3.1-8B2026.02 | 65 | |
| ReMemModel=Llama-3.1-8B2026.02 | 64.9 | |
| ExpRAGModel=Llama-3.1-8B2026.02 | 64.7 | |
| DCModel=Llama-3.1-8B2026.02 | 64.6 | |
| BM25Model=Llama-3.1-8B2026.02 | 64.1 | |
| ZeroModel=Llama-3.1-8B2026.02 | 63.8 |