Knowledge-based Multi-step Reasoning on Mean of WebQSP, CWQ, GSM8K, MWP, and Dr. SPIDER (test)
98.4HIT@1KDCM + Code Module
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| KDCM + Code ModuleAggregation=Mean across datasets2026.01 | 98.4 | 96.83 | 95.51 | |
| Improved knowledge distillation chain-style modelAggregation=Average across experimental datasets2026.01 | 98.4 | 96.83 | 95.51 | |
| LLM-SubKG-SumAggregation=Mean across datasets2026.01 | 92.23 | 91.89 | 90.17 | |
| LLM-SubKG-SumAggregation=Average across experimental datasets2026.01 | 92.23 | 91.89 | 90.17 | |
| Self-CheckAggregation=Mean across datasets2026.01 | 91.25 | 92.35 | 91.27 | |
| Self-CheckAggregation=Average across experimental datasets2026.01 | 91.25 | 92.35 | 91.27 | |
| KG-LLM-PRAggregation=Mean across datasets2026.01 | 91.06 | 91.78 | 90.22 | |
| KG-LLM-PRAggregation=Average across experimental datasets2026.01 | 91.06 | 91.78 | 90.22 | |
| RAGAggregation=Mean across datasets2026.01 | 90.23 | 90.28 | 90.18 | |
| RAGAggregation=Average across experimental datasets2026.01 | 90.23 | 90.28 | 90.18 |