Multi-Hop Question Answering on Synthetic Multi-Hop Benchmark 1-hop
100Avg F1Chain-of-Thought (CoT)
Evaluation Results
| Method | Links | |
|---|---|---|
| Chain-of-Thought (CoT)Context Length=1k, Backbone=Qwen3-8B2025.09 | 100 | |
| DirectContext Length=0.5k, Backbone=Qwen3-8B2025.09 | 98 | |
| Self Consistency (S-C)Context Length=8k, Backbone=Qwen3-8B2025.09 | 91 | |
| Self-Refine (S-R)Context Length=4k, Backbone=Qwen3-8B2025.09 | 81 | |
| InfoQAContext Length=10k, Backbone=Qwen3-8B2025.09 | 78 |