Free Question Answering on SQuAD contextual v1.1
23.2BLEUGCoT-decoding
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GCoT-decodingModel=Qwen2.5-14B, Decoding strategy=Multi-path2026.04 | 23.2 | 71.4 | |
| GCoT-decoding + SpanAlignModel=Qwen2.5-14B, Decoding strategy=Multi-path2026.04 | 21.5 | 69.6 | |
| GreedyModel=Qwen2.5-14B, Decoding strategy=Single-path2026.04 | 21.4 | 67.2 | |
| Beam SearchModel=Qwen2.5-14B, Decoding strategy=Multi-path2026.04 | 20 | 66 | |
| Temperature samplingModel=Qwen2.5-14B, Decoding strategy=Single-path2026.04 | 17.1 | 64.1 | |
| Top-k samplingModel=Qwen2.5-14B, Decoding strategy=Single-path2026.04 | 13.1 | 55.1 | |
| Self-consistency + Prompt-basedModel=Qwen2.5-14B, Decoding strategy=Multi-path2026.04 | 12.1 | 58 | |
| GCoT-decodingModel=Llama-3.1-8B, Decoding strategy=Multi-path2026.04 | 10 | 67.2 | |
| GCoT-decoding + SpanAlignModel=Llama-3.1-8B, Decoding strategy=Multi-path2026.04 | 9.2 | 62 | |
| GreedyModel=Llama-3.1-8B, Decoding strategy=Single-path2026.04 | 8.3 | 60.6 | |
| Beam SearchModel=Llama-3.1-8B, Decoding strategy=Multi-path2026.04 | 7.9 | 59.3 | |
| Temperature samplingModel=Llama-3.1-8B, Decoding strategy=Single-path2026.04 | 7.5 | 57.2 | |
| CoT-decoding + Prompt-basedModel=Qwen2.5-14B, Decoding strategy=Multi-path2026.04 | 5.8 | 50.3 | |
| Top-k samplingModel=Llama-3.1-8B, Decoding strategy=Single-path2026.04 | 5.4 | 51 | |
| GCoT-decodingModel=Gemma-7B, Decoding strategy=Multi-path2026.04 | 4.9 | 54.6 | |
| Self-consistency + Prompt-basedModel=Gemma-7B, Decoding strategy=Multi-path2026.04 | 4.2 | 36.7 | |
| GCoT-decoding + SpanAlignModel=Gemma-7B, Decoding strategy=Multi-path2026.04 | 3.9 | 48.9 | |
| GreedyModel=Gemma-7B, Decoding strategy=Single-path2026.04 | 3.3 | 42.8 | |
| Beam SearchModel=Gemma-7B, Decoding strategy=Multi-path2026.04 | 3.2 | 41.9 | |
| Self-consistency + Prompt-basedModel=Llama-3.1-8B, Decoding strategy=Multi-path2026.04 | 3.2 | 43.2 | |
| Temperature samplingModel=Gemma-7B, Decoding strategy=Single-path2026.04 | 3.1 | 40.1 | |
| Top-k samplingModel=Gemma-7B, Decoding strategy=Single-path2026.04 | 2.8 | 35.2 | |
| CoT-decoding + Prompt-basedModel=Llama-3.1-8B, Decoding strategy=Multi-path2026.04 | 1.3 | 40.9 | |
| CoT-decoding + Prompt-basedModel=Gemma-7B, Decoding strategy=Multi-path2026.04 | 0.2 | 25.7 |