Question Answering on QA ExAnte (test)
1.61d Leakage RateTCFT
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| TCFTBackbone=Qwen-2.5-7B-Instruct, Evaluation Protocol=Ours2026.05 | 1.6 | 0 | 0.65 | 0.75 | |
| TCFTBackbone=Qwen-2.5-14B-Instruct, Evaluation Protocol=Ours2026.05 | 5.99 | 6.02 | 5.39 | 5.8 | |
| Few-shotBackbone=Qwen-2.5-7B-Instruct, Evaluation Protocol=Few-shot2026.05 | 20.8 | 26.23 | 18.95 | 21.75 | |
| SFTBackbone=Qwen-2.5-14B-Instruct, Evaluation Protocol=SFT2026.05 | 20.8 | 25.41 | 17 | 20.75 | |
| Few-shotBackbone=Qwen-2.5-14B-Instruct, Evaluation Protocol=Few-shot2026.05 | 23.35 | 20.48 | 13.77 | 19.2 | |
| Self-verificationBackbone=Qwen-2.5-14B-Instruct, Evaluation Protocol=Self-verification2026.05 | 24.55 | 11.45 | 9.58 | 15.2 | |
| Self-verificationBackbone=Qwen-2.5-7B-Instruct, Evaluation Protocol=Self-verification2026.05 | 30.4 | 33.61 | 25.49 | 29.5 | |
| SFTBackbone=Qwen-2.5-7B-Instruct, Evaluation Protocol=SFT2026.05 | 35.2 | 41.8 | 30.72 | 35.5 | |
| Gemini-3.1-Pro-PreviewModel=Gemini-3.1-Pro-Preview2026.05 | 44 | 41.8 | 40.52 | 42 | |
| Claude-Opus-4.6Model=Claude-Opus-4.62026.05 | 48 | 11.48 | 5.23 | 20.5 | |
| GPT-5.4Model=GPT-5.42026.05 | 49.6 | 13.11 | 7.19 | 22.25 | |
| Chain-of-thoughtBackbone=Qwen-2.5-7B-Instruct, Evaluation Protocol=Chain-of-thought2026.05 | 59.2 | 58.2 | 51.63 | 56 | |
| Chain-of-thoughtBackbone=Qwen-2.5-14B-Instruct, Evaluation Protocol=Chain-of-thought2026.05 | 74.25 | 56.02 | 56.29 | 62.2 | |
| Zero-shotBackbone=Qwen-2.5-7B-Instruct, Evaluation Protocol=Zero-shot2026.05 | 80.8 | 76.23 | 73.86 | 76.75 | |
| Zero-shotBackbone=Qwen-2.5-14B-Instruct, Evaluation Protocol=Zero-shot2026.05 | 81.44 | 65.66 | 62.28 | 69.8 |