Strategy-based Question Answering on StrategyQA (Accuracy, Reusability, Verifiability)
69.11VerifiabilityR1
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| R1Executor=Strong Comm.2026.02 | 69.11 | — | — | |
| GemmaExecutor=Strong Comm.2026.02 | 64.02 | — | — | |
| PhiExecutor=Strong Comm.2026.02 | 63.2 | — | — | |
| DeepSeek-R12026.02 | 62 | 93 | 45 | |
| R1Executor=Full Comm.2026.02 | 61.64 | — | — | |
| Phi4-Reasoning2026.02 | 59 | 95 | 45 | |
| PhiExecutor=Full Comm.2026.02 | 58.52 | — | — | |
| GemmaExecutor=Full Comm.2026.02 | 57.05 | — | — | |
| Gemma32026.02 | 57 | 94 | 42 | |
| R1Executor=Weak Comm.2026.02 | 54.18 | — | — | |
| PhiExecutor=Weak Comm.2026.02 | 53.83 | — | — | |
| LlamaExecutor=Strong Comm.2026.02 | 51.44 | — | — | |
| GemmaExecutor=Weak Comm.2026.02 | 50.07 | — | — | |
| Llama3.12026.02 | 45 | 71 | 43 | |
| LlamaExecutor=Full Comm.2026.02 | 44.98 | — | — | |
| LlamaExecutor=Weak Comm.2026.02 | 38.52 | — | — | |
| CoTNormalized inference cost (T)=2.52026.04 | — | 87.6 | — | |
| CRITICNormalized inference cost (T)=29.32026.04 | — | 90.4 | — | |
| DirectNormalized inference cost (T)=1.02026.04 | — | 82 | — | |
| GoTNormalized inference cost (T)=35.72026.04 | — | 91.7 | — | |
| S^2RNormalized inference cost (T)=18.42026.04 | — | 91.7 | — | |
| SABANormalized inference cost (T)=9.22026.04 | — | 94.4 | — | |
| SC(k=5)Normalized inference cost (T)=12.0, k=52026.04 | — | 90.1 | — | |
| SELF-DISC.Normalized inference cost (T)=5.02026.04 | — | 92.5 | — | |
| Self-RefineNormalized inference cost (T)=6.32026.04 | — | 88.5 | — |