Chain-of-Thought Attribution Faithfulness on GSM8K (test)
232.39AUPCAttriCoT-2x
Evaluation Results
| Method | Links | |
|---|---|---|
| AttriCoT-2xCoT Model=DS-Llama-8B, Forward Pass Count=2(S + T)2026.06 | 232.39 | |
| AttriCoTCoT Model=DS-Llama-8B, Forward Pass Count=S + T2026.06 | 231 | |
| AttriCoT-2xCoT Model=DS-Qwen-14B, Forward Pass Count=2(S + T)2026.06 | 218.55 | |
| AttriCoTCoT Model=DS-Qwen-14B, Forward Pass Count=S + T2026.06 | 216.7 | |
| TA-KL (no log)CoT Model=DS-Llama-8B, Log Transformation=false2026.06 | 195.61 | |
| TA-KLCoT Model=DS-Llama-8B, Log Transformation=true2026.06 | 183.29 | |
| TA-KL (no log)CoT Model=DS-Qwen-14B, Log Transformation=false2026.06 | 181.82 | |
| TA-KLCoT Model=DS-Qwen-14B, Log Transformation=true2026.06 | 169.47 | |
| AttriCoT-2xCoT Model=Qwen3-8B, Forward Pass Count=2(S + T)2026.06 | 168.03 | |
| AttriCoTCoT Model=Qwen3-8B, Forward Pass Count=S + T2026.06 | 166.59 | |
| TA-KL (no log)CoT Model=Qwen3-8B, Log Transformation=false2026.06 | 161.28 | |
| Prompted-ICL (counterpart)CoT Model=DS-Qwen-14B, Prompting Strategy=counterpart, In-Context Learning (ICL)=true2026.06 | 153.51 | |
| Prompted (counterpart)CoT Model=DS-Qwen-14B, Prompting Strategy=counterpart, In-Context Learning (ICL)=false2026.06 | 147.27 | |
| Prompted-ICL (GPT-OSS)CoT Model=DS-Qwen-14B, Prompting Strategy=GPT-OSS-120B, In-Context Learning (ICL)=true2026.06 | 142.58 | |
| AttriCoT-2xCoT Model=DS-Qwen3-8B, Forward Pass Count=2(S + T)2026.06 | 139.73 | |
| AttriCoTCoT Model=DS-Qwen3-8B, Forward Pass Count=S + T2026.06 | 138.64 | |
| Prompted-ICL (GPT-OSS)CoT Model=DS-Llama-8B, Prompting Strategy=GPT-OSS-120B, In-Context Learning (ICL)=true2026.06 | 135.79 | |
| TA-KL (no log)CoT Model=DS-Qwen3-8B, Log Transformation=false2026.06 | 123.42 | |
| Prompted-ICL (self)CoT Model=DS-Qwen-14B, Prompting Strategy=self, In-Context Learning (ICL)=true2026.06 | 117.77 | |
| TA-KLCoT Model=DS-Qwen3-8B, Log Transformation=true2026.06 | 116.05 | |
| Attention-based (trained)CoT Model=DS-Qwen-14B, trained=true2026.06 | 111.39 | |
| Prompted-ICL (self)CoT Model=DS-Llama-8B, Prompting Strategy=self, In-Context Learning (ICL)=true2026.06 | 107.29 | |
| Attention-basedCoT Model=DS-Llama-8B, trained=false2026.06 | 107.06 | |
| Attention-based (trained)CoT Model=DS-Llama-8B, trained=true2026.06 | 101.43 | |
| TA-KLCoT Model=Qwen3-8B, Log Transformation=true2026.06 | 98.53 | |
| Prompted (GPT-OSS)CoT Model=DS-Llama-8B, Prompting Strategy=GPT-OSS-120B, In-Context Learning (ICL)=false2026.06 | 96.23 | |
| Prompted (self)CoT Model=DS-Qwen-14B, Prompting Strategy=self, In-Context Learning (ICL)=false2026.06 | 95.52 | |
| Prompted (GPT-OSS)CoT Model=DS-Qwen-14B, Prompting Strategy=GPT-OSS-120B, In-Context Learning (ICL)=false2026.06 | 94.87 | |
| Prompted (self)CoT Model=DS-Llama-8B, Prompting Strategy=self, In-Context Learning (ICL)=false2026.06 | 88.65 | |
| Attention-basedCoT Model=DS-Qwen-14B, trained=false2026.06 | 87.36 | |
| Prompted-ICL (GPT-OSS)CoT Model=Qwen3-8B, Prompting Strategy=GPT-OSS-120B, In-Context Learning (ICL)=true2026.06 | 72.06 | |
| Prompted (self)CoT Model=Qwen3-8B, Prompting Strategy=self, In-Context Learning (ICL)=false2026.06 | 67.45 | |
| Attention-basedCoT Model=DS-Qwen3-8B, trained=false2026.06 | 65.72 | |
| Prompted-ICL (GPT-OSS)CoT Model=DS-Qwen3-8B, Prompting Strategy=GPT-OSS-120B, In-Context Learning (ICL)=true2026.06 | 62.06 | |
| Attention-based (trained)CoT Model=DS-Qwen3-8B, trained=true2026.06 | 61.88 | |
| Prompted (self)CoT Model=DS-Qwen3-8B, Prompting Strategy=self, In-Context Learning (ICL)=false2026.06 | 61.39 | |
| Prompted-ICL (self)CoT Model=DS-Qwen3-8B, Prompting Strategy=self, In-Context Learning (ICL)=true2026.06 | 60.38 | |
| Prompted (counterpart)CoT Model=DS-Qwen3-8B, Prompting Strategy=counterpart, In-Context Learning (ICL)=false2026.06 | 60.2 | |
| Prompted-ICL (counterpart)CoT Model=DS-Qwen3-8B, Prompting Strategy=counterpart, In-Context Learning (ICL)=true2026.06 | 58.39 | |
| Attention-basedCoT Model=Qwen3-8B, trained=false2026.06 | 58.31 | |
| Attention-based (trained)CoT Model=Qwen3-8B, trained=true2026.06 | 56.02 | |
| Prompted (GPT-OSS)CoT Model=Qwen3-8B, Prompting Strategy=GPT-OSS-120B, In-Context Learning (ICL)=false2026.06 | 54.78 | |
| Prompted-ICL (self)CoT Model=Qwen3-8B, Prompting Strategy=self, In-Context Learning (ICL)=true2026.06 | 50.49 | |
| Prompted (GPT-OSS)CoT Model=DS-Qwen3-8B, Prompting Strategy=GPT-OSS-120B, In-Context Learning (ICL)=false2026.06 | 49.65 |