Mathematical Reasoning on GSM8K (AUROC)
0.651AUROCCoT-UQ
Evaluation Results
| Method | Links | |
|---|---|---|
| CoT-UQModel=Llama 3.1-8B, Strategy=AP, Base Aggregation=TOKENSAR2025.02 | 0.651 | |
| CoT-UQModel=Llama 3.1-8B, Strategy=AP, Base Aggregation=Probas-mean2025.02 | 0.6364 | |
| CoT-UQModel=Llama 3.1-8B, Strategy=AP, Base Aggregation=Probas-min2025.02 | 0.6309 | |
| CoT-UQModel=Llama 2-13B, Strategy=AP, Base Aggregation=TOKENSAR2025.02 | 0.5961 | |
| CoT-UQModel=Llama 2-13B, Strategy=AP, Base Aggregation=Probas-min2025.02 | 0.5863 | |
| CoT-UQModel=Llama 2-13B, Strategy=AP, Base Aggregation=Probas-mean2025.02 | 0.5854 | |
| CoT-UQModel=Llama 2-13B, Strategy=SE, Base Strategy=Self-Probing, Refinement=KEYStep2025.02 | 0.5514 | |
| Probas-minModel=Llama 3.1-8B, Strategy=AP2025.02 | 0.5495 | |
| TOKENSARModel=Llama 2-13B, Strategy=AP2025.02 | 0.5482 | |
| TOKENSARModel=Llama 3.1-8B, Strategy=AP2025.02 | 0.5446 | |
| Probas-meanModel=Llama 2-13B, Strategy=AP2025.02 | 0.5396 | |
| Probas-minModel=Llama 2-13B, Strategy=AP2025.02 | 0.5384 | |
| Probas-meanModel=Llama 3.1-8B, Strategy=AP2025.02 | 0.5317 | |
| Self-ProbingModel=Llama 2-13B, Strategy=SE2025.02 | 0.5272 | |
| CoT-UQModel=Llama 3.1-8B, Strategy=SE, Base Strategy=P(True), Refinement=ALLSteps2025.02 | 0.526 | |
| CoT-UQModel=Llama 2-13B, Strategy=SE, Base Strategy=P(True), Refinement=ALLSteps2025.02 | 0.5259 | |
| CoT-UQModel=Llama 3.1-8B, Strategy=SE, Base Strategy=Self-Probing, Refinement=KEYStep2025.02 | 0.5189 | |
| Self-ProbingModel=Llama 3.1-8B, Strategy=SE2025.02 | 0.4924 | |
| P(True)Model=Llama 3.1-8B, Strategy=SE2025.02 | 0.4815 | |
| P(True)Model=Llama 2-13B, Strategy=SE2025.02 | 0.4606 |