Mathematical Reasoning on ASDiv (AUROC)
66.91AUROCCoT-UQ
Evaluation Results
| Method | Links | |
|---|---|---|
| CoT-UQModel=Llama 3.1-8B, Strategy=AP, Base Aggregation=TOKENSAR2025.02 | 66.91 | |
| CoT-UQModel=Llama 3.1-8B, Strategy=AP, Base Aggregation=Probas-min2025.02 | 64.84 | |
| CoT-UQModel=Llama 3.1-8B, Strategy=AP, Base Aggregation=Probas-mean2025.02 | 64.52 | |
| CoT-UQModel=Llama 2-13B, Strategy=AP, Base Aggregation=TOKENSAR2025.02 | 61.23 | |
| CoT-UQModel=Llama 2-13B, Strategy=AP, Base Aggregation=Probas-min2025.02 | 60.74 | |
| CoT-UQModel=Llama 2-13B, Strategy=AP, Base Aggregation=Probas-mean2025.02 | 59.44 | |
| TOKENSARModel=Llama 3.1-8B, Strategy=AP2025.02 | 58.71 | |
| Probas-minModel=Llama 3.1-8B, Strategy=AP2025.02 | 58.69 | |
| TOKENSARModel=Llama 2-13B, Strategy=AP2025.02 | 58.37 | |
| Probas-meanModel=Llama 3.1-8B, Strategy=AP2025.02 | 58.34 | |
| Probas-meanModel=Llama 2-13B, Strategy=AP2025.02 | 57.73 | |
| Probas-minModel=Llama 2-13B, Strategy=AP2025.02 | 57.7 | |
| CoT-UQModel=Llama 2-13B, Strategy=SE, Base Strategy=Self-Probing, Refinement=KEYStep2025.02 | 57.58 | |
| CoT-UQModel=Llama 2-13B, Strategy=SE, Base Strategy=P(True), Refinement=ALLSteps2025.02 | 56.1 | |
| CoT-UQModel=Llama 3.1-8B, Strategy=SE, Base Strategy=Self-Probing, Refinement=KEYStep2025.02 | 53.79 | |
| CoT-UQModel=Llama 3.1-8B, Strategy=SE, Base Strategy=P(True), Refinement=ALLSteps2025.02 | 53.2 | |
| Self-ProbingModel=Llama 2-13B, Strategy=SE2025.02 | 52.35 | |
| Self-ProbingModel=Llama 3.1-8B, Strategy=SE2025.02 | 50.86 | |
| P(True)Model=Llama 2-13B, Strategy=SE2025.02 | 48.02 | |
| P(True)Model=Llama 3.1-8B, Strategy=SE2025.02 | 47.23 |