Mathematical Reasoning on SVAMP (AUROC)
0.6211AUROCCoT-UQ
Evaluation Results
| Method | Links | |
|---|---|---|
| CoT-UQModel=Llama 3.1-8B, Strategy=AP, Base Aggregation=TOKENSAR2025.02 | 0.6211 | |
| CoT-UQModel=Llama 3.1-8B, Strategy=AP, Base Aggregation=Probas-min2025.02 | 0.6049 | |
| CoT-UQModel=Llama 3.1-8B, Strategy=SE, Base Strategy=P(True), Refinement=ALLSteps2025.02 | 0.6 | |
| CoT-UQModel=Llama 3.1-8B, Strategy=AP, Base Aggregation=Probas-mean2025.02 | 0.5983 | |
| CoT-UQModel=Llama 2-13B, Strategy=AP, Base Aggregation=Probas-min2025.02 | 0.5851 | |
| CoT-UQModel=Llama 2-13B, Strategy=AP, Base Aggregation=TOKENSAR2025.02 | 0.5841 | |
| CoT-UQModel=Llama 2-13B, Strategy=AP, Base Aggregation=Probas-mean2025.02 | 0.5737 | |
| CoT-UQModel=Llama 2-13B, Strategy=SE, Base Strategy=P(True), Refinement=ALLSteps2025.02 | 0.5687 | |
| Probas-minModel=Llama 2-13B, Strategy=AP2025.02 | 0.5509 | |
| TOKENSARModel=Llama 2-13B, Strategy=AP2025.02 | 0.5506 | |
| TOKENSARModel=Llama 3.1-8B, Strategy=AP2025.02 | 0.5501 | |
| Probas-minModel=Llama 3.1-8B, Strategy=AP2025.02 | 0.5479 | |
| Probas-meanModel=Llama 2-13B, Strategy=AP2025.02 | 0.5448 | |
| CoT-UQModel=Llama 3.1-8B, Strategy=SE, Base Strategy=Self-Probing, Refinement=KEYStep2025.02 | 0.5426 | |
| Probas-meanModel=Llama 3.1-8B, Strategy=AP2025.02 | 0.5394 | |
| Self-ProbingModel=Llama 3.1-8B, Strategy=SE2025.02 | 0.5163 | |
| P(True)Model=Llama 3.1-8B, Strategy=SE2025.02 | 0.5158 | |
| CoT-UQModel=Llama 2-13B, Strategy=SE, Base Strategy=Self-Probing, Refinement=KEYStep2025.02 | 0.5053 | |
| Self-ProbingModel=Llama 2-13B, Strategy=SE2025.02 | 0.4727 | |
| P(True)Model=Llama 2-13B, Strategy=SE2025.02 | 0.4636 |