OOD Quality Estimation on Mathematical Reasoning Far-shift OOD
0.161Kendall's TauTV Score
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| TV ScoreModel=Llama2-7B2024.05 | 0.161 | 0.147 | |
| TV Score w/ DiSmoModel=GPT2-XL2024.05 | 0.139 | 0.141 | |
| TV ScoreModel=GPT2-XL2024.05 | 0.138 | 0.123 | |
| TV Score w/ DiSmoModel=Llama2-7B2024.05 | 0.111 | 0.152 | |
| Input EmbeddingModel=Llama2-7B2024.05 | 0.078 | 0.102 | |
| Max Softmax Prob.Model=GPT2-XL2024.05 | 0.066 | 0.044 | |
| Output EmbeddingModel=Llama2-7B2024.05 | 0.058 | 0.025 | |
| Input EmbeddingModel=GPT2-XL2024.05 | 0.058 | 0.068 | |
| PerplexityModel=Llama2-7B2024.05 | 0.05 | 0.045 | |
| Output EmbeddingModel=GPT2-XL2024.05 | 0.05 | 0.016 | |
| PerplexityModel=GPT2-XL2024.05 | 0.036 | 0.038 | |
| Max Softmax Prob.Model=Llama2-7B2024.05 | 0.024 | 0.038 |