OOD Quality Estimation on Mathematical Reasoning OOD (Near-shift)
0.159Kendall TauTV Score
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| TV ScoreModel=Llama2-7B2024.05 | 0.159 | 0.158 | |
| TV ScoreModel=GPT2-XL2024.05 | 0.131 | 0.146 | |
| TV Score w/ DiSmoModel=GPT2-XL2024.05 | 0.123 | 0.154 | |
| TV Score w/ DiSmoModel=Llama2-7B2024.05 | 0.113 | 0.134 | |
| PerplexityModel=Llama2-7B2024.05 | 0.074 | 0.05 | |
| Input EmbeddingModel=GPT2-XL2024.05 | 0.059 | 0.098 | |
| Max Softmax Prob.Model=GPT2-XL2024.05 | 0.057 | 0.057 | |
| Max Softmax Prob.Model=Llama2-7B2024.05 | 0.038 | 0.026 | |
| Output EmbeddingModel=Llama2-7B2024.05 | 0.038 | 0.012 | |
| Input EmbeddingModel=Llama2-7B2024.05 | 0.036 | 0.115 | |
| Output EmbeddingModel=GPT2-XL2024.05 | 0.036 | 0.029 | |
| PerplexityModel=GPT2-XL2024.05 | 0.035 | 0.058 |