Predicting tokenizer outperformance on Pairwise Tokenizer Comparisons GPT-NEOX held-out (val)
100F1 ScoreDeviation from linear function
Evaluation Results
| Method | Links | |
|---|---|---|
| Deviation from linear functionIntrinsic metric=POWER LAW (P), Predictive model=Logistic Regression2025.06 | 100 | |
| COMPRESSIONIntrinsic metric=COMPRESSION, Predictive model=Logistic Regression2025.06 | 86 | |
| Two-stage predictive frameworkIntrinsic metric=C + P + S, Predictive model=SVM (linear kernel)2025.06 | 80 | |
| Slope of linear functionIntrinsic metric=SLOPE (S), Predictive model=Logistic Regression2025.06 | 67 | |
| Rank-frequency AUCIntrinsic metric=AUC, Predictive model=Logistic Regression2025.06 | 40 | |
| CARDINALITYIntrinsic metric=CARDINALITY (C), Predictive model=Logistic Regression2025.06 | 33 |