Model Routing on AI Tutoring student queries
100Alignment Rate (Tolerance 0.5)Premium-only (target)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Premium-only (target)relative token-weighted proxy=premium generation 100× input / 75× output vs. low-cost, evaluator calls=excluded2026.06 | 100 | 100 | 0 | 100 | |
| FairTutorrelative token-weighted proxy=premium generation 100× input / 75× output vs. low-cost, evaluator calls=excluded, escalation threshold (τ)=4.02026.06 | 92 | 28.4 | 71.6 | 18 | |
| Generic cascade routerrelative token-weighted proxy=premium generation 100× input / 75× output vs. low-cost, evaluator calls=excluded2026.06 | 88 | 32 | 68 | 28 | |
| Low-cost-onlyrelative token-weighted proxy=premium generation 100× input / 75× output vs. low-cost, evaluator calls=excluded2026.06 | 78 | 0.8 | 99.2 | 0 | |
| Naive difficulty routerrelative token-weighted proxy=premium generation 100× input / 75× output vs. low-cost, evaluator calls=excluded2026.06 | 72 | 39.7 | 60.3 | 34 |