Multi-fidelity Bandit Optimization on LLM-as-a-judge residual-mismatch Λ=128000 (test)
4,023.4Mean Cost-Weighted Pseudo-RegretTACC
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| TACClambda_H (high-fidelity cost ratio)=500, Total Budget (Lambda)=128000, Number of seeds=2002026.05 | 4,023.4 | 247.3 | |
| UCBlambda_H (high-fidelity cost ratio)=500, Total Budget (Lambda)=128000, Number of seeds=2002026.05 | 5,083.2 | 286.8 | |
| DNClambda_H (high-fidelity cost ratio)=500, Total Budget (Lambda)=128000, Number of seeds=2002026.05 | 5,201 | 289.7 | |
| MF-UCBlambda_H (high-fidelity cost ratio)=500, Total Budget (Lambda)=128000, Number of seeds=2002026.05 | 5,359.1 | 281.8 |