Software Engineering Task Routing on SWE-Bench Verified
73.4Routing Success RateRouteLLM
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| RouteLLMRouter=BERT, t (threshold)=0, Model Pool=strong: GPT-5.3 Codex, weak: GPT-5.4 mini2026.05 | 73.4 | 100 | 0 | — | |
| RouteLLMRouter=BERT, t (threshold)=25, Model Pool=strong: GPT-5.3 Codex, weak: GPT-5.4 mini2026.05 | 73.4 | 100 | 0.47 | — | |
| HyDRAMode=conservative, τ (threshold)=0.150, Model Pool=strong: GPT-5.3 Codex, weak: GPT-5.4 mini2026.05 | 73.2 | 99.73 | 19.9 | 40 | |
| RouteLLMRouter=BERT, t (threshold)=50, Model Pool=strong: GPT-5.3 Codex, weak: GPT-5.4 mini2026.05 | 72 | 98.09 | 24.46 | — | |
| HyDRAMode=aggressive, τ (threshold)=0.300, Model Pool=strong: GPT-5.3 Codex, weak: GPT-5.4 mini2026.05 | 69.8 | 95.1 | 43.1 | 7.2 | |
| RouteLLMRouter=BERT, t (threshold)=75, Model Pool=strong: GPT-5.3 Codex, weak: GPT-5.4 mini2026.05 | 69.4 | 94.55 | 71.27 | — | |
| RouteLLMRouter=BERT, t (threshold)=100, Model Pool=strong: GPT-5.3 Codex, weak: GPT-5.4 mini2026.05 | 69.4 | 94.55 | 71.44 | — |