Aggregate Performance Evaluation on Aggregate 10-Benchmark Suite
79.9Average ScoreFineRouter
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| FineRouterCategory=Routers2026.03 | 79.9 | 65.2 | |
| Claude-Sonnet-4.5Category=Candidate Models2026.03 | 79.6 | 62.1 | |
| DeepSeek-R1Category=Candidate Models2026.03 | 76.9 | 55.1 | |
| IPRCategory=Routers2026.03 | 76.3 | 64.6 | |
| Claude-Haiku-4.5Category=Candidate Models2026.03 | 76 | 57.2 | |
| Qwen3-235B-A22BCategory=Candidate Models2026.03 | 76 | 53 | |
| Llama-4-MaverickCategory=Candidate Models2026.03 | 74.1 | 58 | |
| kNNCategory=Routers2026.03 | 73.6 | 62 | |
| DeepSeek-v3Category=Candidate Models2026.03 | 71.3 | 41.2 | |
| GraphRouterCategory=Routers2026.03 | 70.7 | 38.6 | |
| GPT-OSS-120BCategory=Candidate Models2026.03 | 70.3 | 36.2 | |
| RouteLLMCategory=Routers2026.03 | 67.4 | 38.7 | |
| Llama-3.3-70BCategory=Candidate Models2026.03 | 67 | 44.6 | |
| RouterDCCategory=Routers2026.03 | 66.4 | 35.3 | |
| Mistral-LargeCategory=Candidate Models2026.03 | 64.2 | 48.9 | |
| MLPCategory=Routers2026.03 | 62.9 | 44.9 | |
| Qwen3-32BCategory=Candidate Models2026.03 | 62.6 | 52.8 | |
| Mistral-SmallCategory=Candidate Models2026.03 | 50 | 40.9 | |
| DataAgentRLModel Architecture=LLaMA-DCLM2025.07 | 46.3 | — | |
| DataAgentSFTModel Architecture=LLaMA-DCLM2025.07 | 45.07 | — | |
| RegMixModel Architecture=LLaMA-DCLM2025.07 | 44.85 | — | |
| Base ModelModel Architecture=LLaMA-DCLM2025.07 | 44.52 | — | |
| DBLModel Architecture=LLaMA-DCLM2025.07 | 43.52 | — | |
| DataAgentRLModel Architecture=Pythia-1.4B2025.07 | 41.63 | — | |
| DataAgentSFTModel Architecture=Pythia-1.4B2025.07 | 40.82 | — | |
| Naive TrainingModel Architecture=LLaMA-DCLM2025.07 | 40.08 | — | |
| RegMixModel Architecture=Pythia-1.4B2025.07 | 38.96 | — | |
| Base ModelModel Architecture=Pythia-1.4B2025.07 | 38.91 | — | |
| Naive TrainingModel Architecture=Pythia-1.4B2025.07 | 35.14 | — |