Multi-task Language Understanding on MMLU (Accuracy, Avg.)
84.9AccuracyDAAO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DAAOLLM=LLM Pool2025.09 | 84.9 | 83.26 | |
| MasRouterLLM=LLM Pool2025.09 | 84.25 | 80.66 | |
| MaASLLM=gemini-1.5-flash2025.09 | 83.42 | 80.18 | |
| AFlowLLM=gpt-4o-mini2025.09 | 83.1 | 79.73 | |
| MaASLLM=gpt-4o-mini2025.09 | 83.01 | 80.43 | |
| AFlowLLM=gemini-1.5-flash2025.09 | 82.35 | 77.29 | |
| SC(CoT)LLM=gemini-1.5-flash2025.09 | 81.66 | 74.13 | |
| CoTLLM=gemini-1.5-flash2025.09 | 81.35 | 74.04 | |
| ComplexCoTLLM=gpt-4o-mini2025.09 | 81.05 | 75.57 | |
| SC(CoT)LLM=gpt-4o-mini2025.09 | 81.05 | 75.42 | |
| RouteLLMLLM=LLM Pool2025.09 | 81.04 | 75.5 | |
| ComplexCoTLLM=gemini-1.5-flash2025.09 | 80.74 | 73.39 | |
| VanillaLLM=qwen-2-72b2025.09 | 80.22 | 70.05 | |
| VanillaLLM=gemini-1.5-flash2025.09 | 80.04 | 74.08 | |
| ADASLLM=gemini-1.5-flash2025.09 | 79.68 | 72.05 | |
| ADASLLM=gpt-4o-mini2025.09 | 79.54 | 72.23 | |
| VanillaLLM=llama-3.1-70b2025.09 | 79.08 | 72.01 | |
| CoTLLM=gpt-4o-mini2025.09 | 78.43 | 73.64 | |
| PromptLLMLLM=LLM Pool2025.09 | 78.43 | 75.86 | |
| VanillaLLM=gpt-4o-mini2025.09 | 77.81 | 73.89 |