Mathematical Reasoning on AQUA-RAT
91.73AccuracyQ-Opt + P-Opt
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Q-Opt + P-OptImplementation Model=Gemini-2.5-Pro, Target LLM=Gemini-2.0-Flash2026.03 | 91.73 | — | |
| Q + P-OptImplementation Model=Gemini-2.5-Pro, Target LLM=Gemini-2.0-Flash2026.03 | 90.34 | — | |
| Q-Opt + CoTImplementation Model=Gemini-2.5-Pro, Target LLM=Gemini-2.0-Flash2026.03 | 89.67 | — | |
| MARSImplementation Model=Gemini-2.5-Pro, Target LLM=Gemini-2.0-Flash2026.03 | 89.15 | — | |
| PE2Implementation Model=Gemini-2.5-Pro, Target LLM=Gemini-2.0-Flash2026.03 | 88.23 | — | |
| Q-OptImplementation Model=Gemini-2.5-Pro, Target LLM=Gemini-2.0-Flash2026.03 | 87.86 | — | |
| VecCISC + KMeansLLM Backbone=Llama3.3 70B2026.05 | 87.7 | — | |
| CISCLLM Backbone=Llama3.3 70B2026.05 | 87.6 | — | |
| VecCISC + HACLLM Backbone=Llama3.3 70B2026.05 | 87.6 | — | |
| CoTImplementation Model=Gemini-2.5-Pro, Target LLM=Gemini-2.0-Flash2026.03 | 87.4 | — | |
| VecCISC (random)LLM Backbone=Llama3.3 70B2026.05 | 87.3 | — | |
| OPROImplementation Model=Gemini-2.5-Pro, Target LLM=Gemini-2.0-Flash2026.03 | 86.78 | — | |
| No PromptImplementation Model=Gemini-2.5-Pro, Target LLM=Gemini-2.0-Flash2026.03 | 86.61 | — | |
| SC BaselineLLM Backbone=Llama3.3 70B2026.05 | 86.6 | — | |
| APEImplementation Model=Gemini-2.5-Pro, Target LLM=Gemini-2.0-Flash2026.03 | 85.92 | — | |
| CISCLLM Backbone=Qwen2.5 7B2026.05 | 85.5 | — | |
| VecCISC + KMeansLLM Backbone=Qwen2.5 7B2026.05 | 85.5 | — | |
| VecCISC + HACLLM Backbone=Qwen2.5 7B2026.05 | 85.5 | — | |
| VecCISC + KMeansLLM Backbone=GPT 4o-mini2026.05 | 84.6 | — | |
| SC BaselineLLM Backbone=Qwen2.5 7B2026.05 | 84.5 | — | |
| VecCISC + HACLLM Backbone=GPT 4o-mini2026.05 | 84.3 | — | |
| CISCLLM Backbone=GPT 4o-mini2026.05 | 84 | — | |
| VecCISC (random)LLM Backbone=GPT 4o-mini2026.05 | 83.7 | — | |
| SC BaselineLLM Backbone=GPT 4o-mini2026.05 | 83.5 | — | |
| SCBackbone=Qwen-2.5-72B-Instruct2025.05 | 83.5 | — | |
| VecCISC + KMeansLLM Backbone=Llama3.1 8B2026.05 | 83 | — | |
| VecCISC + HACLLM Backbone=Llama3.1 8B2026.05 | 83 | — | |
| CISCLLM Backbone=Llama3.1 8B2026.05 | 82.9 | — | |
| SC BaselineLLM Backbone=Llama3.1 8B2026.05 | 82.6 | — | |
| VecCISC (random)LLM Backbone=Qwen2.5 7B2026.05 | 82.6 | — | |
| DyLANBackbone=Qwen-2.5-72B-Instruct2025.05 | 82.3 | — | |
| MAS-GPTBackbone=Llama-3.3-70B-Instruct2025.05 | 80.7 | — | |
| SCBackbone=Llama-3.3-70B-Instruct2025.05 | 80.3 | — | |
| DebateBackbone=Llama-3.3-70B-Instruct2025.05 | 80.3 | — | |
| CoTBackbone=Qwen-2.5-72B-Instruct2025.05 | 80.3 | — | |
| DebateBackbone=Qwen-2.5-72B-Instruct2025.05 | 80.3 | — | |
| MADBackbone=Qwen-2.5-72B-Instruct2025.05 | 80.3 | — | |
| MacNetBackbone=Qwen-2.5-72B-Instruct2025.05 | 80.3 | — | |
| AutoGenBackbone=Llama-3.3-70B-Instruct2025.05 | 79.5 | — | |
| AgentVerseBackbone=Llama-3.3-70B-Instruct2025.05 | 79.5 | — | |
| SingleBackbone=Qwen-2.5-72B-Instruct2025.05 | 79.5 | — | |
| AgentVerseBackbone=Qwen-2.5-72B-Instruct2025.05 | 79.5 | — | |
| AFlow-MathBackbone=Qwen-2.5-72B-Instruct2025.05 | 78.7 | — | |
| MADBackbone=Llama-3.3-70B-Instruct2025.05 | 78.3 | — | |
| DyLANBackbone=Llama-3.3-70B-Instruct2025.05 | 78.3 | — | |
| AutoGenBackbone=Qwen-2.5-72B-Instruct2025.05 | 78.3 | — | |
| MAS-GPTBackbone=Qwen-2.5-72B-Instruct2025.05 | 78.3 | — | |
| MacNetBackbone=Llama-3.3-70B-Instruct2025.05 | 77.2 | — | |
| AFlow-MathBackbone=Llama-3.3-70B-Instruct2025.05 | 77.2 | — | |
| CoTBackbone=Llama-3.3-70B-Instruct2025.05 | 76.8 | — | |
| SingleBackbone=Llama-3.3-70B-Instruct2025.05 | 76 | — | |
| Few-shot-CoT-CP (GPT-4) + SCBase Model=GPT-4, Prompting Strategy=Few-shot-CoT-CP, Self-Consistency=true2024.03 | 70.9 | — | |
| VecCISC (random)LLM Backbone=Llama3.1 8B2026.05 | 69.9 | — | |
| CoK + SCBase Model=gpt-3.5-turbo2023.06 | 69.7 | — | |
| Few-shot-CoT-CP (GPT-4)Base Model=GPT-4, Prompting Strategy=Few-shot-CoT-CP2024.03 | 66.9 | — | |
| Manual CoT + SCBase Model=gpt-3.5-turbo2023.06 | 66.8 | — | |
| Zero-shot-CoT + SCBase Model=GPT-3.5-Turbo, Prompting Strategy=Zero-shot-CoT, Self-Consistency=true2024.03 | 66.1 | — | |
| MAVBackbone=Llama-3.3-70B-Instruct2025.05 | 65.8 | — | |
| ComplexCoT + SCBase Model=gpt-3.5-turbo2023.06 | 65 | — | |
| Zero-shot-CoT-CPBase Model=GPT-3.5-Turbo, Prompting Strategy=Zero-shot-CoT-CP2024.03 | 60.6 | — | |
| CoKBase Model=gpt-3.5-turbo2023.06 | 60.2 | — | |
| Few-shot-CoT + SCBase Model=GPT-3.5-Turbo, Prompting Strategy=Few-shot-CoT, Self-Consistency=true2024.03 | 59.4 | — | |
| Low-gap SelectionModel=LLaMA-3.1 (8B), Zero-shot=true2026.01 | 59.06 | — | |
| Persona SwitchModel=LLaMA-3.1 (8B), Zero-shot=true2026.01 | 59.06 | — | |
| Few-shot-CoT (GPT-4)Base Model=GPT-4, Prompting Strategy=Few-shot-CoT2024.03 | 58.7 | — | |
| Few-shot-PoT-SC (Codex)Base Model=Codex, Prompting Strategy=Few-shot-PoT-SC2024.03 | 58.6 | — | |
| Random SelectionModel=LLaMA-3.1 (8B), Zero-shot=true2026.01 | 58.4 | — | |
| NEXABackbone=Qwen2.5-1.5B-Instruct, Number of agents=102026.05 | 57.74 | — | |
| AgentPruneBackbone=Qwen2.5-1.5B-Instruct, Number of agents=102026.05 | 57.58 | — | |
| Contrastive CoTBase Model=GPT-3.5-Turbo2024.03 | 57.5 | — | |
| GDesignerBackbone=Qwen2.5-1.5B-Instruct, Number of agents=102026.05 | 57.13 | — | |
| SCBackbone=Qwen2.5-1.5B-Instruct, Prompting strategy=Self-Consistency2026.05 | 56.82 | — | |
| ComplexCoTBase Model=gpt-3.5-turbo2023.06 | 56.5 | — | |
| SelfOrg*Backbone=Qwen2.5-1.5B-Instruct, Number of agents=10, Communication=Single sequential round2026.05 | 56.46 | — | |
| GPTSwarmBackbone=Qwen2.5-1.5B-Instruct, Number of agents=102026.05 | 55.91 | — | |
| Zero-shot-CoTBase Model=GPT-3.5-Turbo, Prompting Strategy=Zero-shot-CoT2024.03 | 55.9 | — | |
| DECOVECModel=Qwen2-7B, Protocol=Few-Shot (Rand. Std.)2026.04 | 55.51 | 1.38 | |
| Few-shot-CoTBase Model=GPT-3.5-Turbo, Prompting Strategy=Few-shot-CoT2024.03 | 55.5 | — | |
| MultinomialModel=LLaMA-3.1 (8B), Zero-shot=true2026.01 | 55.12 | — | |
| Manual CoTBase Model=gpt-3.5-turbo2023.06 | 55.1 | — | |
| INSTINCTPrompt Mode=zero-shot CoT, Prompt=I have a new solution.2024.03 | 54.724 | — | |
| ZOPOPrompt Mode=zero-shot CoT, Prompt=Let's find the solution by breaking down the problem.2024.03 | 54.724 | — | |
| Top-pModel=LLaMA-3.1 (8B), Zero-shot=true2026.01 | 54.72 | — | |
| CoTBackbone=Qwen2.5-1.5B-Instruct, Prompting strategy=Chain-of-Thought2026.05 | 54.46 | — | |
| InstructZeroPrompt Mode=zero-shot CoT, Prompt=Let's break down the problem.2024.03 | 54.331 | — | |
| Few-shot-PoT (Codex)Base Model=Codex, Prompting Strategy=Few-shot-PoT2024.03 | 54.1 | — | |
| DECOVECModel=Qwen2-7B, Protocol=Few-Shot (BM25)2026.04 | 53.54 | 0.59 | |
| DECOVECModel=Qwen2-7B, Protocol=Few-Shot (Rand. Ext.)2026.04 | 53.15 | 2.15 | |
| EvoPromptPrompt Mode=zero-shot CoT, Prompt=Let's utilize the substitution method to find a solution, then try it out together.2024.03 | 52.756 | — | |
| SingleBackbone=Qwen2.5-1.5B-Instruct, Number of agents=12026.05 | 52.62 | — | |
| hand-craftPrompt Mode=zero-shot CoT, Prompt=Let's think step by step.2024.03 | 52.362 | — | |
| Few-shot-CPBase Model=GPT-3.5-Turbo, Prompting Strategy=Few-shot-CP2024.03 | 52 | — | |
| Self-consistency (Code-davinci-002)Base Model=Code-davinci-0022024.03 | 52 | — | |
| GreedyModel=LLaMA-3.1 (8B), Zero-shot=true2026.01 | 51.97 | — | |
| Top-kModel=LLaMA-3.1 (8B), Zero-shot=true2026.01 | 51.97 | — | |
| DECOVECModel=Qwen2-7B, Protocol=Few-Shot (KATE)2026.04 | 49.21 | 5.49 | |
| Zero-shot-CP + SCBase Model=GPT-3.5-Turbo, Prompting Strategy=Zero-shot-CP, Self-Consistency=true2024.03 | 48.4 | — | |
| Few-shot-CoT + self consistencyModel=PaLM (540B), Prompting Strategy=Few-shot, Chain-of-thought=true, Self-consistency=40 paths2022.05 | 48.3 | — | |
| MAVBackbone=Qwen-2.5-72B-Instruct2025.05 | 48 | — | |
| DECOVECModel=Gemma-2-9B, Protocol=Few-Shot (KATE)2026.04 | 47.24 | 2.36 |