Mathematical Reasoning on SVAMP (Accuracy, Improvement)
83.8Accuracy (SVAMP)Complex CoT with RICP
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Complex CoT with RICPLLM Model=GPT-3.5-Turbo, Prompting Strategy=Complex CoT, Enhancement=RICP2024.07 | 83.8 | 2.6 | |
| Few-shot CoT with RICPLLM Model=Qwen-Turbo, Prompting Strategy=Few-shot CoT, Enhancement=RICP2024.07 | 83.5 | 0.6 | |
| Zero-shot CoTLLM Model=Qwen-Turbo, Prompting Strategy=Zero-shot CoT, Enhancement=Vanilla2024.07 | 83.4 | — | |
| Zero-shot CoT with RICPLLM Model=Qwen-Turbo, Prompting Strategy=Zero-shot CoT, Enhancement=RICP2024.07 | 83.2 | -0.2 | |
| Few-shot CoTLLM Model=Qwen-Turbo, Prompting Strategy=Few-shot CoT, Enhancement=Vanilla2024.07 | 82.9 | — | |
| Auto CoT with RICPLLM Model=Qwen-Turbo, Prompting Strategy=Auto CoT, Enhancement=RICP2024.07 | 82.7 | 1 | |
| Few-shot CoT with RICPLLM Model=GPT-3.5-Turbo, Prompting Strategy=Few-shot CoT, Enhancement=RICP2024.07 | 82.7 | 2.2 | |
| Complex CoT with RICPLLM Model=Qwen-Turbo, Prompting Strategy=Complex CoT, Enhancement=RICP2024.07 | 82.4 | 1.3 | |
| Auto CoTLLM Model=Qwen-Turbo, Prompting Strategy=Auto CoT, Enhancement=Vanilla2024.07 | 81.7 | — | |
| Complex CoTLLM Model=GPT-3.5-Turbo, Prompting Strategy=Complex CoT, Enhancement=Vanilla2024.07 | 81.2 | — | |
| Complex CoTLLM Model=Qwen-Turbo, Prompting Strategy=Complex CoT, Enhancement=Vanilla2024.07 | 81.1 | — | |
| Auto CoT with RICPLLM Model=GPT-3.5-Turbo, Prompting Strategy=Auto CoT, Enhancement=RICP2024.07 | 80.9 | 2.5 | |
| Few-shot CoTLLM Model=GPT-3.5-Turbo, Prompting Strategy=Few-shot CoT, Enhancement=Vanilla2024.07 | 80.5 | — | |
| Standard Prompting with RICPLLM Model=GPT-3.5-Turbo, Prompting Strategy=Standard Prompting, Enhancement=RICP2024.07 | 79.9 | 1.5 | |
| Standard PromptingLLM Model=GPT-3.5-Turbo, Prompting Strategy=Standard Prompting, Enhancement=Vanilla2024.07 | 78.4 | — | |
| Auto CoTLLM Model=GPT-3.5-Turbo, Prompting Strategy=Auto CoT, Enhancement=Vanilla2024.07 | 78.4 | — | |
| Zero-shot CoT with RICPLLM Model=GPT-3.5-Turbo, Prompting Strategy=Zero-shot CoT, Enhancement=RICP2024.07 | 74.4 | 5.6 | |
| Zero-shot CoTLLM Model=GPT-3.5-Turbo, Prompting Strategy=Zero-shot CoT, Enhancement=Vanilla2024.07 | 68.8 | — | |
| Standard Prompting with RICPLLM Model=Qwen-Turbo, Prompting Strategy=Standard Prompting, Enhancement=RICP2024.07 | 62.6 | 0.7 | |
| Standard PromptingLLM Model=Qwen-Turbo, Prompting Strategy=Standard Prompting, Enhancement=Vanilla2024.07 | 61.9 | — |