Reasoning on ARC-E
98.1First-Token AccuracyZephyr-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| Zephyr-7BStrategy=Prefilling2025.05 | 98.1 | |
| Gemma-2-9BStrategy=Prefilling2025.05 | 96.6 | |
| Zephyr-7BStrategy=FTP2025.05 | 95.6 | |
| Llama-3.1-8BStrategy=Prefilling2025.05 | 94.7 | |
| Qwen-2-7BStrategy=Prefilling2025.05 | 94.7 | |
| Qwen-2-7BStrategy=FTP2025.05 | 94.3 | |
| Ministral-8BStrategy=Prefilling2025.05 | 93.7 | |
| Ministral-8BStrategy=FTP2025.05 | 93.6 | |
| Mistral-Nemo-12BStrategy=Prefilling2025.05 | 93.1 | |
| Mistral-Nemo-12BStrategy=FTP2025.05 | 92.6 | |
| Llama-3.1-8BStrategy=FTP2025.05 | 90.2 | |
| Phi-4-14BStrategy=FTP2025.05 | 87.8 | |
| Phi-4-14BStrategy=Prefilling2025.05 | 87.8 | |
| Gemma-7BStrategy=Prefilling2025.05 | 86.7 | |
| Gemma-2-9BStrategy=Prompt Engineering2025.05 | 81.2 | |
| Llama-3.1-8BStrategy=Prompt Engineering2025.05 | 78.7 | |
| Mistral-Nemo-12BStrategy=Prompt Engineering2025.05 | 76.5 | |
| BaseBackbone=LLaMA-2-7B, Compression Ratio=0%2025.10 | 76.3 | |
| Ministral-8BStrategy=Prompt Engineering2025.05 | 76.1 | |
| Qwen-2-7BStrategy=Prompt Engineering2025.05 | 74.1 | |
| Phi-4-14BStrategy=Prompt Engineering2025.05 | 73.1 | |
| PGSVDBackbone=LLaMA-2-7B, Compression Ratio=20%2025.10 | 70.75 | |
| Zephyr-7BStrategy=Prompt Engineering2025.05 | 65.7 | |
| LLM-PrunerBackbone=LLaMA-2-7B, Compression Ratio=20%2025.10 | 64.31 | |
| Gemma-7BStrategy=Prompt Engineering2025.05 | 58.7 | |
| Gemma-7BStrategy=FTP2025.05 | 55.9 | |
| SliceGPTBackbone=LLaMA-2-7B, Compression Ratio=20%2025.10 | 46.09 | |
| Gemma-2-9BStrategy=FTP2025.05 | 38.6 |