Manim Animation Generation on ManimBench
82.2Visual SimilarityQwen 3 Coder 30B (A3B)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen 3 Coder 30B (A3B)Training Strategy=SFT, Inference Strategy=RITL2026.04 | 82.2 | 54.8 | 90 | |
| Qwen 3 Coder 30B (A3B)Training Strategy=GRPO, Inference Strategy=RITL2026.04 | 80.6 | 54.1 | 89 | |
| OpenAI GPT-4.1Training Strategy=Base, Inference Strategy=RITL2026.04 | 76.9 | 56.2 | 86 | |
| Qwen 3 Coder 30B (A3B)Training Strategy=Base, Inference Strategy=RITL2026.04 | 74.5 | 53.3 | 83 | |
| Qwen 3 Coder Next 80BTraining Strategy=Base, Inference Strategy=RITL2026.04 | 72.6 | 53.9 | 82 | |
| SeedCoder 8BTraining Strategy=GRPO, Inference Strategy=RITL2026.04 | 69.2 | 57.8 | 77 | |
| Qwen 2.5 Coder 14BTraining Strategy=SFT, Inference Strategy=RITL2026.04 | 69 | 57.1 | 77 | |
| OpenAI GPT-4.1Training Strategy=Base, Inference=Vanilla2026.04 | 68.6 | 55.9 | 77 | |
| SeedCoder 8BTraining Strategy=SFT, Inference Strategy=RITL2026.04 | 68.4 | 57.9 | 76 | |
| Qwen 3 14BTraining Strategy=Base, Inference Strategy=RITL2026.04 | 67.8 | 56.3 | 76 | |
| SeedCoder 8BTraining Strategy=Base, Inference Strategy=RITL2026.04 | 66.4 | 57.5 | 74 | |
| Qwen 2.5 Coder 14BTraining Strategy=GRPO, Inference Strategy=RITL2026.04 | 66.1 | 56.8 | 72 | |
| Qwen 2.5 Coder 14BTraining Strategy=Base, Inference Strategy=RITL2026.04 | 65.2 | 57 | 72 | |
| SeedCoder 8BTraining Strategy=GRPO, Inference=Vanilla2026.04 | 64.8 | 57.8 | 72 | |
| Qwen 3 14BTraining Strategy=GRPO, Inference Strategy=RITL2026.04 | 64 | 55.4 | 71 | |
| Qwen 3 Coder Next 80BTraining Strategy=Base, Inference=Vanilla2026.04 | 63.7 | 54.6 | 72 | |
| Mistral Small 3.2 24BTraining Strategy=GRPO, Inference Strategy=RITL2026.04 | 63.4 | 56.4 | 68 | |
| Qwen 3 Coder 30B (A3B)Training Strategy=SFT, Inference=Vanilla2026.04 | 63.2 | 54.5 | 70 | |
| Mistral Small 3.2 24BTraining Strategy=Base, Inference Strategy=RITL2026.04 | 63.1 | 55.1 | 71 | |
| Qwen 3 14BTraining Strategy=SFT, Inference Strategy=RITL2026.04 | 62.7 | 55.8 | 69 | |
| SeedCoder 8BTraining Strategy=SFT, Inference=Vanilla2026.04 | 62.3 | 58 | 69 | |
| Qwen 3 Coder 30B (A3B)Training Strategy=GRPO, Inference=Vanilla2026.04 | 61.7 | 54.6 | 68 | |
| Qwen 2.5 Coder 7BTraining Strategy=Base, Inference Strategy=RITL2026.04 | 61 | 54.6 | 66 | |
| Mistral Small 3.2 24BTraining Strategy=SFT, Inference Strategy=RITL2026.04 | 60.4 | 57.2 | 65 | |
| SeedCoder 8BTraining Strategy=Base, Inference=Vanilla2026.04 | 60.3 | 57.7 | 67 | |
| Mistral Small 3.2 24BTraining Strategy=GRPO, Inference=Vanilla2026.04 | 59.4 | 56.7 | 64 | |
| Qwen 3 Coder 30B (A3B)Training Strategy=Base, Inference=Vanilla2026.04 | 59.2 | 53.2 | 66 | |
| Qwen 2.5 Coder 7BTraining Strategy=GRPO, Inference Strategy=RITL2026.04 | 58.8 | 54.9 | 64 | |
| Qwen 3 14BTraining Strategy=SFT, Inference=Vanilla2026.04 | 58.3 | 55.6 | 65 | |
| Qwen 2.5 Coder 7BTraining Strategy=SFT, Inference Strategy=RITL2026.04 | 58 | 55.6 | 63 | |
| Ministral 3 14BTraining Strategy=GRPO, Inference Strategy=RITL2026.04 | 57.8 | 55.9 | 63 | |
| Qwen 3 14BTraining Strategy=Base, Inference=Vanilla2026.04 | 57.1 | 56.2 | 64 | |
| Mistral Small 3.2 24BTraining Strategy=Base, Inference=Vanilla2026.04 | 55.9 | 55.2 | 63 | |
| Ministral 3 14BTraining Strategy=SFT, Inference Strategy=RITL2026.04 | 55.9 | 55.8 | 61 | |
| Mistral Small 3.2 24BTraining Strategy=SFT, Inference=Vanilla2026.04 | 54.9 | 56.8 | 59 | |
| Qwen 2.5 Coder 14BTraining Strategy=SFT, Inference=Vanilla2026.04 | 54.8 | 57.1 | 60 | |
| Qwen 3 14BTraining Strategy=GRPO, Inference=Vanilla2026.04 | 54.6 | 56.6 | 60 | |
| Qwen 2.5 Coder 14BTraining Strategy=Base, Inference=Vanilla2026.04 | 52 | 56.9 | 57 | |
| Qwen 2.5 Coder 14BTraining Strategy=GRPO, Inference=Vanilla2026.04 | 51.8 | 57.1 | 57 | |
| Qwen 2.5 Coder 7BTraining Strategy=SFT, Inference=Vanilla2026.04 | 50.8 | 54.3 | 55 | |
| Ministral 3 14BTraining Strategy=Base, Inference Strategy=RITL2026.04 | 49.9 | 52.2 | 56 | |
| Qwen 2.5 Coder 7BTraining Strategy=GRPO, Inference=Vanilla2026.04 | 48.4 | 55.5 | 53 | |
| Ministral 3 8BTraining Strategy=GRPO, Inference Strategy=RITL2026.04 | 48 | 56.1 | 52 | |
| Ministral 3 8BTraining Strategy=SFT, Inference Strategy=RITL2026.04 | 47.3 | 56.9 | 51 | |
| Qwen 3 8BTraining Strategy=GRPO, Inference Strategy=RITL2026.04 | 47.3 | 56.6 | 53 | |
| Ministral 3 8BTraining Strategy=Base, Inference Strategy=RITL2026.04 | 45.9 | 52 | 53 | |
| Ministral 3 14BTraining Strategy=GRPO, Inference=Vanilla2026.04 | 44.6 | 54.5 | 48 | |
| Ministral 3 14BTraining Strategy=SFT, Inference=Vanilla2026.04 | 42.7 | 54.4 | 46 | |
| Qwen 2.5 Coder 7BTraining Strategy=Base, Inference=Vanilla2026.04 | 42 | 55.2 | 46 | |
| Ministral 3 8BTraining Strategy=SFT, Inference=Vanilla2026.04 | 40.8 | 56.8 | 44 | |
| Qwen 3 8BTraining Strategy=SFT, Inference Strategy=RITL2026.04 | 40.4 | 55.9 | 44 | |
| Ministral 3 8BTraining Strategy=GRPO, Inference=Vanilla2026.04 | 40 | 56.7 | 43 | |
| Ministral 3 14BTraining Strategy=Base, Inference=Vanilla2026.04 | 39.9 | 47.7 | 45 | |
| Meta LLaMA 3.1 8BTraining Strategy=GRPO, Inference Strategy=RITL2026.04 | 39.8 | 54.7 | 45 | |
| Qwen 3 8BTraining Strategy=Base, Inference Strategy=RITL2026.04 | 37 | 56.2 | 41 | |
| Qwen 2.5 Coder 3BTraining Strategy=SFT, Inference Strategy=RITL2026.04 | 36.8 | 53.9 | 40 | |
| Qwen 3 8BTraining Strategy=GRPO, Inference=Vanilla2026.04 | 36.4 | 54.4 | 40 | |
| Qwen 2.5 Coder 3BTraining Strategy=SFT, Inference=Vanilla2026.04 | 35.8 | 54.8 | 39 | |
| Qwen 3 4BTraining Strategy=SFT, Inference Strategy=RITL2026.04 | 35.8 | 54.8 | 40 | |
| Qwen 2.5 Coder 3BTraining Strategy=GRPO, Inference Strategy=RITL2026.04 | 35.5 | 54.8 | 39 | |
| Ministral 3 8BTraining Strategy=Base, Inference=Vanilla2026.04 | 34.5 | 51.4 | 40 | |
| Qwen 3 8BTraining Strategy=SFT, Inference=Vanilla2026.04 | 34.4 | 55.2 | 38 | |
| Meta LLaMA 3.1 8BTraining Strategy=Base, Inference Strategy=RITL2026.04 | 34.1 | 54.5 | 39 | |
| Qwen 2.5 Coder 3BTraining Strategy=Base, Inference Strategy=RITL2026.04 | 34 | 54.7 | 38 | |
| Qwen 3 4BTraining Strategy=Base, Inference Strategy=RITL2026.04 | 33.5 | 54.6 | 37 | |
| Qwen 2.5 Coder 3BTraining Strategy=GRPO, Inference=Vanilla2026.04 | 32.4 | 54.4 | 36 | |
| Meta LLaMA 3.1 8BTraining Strategy=SFT, Inference Strategy=RITL2026.04 | 31.3 | 55 | 37 | |
| Qwen 2.5 Coder 3BTraining Strategy=Base, Inference=Vanilla2026.04 | 30.8 | 54.5 | 34 | |
| Meta LLaMA 3.1 8BTraining Strategy=GRPO, Inference=Vanilla2026.04 | 30.1 | 54.3 | 33 | |
| Meta LLaMA 3.1 8BTraining Strategy=SFT, Inference=Vanilla2026.04 | 30 | 54.4 | 33 | |
| Qwen 3 4BTraining Strategy=GRPO, Inference Strategy=RITL2026.04 | 29.8 | 53.7 | 33 | |
| Meta LLaMA 3.1 8BTraining Strategy=Base, Inference=Vanilla2026.04 | 29.7 | 54.4 | 33 | |
| Qwen 3 8BTraining Strategy=Base, Inference=Vanilla2026.04 | 29.7 | 56.4 | 33 | |
| Qwen 3 4BTraining Strategy=SFT, Inference=Vanilla2026.04 | 28.4 | 54 | 31 | |
| Qwen 3 4BTraining Strategy=GRPO, Inference=Vanilla2026.04 | 28.4 | 53.9 | 31 | |
| Qwen 3 4BTraining Strategy=Base, Inference=Vanilla2026.04 | 26.3 | 54.3 | 29 | |
| Qwen 2.5 Coder 1.5BTraining Strategy=GRPO, Inference Strategy=RITL2026.04 | 25.2 | 52.2 | 29 | |
| Ministral 3 3BTraining Strategy=GRPO, Inference Strategy=RITL2026.04 | 24.4 | 55.2 | 27 | |
| Qwen 2.5 Coder 1.5BTraining Strategy=SFT, Inference=Vanilla2026.04 | 24.1 | 52.5 | 27 | |
| Qwen 2.5 Coder 1.5BTraining Strategy=Base, Inference Strategy=RITL2026.04 | 24 | 51.4 | 27 | |
| Qwen 2.5 Coder 1.5BTraining Strategy=Base, Inference=Vanilla2026.04 | 23.8 | 50.6 | 27 | |
| Qwen 2.5 Coder 1.5BTraining Strategy=SFT, Inference Strategy=RITL2026.04 | 22.8 | 51.3 | 26 | |
| Ministral 3 3BTraining Strategy=Base, Inference Strategy=RITL2026.04 | 21.2 | 28.4 | 24 | |
| Ministral 3 3BTraining Strategy=GRPO, Inference=Vanilla2026.04 | 20.9 | 53.6 | 23 | |
| Qwen 2.5 Coder 1.5BTraining Strategy=GRPO, Inference=Vanilla2026.04 | 20.6 | 52.7 | 23 | |
| Ministral 3 3BTraining Strategy=SFT, Inference=Vanilla2026.04 | 19.4 | 54.3 | 22 | |
| Ministral 3 3BTraining Strategy=SFT, Inference Strategy=RITL2026.04 | 19.4 | 53.4 | 22 | |
| LLaMA 3.2 3BTraining Strategy=GRPO, Inference Strategy=RITL2026.04 | 17.2 | 53.5 | 19 | |
| LLaMA 3.2 3BTraining Strategy=SFT, Inference=Vanilla2026.04 | 16.1 | 51 | 18 | |
| Ministral 3 3BTraining Strategy=Base, Inference=Vanilla2026.04 | 15.9 | 52.9 | 18 | |
| LLaMA 3.2 3BTraining Strategy=SFT, Inference Strategy=RITL2026.04 | 15.8 | 53.6 | 18 | |
| LLaMA 3.2 3BTraining Strategy=Base, Inference Strategy=RITL2026.04 | 14.3 | 53.2 | 16 | |
| LLaMA 3.2 3BTraining Strategy=Base, Inference=Vanilla2026.04 | 12.8 | 52.5 | 14 | |
| Qwen 2.5 0.5BTraining Strategy=GRPO, Inference Strategy=RITL2026.04 | 12.4 | 44.5 | 14 | |
| LLaMA 3.2 1BTraining Strategy=GRPO, Inference Strategy=RITL2026.04 | 12.4 | 47.7 | 14 | |
| LLaMA 3.2 1BTraining Strategy=SFT, Inference=Vanilla2026.04 | 11.8 | 44.6 | 13 | |
| Qwen 2.5 0.5BTraining Strategy=GRPO, Inference=Vanilla2026.04 | 11.6 | 41.3 | 13 | |
| Qwen 2.5 0.5BTraining Strategy=SFT, Inference Strategy=RITL2026.04 | 11.3 | 45.1 | 13 | |
| Qwen 2.5 0.5BTraining Strategy=SFT, Inference=Vanilla2026.04 | 10.9 | 40.3 | 12 | |
| LLaMA 3.2 1BTraining Strategy=SFT, Inference Strategy=RITL2026.04 | 10.8 | 47.6 | 12 |