Multi-Agent Pathfinding on 5-agent evaluation set (9x9)
10Valid RateQwen3-4B + GRPO + LLM-as-Environment-Engineer
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen3-4B + GRPO + LLM-as-Environment-Engineer2026.06 | 10 | 2.67 | |
| Qwen3-4B + GRPOtraining_config=random2026.06 | 8.67 | 2.67 | |
| Kimi-K2.52026.06 | 8 | 4.67 | |
| Grok-4.22026.06 | 7.33 | 5.33 | |
| GPT-5.42026.06 | 5.33 | 4 | |
| Gemini-3.1-Pro2026.06 | 2 | 1.33 | |
| Qwen3-4Bvariant=base2026.06 | 0.67 | 0.67 |