Multi-agent policy synthesis on Cleanup
2.75U ScoreGemini 3.1 Pro
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Gemini 3.1 ProFeedback=reward+social2026.03 | 2.75 | 0.54 | 432.6 | |
| Gemini 3.1 ProFeedback=reward-only2026.03 | 1.79 | 0.13 | 386 | |
| Claude Sonnet 4.6Feedback=reward+social2026.03 | 1.37 | 0.09 | 294.6 | |
| Claude Sonnet 4.6Feedback=reward-only2026.03 | 1.14 | 0.47 | 233 | |
| Claude Sonnet 4.6Feedback=zero-shot2026.03 | 1.01 | 3.06 | 137 | |
| GEPA (Gemini 3.1 Pro)2026.03 | 0.77 | 1.75 | 209.5 | |
| Gemini 3.1 ProFeedback=zero-shot2026.03 | 0.45 | 0.45 | 274.1 | |
| Q-learner2026.03 | 0.16 | 0.2 | 208.6 | |
| BFS Collector2026.03 | 0.1 | 0.61 | 16.4 |