Social Deduction Game Gameplay on Avalon
61.3Overall Win RateOurs + Strategist
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Ours + StrategistBackend LLM=Gemini-2.5-Flash2025.10 | 61.3 | 78.5 | 35.2 | |
| StrategistBackend LLM=Gemini-2.5-Flash2025.10 | 57.4 | 77.6 | 27 | |
| Ours + ReActBackend LLM=Gemini-2.5-Flash2025.10 | 56.4 | 73.1 | 30.9 | |
| D-BOS (Est.)k (belief-planning horizon)=1, Estimation variant (Est.)=true, Evaluation protocol=against frozen policies2026.05 | 55.7 | — | — | |
| LASIBackend LLM=Gemini-2.5-Flash2025.10 | 53 | 71.8 | 25 | |
| ReConBackend LLM=Gemini-2.5-Flash2025.10 | 51.6 | 73.2 | 19.5 | |
| D-BOS (Est.)k (belief-planning horizon)=3, Estimation variant (Est.)=true, Evaluation protocol=against frozen policies2026.05 | 50.7 | — | — | |
| ReActBackend LLM=Gemini-2.5-Flash2025.10 | 48.7 | 70.5 | 16 | |
| D-BOSk (belief-planning horizon)=3, Estimation variant (Est.)=false, Evaluation protocol=against frozen policies2026.05 | 42.3 | — | — | |
| D-BOSk (belief-planning horizon)=5, Estimation variant (Est.)=false, Evaluation protocol=against frozen policies2026.05 | 41.3 | — | — | |
| D-BOSk (belief-planning horizon)=1, Estimation variant (Est.)=false, Evaluation protocol=against frozen policies2026.05 | 36.3 | — | — | |
| PPO (No shaping)Evaluation protocol=against frozen policies2026.05 | 34.3 | — | — | |
| BBMEvaluation protocol=against frozen policies2026.05 | 24.3 | — | — |