Strategic Game Playing on 13 Strategic Scenarios (test)
0.336TicTacToe ScoreMAFP
Evaluation Results
| Method | Links | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MAFPEvaluation Setting=Multiple Round, Attacker Model=Q3.5-35B2026.06 | 0.336 | 0.34 | 0.125 | 0.41 | 0.449 | 0.461 | 0.324 | 0.469 | 0.578 | 0.393 | 0.477 | 0.605 | 0.508 | 0.421 | |
| SREvaluation Setting=Multiple Round, Attacker Model=Q3.5-35B2026.06 | 0.328 | 0.395 | 0 | 0.359 | 0.332 | 0.234 | 0.273 | 0.227 | 0.484 | 0.438 | 0.436 | 0.516 | 0.498 | 0.348 | |
| ToMEvaluation Setting=Multiple Round, Attacker Model=Q3.5-35B2026.06 | 0.307 | 0.293 | 0 | 0.391 | 0.281 | 0.359 | 0.328 | 0.406 | 0.516 | 0.293 | 0.521 | 0.602 | 0.504 | 0.369 | |
| MAFP-LastEvaluation Setting=Multiple Round, Attacker Model=Q3.5-35B2026.06 | 0.287 | 0.344 | 0.469 | 0.297 | 0.268 | 0.234 | 0.305 | 0.5 | 0.477 | 0.148 | 0.527 | 0.502 | 0.5 | 0.374 | |
| Q3.5-35BEvaluation Setting=Single Round2026.06 | 0.258 | 0.297 | 0.031 | 0.406 | 0.34 | 0.266 | 0.277 | 0.328 | 0.531 | 0.406 | 0.385 | 0.584 | 0.5 | 0.355 | |
| Q3-1.7BEvaluation Setting=Single Round2026.06 | 0.232 | 0.477 | 0.133 | 0.406 | 0.174 | 0.531 | 0.305 | 0.348 | 0.363 | 0.496 | 0.234 | 0.473 | 0.5 | 0.359 | |
| DebateEvaluation Setting=Multiple Round, Attacker Model=Q3.5-35B2026.06 | 0.229 | 0.328 | 0.023 | 0.434 | 0.258 | 0.297 | 0.266 | 0.367 | 0.516 | 0.398 | 0.559 | 0.549 | 0.502 | 0.363 | |
| Llama-3.1-8BEvaluation Setting=Single Round2026.06 | 0.227 | 0.312 | 0.018 | 0.434 | 0.281 | 0.203 | 0.246 | 0.328 | 0.402 | 0.467 | 0.334 | 0.354 | 0.5 | 0.316 | |
| GPT-5-nanoEvaluation Setting=Single Round2026.06 | 0.164 | 0.078 | 0.031 | 0.32 | 0.281 | 0.312 | 0.219 | 0.234 | 0.43 | 0.443 | 0.447 | 0.562 | 0.5 | 0.309 |