User Simulation on MirrorBench
0.713Realism Score (LLM-judge)DITTO
Evaluation Results
| Method | Links | |
|---|---|---|
| DITTOBackbone=Qwen3-VL-8B-Instruct2026.05 | 0.713 | |
| GRPOBackbone=Qwen3-VL-8B-Instruct2026.05 | 0.683 | |
| Qwen3-VL-8B-InstructRole=Base2026.05 | 0.547 | |
| GPT-5.42026.05 | 0.536 | |
| HumanLM-8B2026.05 | 0.481 | |
| GPT-5-nano2026.05 | 0.358 |