User Simulation on HUMANUAL (test)
40.58News ScoreHUMANLM
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| HUMANLMTraining Algorithm=GRPO, Group size=4, Batch size=32, LLM-judge=gpt-5-mini, Sampling temperature=0.4, no-repeat n-gram constraint=n = 4, max response length=10242026.02 | 40.58 | 57.1 | 46.21 | 40.68 | 46.21 | 43.63 | 45.7 | |
| GRPO-thinkTraining Algorithm=GRPO, Sampling temperature=0.4, no-repeat n-gram constraint=n = 4, max response length=10242026.02 | 38.07 | 55.48 | 46.33 | 40.06 | 39.7 | 42.3 | 43.7 | |
| Qwen3-8b-thinkBackbone=Qwen3-8b, Sampling temperature=0.4, no-repeat n-gram constraint=n = 4, max response length=10242026.02 | 36.33 | 55.35 | 44.5 | 39.78 | 38.17 | 40.7 | 42.5 |