Summarization on SAMSum (ROUGE, BERTScore, Aligned Pref., Faithfulness)
50.78ROUGE-1qwen3
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| qwen3Inference mode=slow (thinking), Reward configuration=BRL2026.04 | 50.78 | 25.03 | 47.58 | 74.67 | 48.83 | 68.52 | |
| qwen3Inference mode=slow (thinking), Reward configuration=BRLP2026.04 | 49.78 | 24.12 | 46.43 | 73.32 | 66.83 | 83.52 | |
| qwen3Inference mode=fast (no-thinking), Reward configuration=BRL2026.04 | 48.85 | 22.89 | 45.86 | 72.64 | 44.35 | 66.59 | |
| gpt-4o2026.04 | 48.52 | 19.59 | 36.55 | 79.51 | 52.23 | 63.54 | |
| qwen3Inference mode=fast (no-thinking), Reward configuration=BRLP2026.04 | 47.4 | 21.79 | 44.73 | 70.34 | 55.63 | 79.59 | |
| gpt-4o-mini2026.04 | 45.87 | 16.93 | 33.82 | 78.28 | 54.34 | 61.32 | |
| gpt-4.12026.04 | 42.91 | 14.08 | 32.04 | 77.37 | 50.12 | 61.24 | |
| gpt-4.1-mini2026.04 | 42.78 | 14.11 | 32.31 | 77.4 | 49.83 | 60.21 | |
| gpt-5-mini2026.04 | 41.54 | 12.95 | 30.15 | 76.68 | 48.32 | 53.44 | |
| gpt-52026.04 | 39.72 | 11.47 | 28.98 | 75.63 | 56.35 | 61.44 |