Health-domain instruction following on HealthBench 1K-example (eval)
78.8ScoreAlternating RL
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Alternating RLActor Model=Qwen3-14B, RL Algorithm=Alternating RL (ARL)2026.03 | 78.8 | 180 | |
| Scalarized RLActor Model=Qwen3-14B, RL Algorithm=Scalarized aggregation (SRL)2026.03 | 76.2 | 340 | |
| Alternating RLActor Model=Qwen3-8B, RL Algorithm=Alternating RL (ARL)2026.03 | 76.1 | 140 | |
| Scalarized RLActor Model=Qwen3-8B, RL Algorithm=Scalarized aggregation (SRL)2026.03 | 75 | 290 | |
| Alternating RLActor Model=Qwen3-4B, RL Algorithm=Alternating RL (ARL)2026.03 | 71.6 | 120 | |
| Scalarized RLActor Model=Qwen3-4B, RL Algorithm=Scalarized aggregation (SRL)2026.03 | 69 | 260 | |
| Alternating RLActor Model=Qwen3-1.7B, RL Algorithm=Alternating RL (ARL)2026.03 | 57.6 | 90 | |
| Scalarized RLActor Model=Qwen3-1.7B, RL Algorithm=Scalarized aggregation (SRL)2026.03 | 55.6 | 220 |