Tactical Deconfliction on BlueSky Simulation Scenario C
1.9NMACs (All)SFT-LoRA
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| SFT-LoRAAlignment Strategy=SFT, Base Model=Qwen-Math-7B2026.03 | 1.9 | 0.8 | 1.1 | 52 | 7.5 | |
| GRPO-LoRAAlignment Strategy=GRPO, Base Model=Qwen-Math-7B2026.03 | 2.3 | 0.6 | 1.7 | 42 | 6.6 | |
| Qwen-Math-7BAlignment Strategy=Base2026.03 | 4 | 2.5 | 1.5 | 5 | 1.6 |