Financial Advisory Preference Ranking on S&P 500 2017 (held-out)
56.8Win Rate vs Qwen3-VL-235BQwen3-VL-4B + Hindsight DPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen3-VL-4B + Hindsight DPOTraining Stage=Hindsight Direct Preference Optimization, Base Model=Qwen3-VL-4B, Evaluation protocol=LLM judge, Number of evaluation runs=52026.04 | 56.8 | 52.6 | 71.4 | |
| Qwen3-VL-4B + SFT (Stage 1)Training Stage=Supervised Fine-Tuning (Stage 1), Base Model=Qwen3-VL-4B, Evaluation protocol=LLM judge, Number of evaluation runs=52026.04 | 48.7 | 46.7 | 62.5 |