Card Games on Point24
54SRGTR-Turbo
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GTR-TurboGuidance Mode=KL, Training Time=89h, Cost Estimation=$114.812025.12 | 54 | — | |
| GTR-TurboTraining strategy=KL, Base Model=Qwen2.5-VL-7B-Instruct2025.12 | 53.5 | 2.39 | |
| GTR-TurboGuidance Mode=SFT, Training Time=168h, Cost Estimation=$216.722025.12 | 48 | — | |
| GTR-TurboTraining strategy=SFT, Base Model=Qwen2.5-VL-7B-Instruct2025.12 | 48 | 1.32 | |
| GTRTraining strategy=Online imitation learning with thought guidance2025.12 | 44.5 | 0.53 | |
| GTRTeacher Model=GPT-4o, Training Time=191h, Cost Estimation=$307.78 / 70.35M tokens2025.12 | 41 | — | |
| Qwen2.5-VLModel size=7B, Training strategy=SFT2025.12 | 22 | -3.2 | |
| GPT-4oTool use=Yes2025.12 | 13.5 | -3.59 | |
| Qwen2.5-VLModel size=72B2025.12 | 5.6 | -5.69 | |
| Qwen2-VL-72BModel Parameters=72B2024.09 | 4.5 | — | |
| RL4VLMTraining Time=86h, Cost Estimation=$02025.12 | 4 | — | |
| RL4VLMTraining strategy=PPO2025.12 | 3.5 | -13.3 | |
| GPT-4o2024.09 | 3 | — | |
| Previous SOTA2024.09 | 2.6 | — | |
| GPT-4oTool use=No2025.12 | 2.5 | -6.35 | |
| CNN+RLNote=Reported in previous work2025.12 | 0 | -1.12 | |
| Qwen2.5-VLModel size=32B2025.12 | 0 | -7.25 |