Aggregated Performance on Average 10 Tasks
100.6Average AccuracyFlashVLM
Evaluation Results
| Method | Links | |
|---|---|---|
| FlashVLMToken Budget=128 Tokens, Pruning Rate=77.8%2025.12 | 100.6 | |
| Upper Bound2025.12 | 100 | |
| FlashVLMToken Budget=64 Tokens, Pruning Rate=88.9%2025.12 | 97.9 | |
| FlashVLMToken Budget=32 Tokens, Pruning Rate=94.4%2025.12 | 92.8 | |
| ReMiTModel Family=SmolLM3-3B, Mid-Training=True, Few-shot=True2026.02 | 42.97 | |
| Vanilla NTPModel Family=SmolLM3-3B, Mid-Training=True, Few-shot=True2026.02 | 41.13 | |
| MiniPLMModel Family=SmolLM3-3B, Mid-Training=True, Few-shot=True2026.02 | 40.65 | |
| RHO-1Model Family=SmolLM3-3B, Mid-Training=True, Few-shot=True2026.02 | 39.8 | |
| ReMiTModel Family=Youtu-LLM-2B, Mid-Training=True, Few-shot=True2026.02 | 38.58 | |
| Vanilla NTPModel Family=Youtu-LLM-2B, Mid-Training=True, Few-shot=True2026.02 | 36.8 | |
| MiniPLMModel Family=Youtu-LLM-2B, Mid-Training=True, Few-shot=True2026.02 | 36.69 | |
| RHO-1Model Family=Youtu-LLM-2B, Mid-Training=True, Few-shot=True2026.02 | 32.24 | |
| Pre-TrainedModel Family=SmolLM3-3B, Mid-Training=Standard Checkpoint, Few-shot=True2026.02 | 31.18 | |
| Pre-TrainedModel Family=Youtu-LLM-2B, Mid-Training=Standard Checkpoint, Few-shot=True2026.02 | 30.34 | |
| ReMiTModel Family=OLMo-1B, Mid-Training=True, Few-shot=True2026.02 | 27.56 | |
| RHO-1Model Family=OLMo-1B, Mid-Training=True, Few-shot=True2026.02 | 23.1 | |
| MiniPLMModel Family=OLMo-1B, Mid-Training=True, Few-shot=True2026.02 | 22.41 | |
| Vanilla NTPModel Family=OLMo-1B, Mid-Training=True, Few-shot=True2026.02 | 22.35 | |
| Pre-TrainedModel Family=OLMo-1B, Mid-Training=Standard Checkpoint, Few-shot=True2026.02 | 16.44 |