Whole-slide Image Visual Question Answering on SlideBench TCGA
75.36AccuracySlideChat-7B (Baseline)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| SlideChat-7B (Baseline)Input Type=Slide2026.01 | 75.36 | — | — | — | |
| PathReasoner-R1-7BInput Type=Slide2026.01 | 74.68 | — | — | — | |
| PathReasoner-SFT-7BInput Type=Slide2026.01 | 67.98 | — | — | — | |
| WSI-LLaVA-7BInput Type=Slide2026.01 | 60.2 | — | — | — | |
| Patho-R1-7BInput Type=Thumbnail2026.01 | 52.34 | — | — | — | |
| InternVL3.5-8BInput Type=Thumbnail2026.01 | 49.82 | — | — | — | |
| MedVLThinker-7BInput Type=Thumbnail2026.01 | 46.98 | — | — | — | |
| Qwen3-VL-8B-InstructInput Type=Thumbnail2026.01 | 46.16 | — | — | — | |
| HuatuoGPT-Vision-7BInput Type=Thumbnail2026.01 | 45.89 | — | — | — | |
| MedGemma-4B-ITInput Type=Thumbnail2026.01 | 41.55 | — | — | — | |
| Qwen2.5-VL-8B-InstructInput Type=Thumbnail2026.01 | 41.48 | — | — | — | |
| Qwen3-VL-8B-ThinkingInput Type=Thumbnail2026.01 | 38.49 | — | — | — | |
| Quilt-LLaVA-7BInput Type=Thumbnail2026.01 | 28.72 | — | — | — | |
| LLaVA-Med-7BInput Type=Thumbnail2026.01 | 26.27 | — | — | — | |
| GPT-4oToken pruning ratio=~ 60×2026.06 | — | 62.89 | 46.69 | 57.91 | |
| LLaVA-MedFLOPs=1.70T, Token pruning ratio=~ 60×2026.06 | — | 47.34 | 32.78 | 42 | |
| MedDrFLOPs=1.70T, Token pruning ratio=~ 60×2026.06 | — | 73.3 | 57.78 | 67.7 | |
| Ours (SparseLearn)FLOPs=1.73T, Token pruning ratio=~ 60×2026.06 | — | 81.68 | 70.09 | 73.32 | |
| Quilt-LLaVAFLOPs=1.70T, Token pruning ratio=~ 60×2026.06 | — | 57.76 | 35.96 | 48.07 | |
| Random BaselineToken pruning ratio=~ 60×2026.06 | — | 24.44 | 24.91 | 25.02 | |
| SlideChatFLOPs=133.3T, Constraint=Upper Bound2026.06 | — | 87.64 | 73.27 | 81.17 |