Visual Question Answering on SlideBench-VQA TCGA
87.64Microscopy ScoreSlideChat
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| SlideChatInput=Slide2024.10 | 87.64 | 73.27 | 84.26 | 81.17 | |
| SlideChat2025.12 | 87.64 | 73.27 | 84.26 | — | |
| SlideChatFlops=133.3T, Status=Upper Bound, Uncompressed=true2026.03 | 87.64 | 73.27 | — | 81.17 | |
| PathFLIPInstruction fine-tuning setting=Qwen3-0.6B2025.12 | 86.11 | 76.28 | 89.41 | — | |
| HistoSelectInput=WSI2026.02 | 84.62 | 73.09 | 77.3 | — | |
| SlideChatInput=WSI2026.02 | 83.15 | 71.36 | 75.33 | — | |
| TC-SSAFlops=1.72T, Token Compression=~ 60×2026.03 | 81.94 | 77.14 | — | 78.34 | |
| SlideChat-7BModel Type=S, Reasoning Ability=Non-reasoning2026.01 | 81.68 | 73.21 | 72.45 | 75.36 | |
| PathReasoner-R1-7BModel Type=S, Reasoning Ability=Reasoning2026.01 | 78.52 | 73.43 | 72.39 | 74.68 | |
| PathReasoner-SFT-7BModel Type=S, Reasoning Ability=Reasoning2026.01 | 75.73 | 65.05 | 67.35 | 67.98 | |
| MedDrInput=Patch2024.10 | 73.3 | 57.78 | 74.25 | 67.7 | |
| MedDrInput sampling protocol=30 randomly sampled patches2025.12 | 73.3 | 57.78 | 74.25 | — | |
| MedDrFlops=1.70T, Token Compression=~ 60×2026.03 | 73.3 | 57.78 | — | 67.7 | |
| MedDrInput=Slide (T)2024.10 | 70.48 | 52.47 | 72.8 | 64.25 | |
| PathGen-LLavaInput sampling protocol=30 randomly sampled patches2025.12 | 68.67 | 49.69 | 63.36 | — | |
| CPath-Omni2025.12 | 63.7 | 52.41 | 59.26 | — | |
| Patho-R1-7BModel Type=T, Reasoning Ability=Reasoning2026.01 | 63.61 | 47.53 | 57.14 | 52.34 | |
| GPT-4oInput=Patch2024.10 | 62.89 | 46.69 | 66.77 | 57.91 | |
| GPT-4oToken Compression=~ 60×2026.03 | 62.89 | 46.69 | — | 57.91 | |
| HuatuoGPT-Vision-7BModel Type=T, Reasoning Ability=Non-reasoning2026.01 | 58.64 | 39.58 | 60.2 | 45.89 | |
| Qwen3-VL-8B-InstructModel Type=T, Reasoning Ability=Non-reasoning2026.01 | 57.85 | 39.38 | 69.39 | 46.16 | |
| Quilt-LLaVAInput=Patch2024.10 | 57.76 | 35.96 | 53.07 | 48.07 | |
| Quilt-LLavaInput sampling protocol=30 randomly sampled patches2025.12 | 57.76 | 35.96 | 53.07 | — | |
| Quilt-LLaVAFlops=1.70T, Token Compression=~ 60×2026.03 | 57.76 | 35.96 | — | 48.07 | |
| InternVL3.5-8BModel Type=T, Reasoning Ability=Reasoning2026.01 | 56.28 | 45.52 | 68.37 | 49.83 | |
| WSI-LLaVA-7BModel Type=S, Reasoning Ability=Non-reasoning2026.01 | 56.08 | 64.14 | 52.57 | 60.2 | |
| Quilt-LLaVAInput=Thumbnail2026.02 | 52.39 | 30.19 | 49.33 | — | |
| LLaVA-MedInput=Thumbnail2026.02 | 52.15 | 29.97 | 47.33 | — | |
| Qwen2.5-VL-7B-InstructModel Type=T, Reasoning Ability=Non-reasoning2026.01 | 49.74 | 37.16 | 53.06 | 41.48 | |
| MedGemma-4B-ITModel Type=T, Reasoning Ability=Non-reasoning2026.01 | 49.48 | 36.96 | 57.14 | 41.55 | |
| Qwen3-VL-8B-ThinkingModel Type=T, Reasoning Ability=Reasoning2026.01 | 49.48 | 33.23 | 48.98 | 38.49 | |
| Quilt-LLaVAInput=Slide (T)2024.10 | 49.12 | 26.97 | 44.75 | 39.39 | |
| MedVLThinker-7BModel Type=T, Reasoning Ability=Reasoning2026.01 | 48.43 | 44.61 | 65.31 | 46.98 | |
| LLaVA-MedInput=Patch2024.10 | 47.34 | 32.78 | 47.96 | 42 | |
| LLava-MedInput sampling protocol=30 randomly sampled patches2025.12 | 47.34 | 32.78 | 47.96 | — | |
| LLaVA-MedFlops=1.70T, Token Compression=~ 60×2026.03 | 47.34 | 32.78 | — | 42 | |
| LLaVA-MedInput=Slide (T)2024.10 | 45.82 | 27.58 | 40.84 | 37.39 | |
| Quilt-LLaVA-7BModel Type=T, Reasoning Ability=Non-reasoning2026.01 | 44.76 | 20.24 | 52.04 | 28.72 | |
| GPT-4oInput=Thumbnail2026.02 | 39.24 | 24.12 | 44.67 | — | |
| GPT-4Input=Text2024.10 | 38.28 | 29.09 | 45 | 37.25 | |
| GPT-4oInput=Slide (T)2024.10 | 38.28 | 23.1 | 43.42 | 34.07 | |
| LLaVA-Med-7BModel Type=T, Reasoning Ability=Non-reasoning2026.01 | 35.6 | 21.05 | 42.86 | 26.27 | |
| RandomInput=Text2024.10 | 24.44 | 24.91 | 26.44 | 25.02 | |
| Random BaselineToken Compression=~ 60×2026.03 | 24.44 | 24.91 | — | 25.02 |