Figurative Language Detection on IRFL
46.4Idioms AccuracyGRU
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GRUInput Modalities=Image vs Caption ⊕ Definition2026.06 | 46.4 | 57.4 | 44.92 | 49.56 | |
| gMLPInput Modalities=Image vs Caption ⊕ Definition2026.06 | 43.8 | 64.55 | 44.92 | 51.36 | |
| RePercENTInput Modalities=Image vs Caption ⊕ Definition2026.06 | 42.5 | 65.49 | 44.98 | 51.38 | |
| CLIPBackbone=ViT-B/32, Protocol=finetune-all, Input Modalities=Image vs Caption ⊕ Definition2026.06 | 38.1 | 70.4 | 42.16 | 50.81 | |
| GRU2026.06 | 36.1 | 48.81 | 33.69 | 39.46 | |
| CLIPBackbone=ViT-B/32, Protocol=finetune-proj, Input Modalities=Image vs Caption ⊕ Definition2026.06 | 36 | 66.5 | 41.44 | 48.67 | |
| gMLP2026.06 | 33.7 | 54.37 | 28.23 | 38.52 | |
| RePercENTSE (Semantic Encodings)=true, GSA (Group Slot Attention)=true2026.06 | 33.6 | 61.88 | 35.56 | 44.07 | |
| RePercENTSE (Semantic Encodings)=true, GSA (Group Slot Attention)=false2026.06 | 32.6 | 61.88 | 31.95 | 42.35 | |
| RePercENTSE (Semantic Encodings)=false, GSA (Group Slot Attention)=true2026.06 | 30.2 | 63.11 | 30.45 | 41.56 | |
| RePercENTSE (Semantic Encodings)=false, GSA (Group Slot Attention)=false2026.06 | 28.2 | 63.97 | 27.87 | 40.3 | |
| CLIPBackbone=ViT-B/32, Protocol=zero-shot, Input Modalities=Image vs Caption ⊕ Definition2026.06 | 26 | 48.38 | 36.04 | 38.78 | |
| CLIP-ViT-B/32Evaluation Protocol=end-to-end FT2026.06 | 22.5 | 70.18 | 29.61 | 41.73 | |
| CLIP-ViT-B/32Evaluation Protocol=projection-only FT2026.06 | 20.4 | 68.38 | 31.59 | 41.41 | |
| CLIP-ViT-B/32Evaluation Protocol=zero-shot2026.06 | 16 | 45.49 | 23.42 | 29.14 |