Binary classification on EXIST Task 1.1 2025 (test)
69.63ICMRouting to CEJ (L to C)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Routing to CEJ (L to C)Variant=C4, persona_model=LLAMA-3.3-70B-INSTRUCT, judge_model=COGITO-70B (reasoning)2025.12 | 69.63 | 85 | 82.33 | 83.89 | |
| Routing to CEJ (Q to C)Variant=C3, persona_model=QWEN2.5-72B-INSTRUCT, judge_model=COGITO-70B (reasoning)2025.12 | 68.31 | 84.33 | 81.91 | 83.48 | |
| LLaMA-3.1-8B-Instructsource=EXIST 2025 Leaderboard2025.12 | 67.74 | 84.05 | 81.67 | — | |
| LLAMA-3.2-3B + Task-specific optimizationsVariant=C2, optimizations=multi-dataset training, class-balanced loss, threshold calibration2025.12 | 65.95 | 83.15 | 80.89 | 82.57 | |
| XLM-RoBERTasource=EXIST 2025 Leaderboard, author=Villarreal-Haro et al.2025.12 | 62.97 | 81.65 | 79.96 | — | |
| Ensemble Approachsource=EXIST 2025 Leaderboard2025.12 | 62.49 | 81.41 | 79.91 | — | |
| DeepSeek-R1-Distill-Llama-8Bsource=EXIST 2025 Leaderboard2025.12 | 61.27 | 80.79 | 79.45 | — | |
| Dual-Transformer Fusion Networksource=EXIST 2025 Leaderboard2025.12 | 58.06 | 79.18 | 78.37 | — | |
| XLM-RoBERTasource=EXIST 2025 Leaderboard, author=Pan et al.2025.12 | 57.99 | 79.15 | 78.24 | — | |
| BERTsource=EXIST 2025 Leaderboard2025.12 | 57.27 | 78.78 | 78.02 | — | |
| LLAMA-3.2-3BVariant=C1, Adaptation=LoRA fine-tuning2025.12 | 57.09 | 79.13 | 75.96 | 79.34 | |
| COGITO-70B (reasoning)mode=zero-shot, judge_model=C, variant=Zero-shot Baselines2025.12 | 37.72 | 68.96 | 71.43 | 73.84 | |
| QWEN2.5-72B-INSTRUCTmode=zero-shot, variant=Zero-shot Baselines2025.12 | 37.65 | 68.92 | 71.12 | 73.84 | |
| COGITO-70Bmode=zero-shot, variant=Zero-shot Baselines2025.12 | 36.56 | 68.37 | 71.09 | 73.48 | |
| LLAMA-3.3-70B-INSTRUCTmode=zero-shot, persona=L, variant=Zero-shot Baselines2025.12 | 30.02 | 65.09 | 68.68 | 71.45 |