Scientific-source retrieval on CheckThat! Task 1 (German) 2026 (dev)
0.6601MRR@5With LLM judge
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| With LLM judgeStage=Final selective disagreement pipeline, LLM Judge Model=gpt-5.52026.05 | 0.6601 | 0.0618 | |
| Plain rerankerStage=Reranking baseline2026.05 | 0.5983 | — |