General-purpose RAG evaluation on KILT, SuperGLUE, and AIS
0.132Kendall's TauARES (answer relevance)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ARES (answer relevance)Model=Fine-tuned lightweight LM judges + PPI, Evaluation=vs. RAGAS2026.06 | 0.132 | — | |
| ARES (context relevance)Model=Fine-tuned lightweight LM judges + PPI, Evaluation=vs. RAGAS2026.06 | 0.065 | — |