Retrieval Judgment on RAL2M covidqa, expertqa, hagrid, hotpotqa, msmarco (test)
70.7AccuracyLatent Model
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Latent ModelCategory=LLM Ensemble2026.01 | 70.7 | 13.9 | 0.73 | 0.634 | |
| Neural ModelCategory=LLM Ensemble2026.01 | 64.2 | 32.5 | 0.58 | 0.588 | |
| Majority VoteCategory=LLM Ensemble2026.01 | 60.2 | 49.2 | 0.526 | 0.611 | |
| LLM DebateCategory=LLM Inference2026.01 | 59.7 | 1.1 | 0.838 | 0.135 | |
| Weighted VoteCategory=LLM Ensemble2026.01 | 59.6 | 41.1 | 0.525 | 0.563 | |
| BGE+KeyCategory=Retrieval with Data Augmentation2026.01 | 58.3 | 34 | 0.51 | 0.494 | |
| GPT-5Category=LLM Inference2026.01 | 55.7 | 61.3 | 0.489 | 0.602 | |
| LLM AverageCategory=LLM Inference2026.01 | 55.5 | 50.6 | 0.486 | 0.552 | |
| BGE@0.85Category=Retrieval with Data Augmentation2026.01 | 53.1 | 38.3 | 0.449 | 0.432 | |
| BGE+IntentCategory=Retrieval with Data Augmentation2026.01 | 49.5 | 98.8 | 0.496 | 0.56 |