Multi-Label Verdict Prediction on PUBHEALTH supplementary experiments
2.786OILlama3-8B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Llama3-8BEvaluation Scenario=Correction Performance, Knowledge-Evidence Category=Known-Unknows2026.05 | 2.786 | — | 73.59 | |
| Phi-4Evaluation Scenario=Persistence Performance, Knowledge-Evidence Category=Known-Knows2026.05 | 2.389 | 29.5 | — | |
| Llama3-8BEvaluation Scenario=Persistence Performance, Knowledge-Evidence Category=Known-Knows2026.05 | 2.256 | 30.71 | — | |
| Phi-4Evaluation Scenario=Correction Performance, Knowledge-Evidence Category=Known-Unknows2026.05 | 1.534 | — | 60.53 | |
| Llama3-8BEvaluation Scenario=Persistence Performance, Knowledge-Evidence Category=Unknown-Knows2026.05 | 1.524 | 39.62 | — | |
| Llama3-8BEvaluation Scenario=Correction Performance, Knowledge-Evidence Category=Unknown-Unknows2026.05 | 1.242 | — | 55.39 | |
| Phi-4Evaluation Scenario=Correction Performance, Knowledge-Evidence Category=Unknown-Unknows2026.05 | 1.237 | — | 55.29 | |
| Phi-4Evaluation Scenario=Persistence Performance, Knowledge-Evidence Category=Unknown-Knows2026.05 | 0.832 | 54.6 | — |