Factuality Correction on BIO dataset
93Factual PrecisionLLM2
Evaluation Results
| Method | Links | |
|---|---|---|
| LLM2Base Model=llama-3.3-70b-instruct, Evaluation Stage=After Correction2026.01 | 93 | |
| RACBase Model=llama-3.3-70b-instruct, Evaluation Stage=After Correction2026.01 | 91 | |
| FACTCORRECTORBase Model=llama-3.3-70b-instruct, Evaluation Stage=After Correction2026.01 | 91 | |
| FACTCORRECTORBase Model=granite-4.0-h-small, Evaluation Stage=After Correction2026.01 | 90 | |
| FACTCORRECTORBase Model=mixtral-8x22b-instruct, Evaluation Stage=After Correction2026.01 | 87 | |
| CRITICBase Model=llama-3.3-70b-instruct, Evaluation Stage=After Correction2026.01 | 86 | |
| LLM2Base Model=mixtral-8x22b-instruct, Evaluation Stage=After Correction2026.01 | 86 | |
| RACBase Model=granite-4.0-h-small, Evaluation Stage=After Correction2026.01 | 86 | |
| RACBase Model=gpt-oss-120b, Evaluation Stage=After Correction2026.01 | 86 | |
| RACBase Model=mixtral-8x22b-instruct, Evaluation Stage=After Correction2026.01 | 84 | |
| CRITICBase Model=granite-4.0-h-small, Evaluation Stage=After Correction2026.01 | 80 | |
| LLM2Base Model=granite-4.0-h-small, Evaluation Stage=After Correction2026.01 | 79 | |
| LLM2Base Model=gpt-oss-120b, Evaluation Stage=After Correction2026.01 | 79 | |
| LLM1Base Model=llama-3.3-70b-instruct, Evaluation Stage=After Correction2026.01 | 76 | |
| CRITICBase Model=mixtral-8x22b-instruct, Evaluation Stage=After Correction2026.01 | 75 | |
| FACTCORRECTORBase Model=gpt-oss-120b, Evaluation Stage=After Correction2026.01 | 75 | |
| granite-4.0-h-smallEvaluation Stage=Before Correction2026.01 | 72 | |
| LLM1Base Model=mixtral-8x22b-instruct, Evaluation Stage=After Correction2026.01 | 72 | |
| LLM1Base Model=granite-4.0-h-small, Evaluation Stage=After Correction2026.01 | 71 | |
| llama-3.3-70b-instructEvaluation Stage=Before Correction2026.01 | 65 | |
| mixtral-8x22b-instructEvaluation Stage=Before Correction2026.01 | 64 | |
| gpt-oss-120bEvaluation Stage=Before Correction2026.01 | 56 | |
| CRITICBase Model=gpt-oss-120b, Evaluation Stage=After Correction2026.01 | 56 | |
| LLM1Base Model=gpt-oss-120b, Evaluation Stage=After Correction2026.01 | 33 |