Factuality Generation on FActScore (test)
20.4Number of FactsLlama3-8B-chat + E/R + Coarse
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Llama3-8B-chat + E/R + CoarseBackbone=Llama3-8B-chat, Evaluator=FENCE, Recipe=Edit/Remove + Coarse2024.10 | 20.4 | 10.79 | 65.41 | |
| Llama3-8B-chat + EVER-PrefBackbone=Llama3-8B-chat, Evaluator=Zero-shot, Recipe=EVER-Pref2024.10 | 20.25 | 15.16 | 57.18 | |
| Llama3-8B-chat + SFTBackbone=Llama3-8B-chat, Evaluator=Zero-shot, Recipe=SFT2024.10 | 20.05 | 18.13 | 52.52 | |
| Llama3-8B-chat + FactTune-FSBackbone=Llama3-8B-chat, Evaluator=Zero-shot, Recipe=FactTune-FS2024.10 | 18.77 | 13.34 | 58.45 | |
| Llama3-8B-chat + Self-Eval-SKTBackbone=Llama3-8B-chat, Evaluator=Zero-shot, Recipe=Self-Eval-SKT2024.10 | 18.69 | 14.22 | 56.8 | |
| Llama3-8B-chatBackbone=Llama3-8B-chat, Evaluator=Zero-shot2024.10 | 17.83 | 17.16 | 50.96 | |
| Llama2-7B-chat + EVER-PrefBackbone=Llama2-7B-chat, Evaluator=Zero-shot, Recipe=EVER-Pref2024.10 | 11.24 | 15.11 | 42.66 | |
| Llama2-7B-chat + FactTune-FSBackbone=Llama2-7B-chat, Evaluator=Zero-shot, Recipe=FactTune-FS2024.10 | 11.23 | 12.87 | 46.6 | |
| Llama2-7B-chat + Self-Eval-SKTBackbone=Llama2-7B-chat, Evaluator=Zero-shot, Recipe=Self-Eval-SKT2024.10 | 11.02 | 14.18 | 43.73 | |
| Llama2-7B-chat + E/R + CoarseBackbone=Llama2-7B-chat, Evaluator=FENCE, Recipe=Edit/Remove + Coarse2024.10 | 10.84 | 8.72 | 55.43 | |
| Llama2-7B-chat + SFTBackbone=Llama2-7B-chat, Evaluator=Zero-shot, Recipe=SFT2024.10 | 10.76 | 15.59 | 40.83 | |
| Llama2-7B-chatBackbone=Llama2-7B-chat, Evaluator=Zero-shot2024.10 | 10.7 | 17.04 | 38.57 |