Faithfulness Evaluation on tldr_news (800 samples)
79.5BLEURandom
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| RandomModel=Llama-2 (7B-Chat)2024.05 | 79.5 | 73.8 | 73.9 | 73.8 | 0.92 | 94.3 | 0.04 | |
| AttentionModel=Llama-2 (7B-Chat)2024.05 | 76.1 | 69.5 | 69.8 | 69.5 | 0.9 | 90.3 | 0.074 | |
| Last-AttentionModel=Llama-2 (7B-Chat)2024.05 | 74.9 | 67.8 | 68.2 | 67.9 | 0.879 | 86.7 | 0.11 | |
| Integrated-GradientModel=Llama-2 (7B-Chat)2024.05 | 69.3 | 60.5 | 61.1 | 60.7 | 0.849 | 78.3 | 0.175 | |
| JoPAModel=Llama-2 (7B-Chat)2024.05 | 68.7 | 59.5 | 59.1 | 59 | 0.832 | 58.9 | 0.419 |