RDF-to-text generation on WebNLG All standard (test)
0.4352BLEUFine-tuned BART
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Fine-tuned BARTInter-pretability=false, training=Fine-tuned2025.12 | 0.4352 | 0.6791 | 0.9308 | 0.1275 | |
| Rule-based NLG (trained by Qwen 3 235B)Inter-pretability=true, Training source=Qwen 3 235B2025.12 | 0.3939 | 0.6759 | 0.929 | 0.1767 | |
| Rule-based NLG (trained by GPT-4.1)Inter-pretability=true, Training source=GPT-4.12025.12 | 0.3934 | 0.7069 | 0.9291 | 0.1841 | |
| Prompted Llama 3.3 70BInter-pretability=false, training=Prompted2025.12 | 0.3616 | 0.6887 | 0.9255 | 0.1058 | |
| Rule-based NLG (trained by Qwen 2.5 72B)Inter-pretability=true, Training source=Qwen 2.5 72B2025.12 | 0.3309 | 0.6531 | 0.9224 | 0.1193 | |
| Rule-based NLG (trained by Llama 3.3 70B)Inter-pretability=true, Training source=Llama 3.3 70B2025.12 | 0.2858 | 0.6578 | 0.9179 | 0.0762 |