RDF-to-text generation on WebNLG OOD standard (test)
37.72BLEURule-based NLG (trained by Qwen 3 235B)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Rule-based NLG (trained by Qwen 3 235B)Inter-pretability=true, Training source=Qwen 3 235B2025.12 | 37.72 | 69.8 | 92.81 | 0.1645 | |
| Rule-based NLG (trained by GPT-4.1)Inter-pretability=true, Training source=GPT-4.12025.12 | 36.15 | 71.24 | 92.51 | 0.1483 | |
| Prompted Llama 3.3 70BInter-pretability=false, training=Prompted2025.12 | 33.27 | 69.89 | 92.43 | 0.0969 | |
| Fine-tuned BARTInter-pretability=false, training=Fine-tuned2025.12 | 30.52 | 63.43 | 91.83 | -0.0261 | |
| Rule-based NLG (trained by Llama 3.3 70B)Inter-pretability=true, Training source=Llama 3.3 70B2025.12 | 28.58 | 66.06 | 91.87 | 0.0618 | |
| Rule-based NLG (trained by Qwen 2.5 72B)Inter-pretability=true, Training source=Qwen 2.5 72B2025.12 | 26.09 | 64.56 | 91.75 | 0.0655 |