Scientific Research Title Generation on CSPubSum 10 selected examples (test)
51.18ROUGE-1Author vs PEGASUS-large
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Author vs PEGASUS-largeFine-tuned on=CSPubSum training set, Evaluation Type=Model vs Human2026.06 | 51.18 | 34.11 | 50.38 | 49.14 | 90.85 | |
| Expert-3 vs PEGASUS-largeFine-tuned on=CSPubSum training set, Evaluation Type=Model vs Human2026.06 | 48.1 | 29.11 | 43.99 | 38.08 | 90.25 | |
| Author vs LLaMA-3-8BFine-tuned on=CSPubSum training set, Evaluation Type=Model vs Human2026.06 | 46.23 | 30.07 | 45.43 | 43.71 | 90.11 | |
| Author vs Expert-3Evaluation Type=Human vs Human2026.06 | 41.48 | 20.71 | 37.81 | 45.9 | 89.14 | |
| Expert-1 vs PEGASUS-largeFine-tuned on=CSPubSum training set, Evaluation Type=Model vs Human2026.06 | 41.17 | 22.49 | 38.11 | 37.11 | 90.42 | |
| Expert-1 vs LLaMA-3-8BFine-tuned on=CSPubSum training set, Evaluation Type=Model vs Human2026.06 | 41.15 | 19.51 | 38.31 | 36.34 | 91 | |
| Expert-2 vs Expert-3Evaluation Type=Human vs Human2026.06 | 39.51 | 13.73 | 31.36 | 37.84 | 88.56 | |
| Author vs Expert-1Evaluation Type=Human vs Human2026.06 | 39.32 | 20.75 | 36.04 | 39.76 | 89.14 | |
| Expert-1 vs Expert-3Evaluation Type=Human vs Human2026.06 | 39 | 15.69 | 34.34 | 38.41 | 89.15 | |
| Expert-3 vs LLaMA-3-8BFine-tuned on=CSPubSum training set, Evaluation Type=Model vs Human2026.06 | 36.67 | 18.97 | 34.22 | 29.6 | 89.35 | |
| Expert-2 vs PEGASUS-largeFine-tuned on=CSPubSum training set, Evaluation Type=Model vs Human2026.06 | 36.59 | 13.8 | 32.77 | 29.11 | 88.56 | |
| Expert-1 vs Expert-2Evaluation Type=Human vs Human2026.06 | 34.97 | 14.39 | 31.46 | 26.73 | 87.71 | |
| Author vs Expert-2Evaluation Type=Human vs Human2026.06 | 27.18 | 10.18 | 24.32 | 27.98 | 88.3 | |
| Expert-2 vs LLaMA-3-8BFine-tuned on=CSPubSum training set, Evaluation Type=Model vs Human2026.06 | 25.38 | 5.24 | 22.27 | 18.11 | 87.74 |