Seq-to-Seq Learning on Long-sequence task orders L1-L6 (test)
4.27L1 ScoreGRID
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| GRIDModel=Flan-T5-base, Setting=GRID2025.07 | 4.27 | 5.1 | 4.04 | 5.32 | 0.18 | -7.9 | 1.84 | |
| GRIDModel=Flan-T5-large, Setting=GRID2025.07 | 2.35 | 0.75 | -1.48 | -0.64 | -0.6 | -0.9 | -0.09 | |
| Progressive PromptsModel=Flan-T5-large, Setting=PP2025.07 | -0.73 | -3.17 | -6.65 | -1.43 | -1.53 | -0.99 | -2.42 | |
| GRIDModel=Flan-T5-small, Setting=GRID2025.07 | -2.53 | -4.09 | 3.33 | 0.25 | -3.02 | -4.64 | -1.78 | |
| GRIDModel=T5-large, Setting=GRID2025.07 | -2.82 | -1.2 | -0.9 | -3.11 | -2.2 | -6.28 | -2.75 | |
| Progressive PromptsModel=T5-base, Setting=PP2025.07 | -3.02 | -7.02 | -4.8 | -4.55 | -1.58 | -4.23 | -4.2 | |
| Progressive PromptsModel=Flan-T5-small, Setting=PP2025.07 | -3.57 | -5.26 | -4.98 | -5.33 | -4.92 | -4.64 | -4.78 | |
| Progressive PromptsModel=T5-large, Setting=PP2025.07 | -4.14 | -1.19 | -5.58 | -2.66 | -2.67 | -3.81 | -3.34 | |
| GRIDModel=T5-base, Setting=GRID2025.07 | -4.93 | -6.68 | -5.78 | -11.32 | -2.21 | -13.28 | -7.37 | |
| Progressive PromptsModel=T5-small, Setting=PP2025.07 | -5.07 | -11.28 | -1.48 | -6.09 | -9.58 | -8.63 | -7.02 | |
| GRIDModel=T5-small, Setting=GRID2025.07 | -5.82 | -0.11 | -5.55 | -0.96 | -1.11 | -3.33 | -2.81 | |
| Progressive PromptsModel=Flan-T5-base, Setting=PP2025.07 | -13.91 | -2.08 | -1.48 | -4.9 | -2.6 | -4.53 | -4.92 |