Collector instantiation on 138 tasks with natural-language descriptions (test)
53.69Avg. QualityOurs
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Ours2026.06 | 53.69 | 55.07 | 85.51 | |
| Unconstrained LLM generation2026.06 | 48.76 | 52.17 | 18.84 | |
| Rule-template baseline (no LLM)LLM usage=false2026.06 | 31.84 | 9.42 | 75.36 |