Semantic Parsing on SMCalFlow
88.9Program AccuracyGrammar Prompting (w. oracle grammar)
Evaluation Results
| Method | Links | |
|---|---|---|
| Grammar Prompting (w. oracle grammar)# ICL examples=16, # retrieval set=1282023.05 | 88.9 | |
| Grammar Prompting# ICL examples=16, # retrieval set=1282023.05 | 62.8 | |
| Previous Work# ICL examples=16, # retrieval set=1282023.05 | 60.7 | |
| Standard Prompting# ICL examples=16, # retrieval set=1282023.05 | 60 | |
| CLGShot count=1024-shot, Model=Qwen2.5-72B2025.06 | 52.82 | |
| RandomShot count=1024-shot, Model=Qwen2.5-72B2025.06 | 52.67 | |
| BGE-KMeansModel=Qwen2.5-72B, Shots=1282025.06 | 41.29 | |
| CLGModel=Llama3-70B, Shots=1282025.06 | 40.35 | |
| CLGModel=Qwen2.5-72B, Shots=1282025.06 | 40.16 | |
| CLGShot count=128-shot, Model=Qwen2.5-72B2025.06 | 40.16 | |
| EPR-KMeansModel=Qwen2.5-72B, Shots=1282025.06 | 39.65 | |
| Best-of-NModel=Qwen2.5-72B, Shots=1282025.06 | 39.22 | |
| BGE-KMeansModel=Llama3-70B, Shots=1282025.06 | 38.73 | |
| RandomModel=Qwen2.5-72B, Shots=1282025.06 | 38.23 | |
| RandomShot count=128-shot, Model=Qwen2.5-72B2025.06 | 38.23 | |
| EPR-KMeansModel=Llama3-70B, Shots=1282025.06 | 38.1 | |
| Best-of-NModel=Llama3-70B, Shots=1282025.06 | 37.66 | |
| RandomModel=Llama3-70B, Shots=1282025.06 | 36.68 | |
| Latent-BayesianModel=Llama3-70B, Shots=1282025.06 | 23.13 | |
| Latent-BayesianModel=Qwen2.5-72B, Shots=1282025.06 | 22.61 | |
| BM25-MajorModel=Llama3-70B, Shots=1282025.06 | 14.38 | |
| BM25-MajorModel=Qwen2.5-72B, Shots=1282025.06 | 11.57 |