Semantic Parsing on Break
42.21AccuracyCLG
Evaluation Results
| Method | Links | |
|---|---|---|
| CLGShot count=1024-shot, Model=Qwen2.5-72B2025.06 | 42.21 | |
| RandomShot count=1024-shot, Model=Qwen2.5-72B2025.06 | 42.08 | |
| EPR-KMeansModel=Qwen2.5-72B, Shots=1282025.06 | 39.12 | |
| EPR-KMeansModel=Llama3-70B, Shots=1282025.06 | 38.39 | |
| CLGModel=Llama3-70B, Shots=1282025.06 | 37.33 | |
| CLGModel=Qwen2.5-72B, Shots=1282025.06 | 37.07 | |
| CLGShot count=128-shot, Model=Qwen2.5-72B2025.06 | 37.07 | |
| RandomModel=Qwen2.5-72B, Shots=1282025.06 | 36.67 | |
| RandomShot count=128-shot, Model=Qwen2.5-72B2025.06 | 36.67 | |
| Best-of-NModel=Llama3-70B, Shots=1282025.06 | 35.86 | |
| BGE-KMeansModel=Qwen2.5-72B, Shots=1282025.06 | 35.84 | |
| Best-of-NModel=Qwen2.5-72B, Shots=1282025.06 | 35.73 | |
| RandomModel=Llama3-70B, Shots=1282025.06 | 35.55 | |
| Latent-BayesianModel=Qwen2.5-72B, Shots=1282025.06 | 35.35 | |
| BGE-KMeansModel=Llama3-70B, Shots=1282025.06 | 35.13 | |
| Latent-BayesianModel=Llama3-70B, Shots=1282025.06 | 32.59 | |
| BM25-MajorModel=Qwen2.5-72B, Shots=1282025.06 | 22.81 | |
| BM25-MajorModel=Llama3-70B, Shots=1282025.06 | 15.05 |