Compositional Generalization on Evaluation Dataset Unseen (Fold 3)
0.4022ScoreQwen2.5-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-7BFine-tuning Strategy=LoRA, Model Backbone=Qwen2.5-7B2026.01 | 0.4022 | |
| Mistral-7BFine-tuning Strategy=LoRA, Model Backbone=Mistral-7B2026.01 | 0.3977 | |
| Claude 3.5 Sonnet + FS*Evaluation Protocol=Few-shot, Model Backbone=Claude 3.5 Sonnet2026.01 | 0.3948 | |
| Qwen2.5-7B + COGLMFine-tuning Strategy=LoRA, Model Backbone=Qwen2.5-7B, Method Components=COGLM2026.01 | 0.3902 | |
| Mistral-7B + COGLMFine-tuning Strategy=LoRA, Model Backbone=Mistral-7B, Method Components=COGLM2026.01 | 0.3864 | |
| Claude 3.5 SonnetEvaluation Protocol=Zero-shot, Model Backbone=Claude 3.5 Sonnet2026.01 | 0.3822 | |
| LLaMA-3-8B + COGLMFine-tuning Strategy=LoRA, Model Backbone=LLaMA-3-8B, Method Components=COGLM2026.01 | 0.3665 | |
| Gemma-7BFine-tuning Strategy=LoRA, Model Backbone=Gemma-7B2026.01 | 0.358 | |
| Gemma-7B + COGLMFine-tuning Strategy=LoRA, Model Backbone=Gemma-7B, Method Components=COGLM2026.01 | 0.3497 | |
| DeepSeek V3Evaluation Protocol=Zero-shot, Model Backbone=DeepSeek V32026.01 | 0.3484 | |
| LLaMA-3-8BFine-tuning Strategy=LoRA, Model Backbone=LLaMA-3-8B2026.01 | 0.3481 | |
| DeepSeek V3 + FS*Evaluation Protocol=Few-shot, Model Backbone=DeepSeek V32026.01 | 0.3443 | |
| GPT-4o + FS*Evaluation Protocol=Few-shot, Model Backbone=GPT-4o2026.01 | 0.3346 | |
| GPT-4oEvaluation Protocol=Zero-shot, Model Backbone=GPT-4o2026.01 | 0.3235 | |
| DeepSeek-7BFine-tuning Strategy=LoRA, Model Backbone=DeepSeek-7B2026.01 | 0.3065 | |
| DeepSeek-7B + COGLMFine-tuning Strategy=LoRA, Model Backbone=DeepSeek-7B, Method Components=COGLM2026.01 | 0.3053 | |
| Llama 3 (70B) + FS*Evaluation Protocol=Few-shot, Model Backbone=Llama 3 (70B)2026.01 | 0.2665 | |
| Llama 3 (70B)Evaluation Protocol=Zero-shot, Model Backbone=Llama 3 (70B)2026.01 | 0.1843 |