Compositional Generalization on Evaluation Dataset Unseen (Fold 1)
0.4818ScoreDeepSeek V3
Evaluation Results
| Method | Links | |
|---|---|---|
| DeepSeek V3Evaluation Protocol=Zero-shot, Model Backbone=DeepSeek V32026.01 | 0.4818 | |
| GPT-4o + FS*Evaluation Protocol=Few-shot, Model Backbone=GPT-4o2026.01 | 0.4668 | |
| Claude 3.5 Sonnet + FS*Evaluation Protocol=Few-shot, Model Backbone=Claude 3.5 Sonnet2026.01 | 0.4627 | |
| LLaMA-3-8BFine-tuning Strategy=LoRA, Model Backbone=LLaMA-3-8B2026.01 | 0.4626 | |
| Gemma-7BFine-tuning Strategy=LoRA, Model Backbone=Gemma-7B2026.01 | 0.451 | |
| GPT-4oEvaluation Protocol=Zero-shot, Model Backbone=GPT-4o2026.01 | 0.4298 | |
| LLaMA-3-8B + COGLMFine-tuning Strategy=LoRA, Model Backbone=LLaMA-3-8B, Method Components=COGLM2026.01 | 0.4109 | |
| Qwen2.5-7BFine-tuning Strategy=LoRA, Model Backbone=Qwen2.5-7B2026.01 | 0.4079 | |
| Gemma-7B + COGLMFine-tuning Strategy=LoRA, Model Backbone=Gemma-7B, Method Components=COGLM2026.01 | 0.3974 | |
| Claude 3.5 SonnetEvaluation Protocol=Zero-shot, Model Backbone=Claude 3.5 Sonnet2026.01 | 0.3944 | |
| Mistral-7BFine-tuning Strategy=LoRA, Model Backbone=Mistral-7B2026.01 | 0.388 | |
| Llama 3 (70B) + FS*Evaluation Protocol=Few-shot, Model Backbone=Llama 3 (70B)2026.01 | 0.3811 | |
| Mistral-7B + COGLMFine-tuning Strategy=LoRA, Model Backbone=Mistral-7B, Method Components=COGLM2026.01 | 0.3786 | |
| Qwen2.5-7B + COGLMFine-tuning Strategy=LoRA, Model Backbone=Qwen2.5-7B, Method Components=COGLM2026.01 | 0.344 | |
| DeepSeek-7BFine-tuning Strategy=LoRA, Model Backbone=DeepSeek-7B2026.01 | 0.3082 | |
| DeepSeek V3 + FS*Evaluation Protocol=Few-shot, Model Backbone=DeepSeek V32026.01 | 0.3016 | |
| DeepSeek-7B + COGLMFine-tuning Strategy=LoRA, Model Backbone=DeepSeek-7B, Method Components=COGLM2026.01 | 0.2953 | |
| Llama 3 (70B)Evaluation Protocol=Zero-shot, Model Backbone=Llama 3 (70B)2026.01 | 0.2524 |