Natural Language Inference on OP v2 (test)
54.65 Persp Average ScoreQwen2.5-3B-Instruct (Implicit OP-GRPO)
Evaluation Results
| Method | Links | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Qwen2.5-3B-Instruct (Implicit OP-GRPO)Backbone=Qwen2.5-3B-Instruct, Training Algorithm=GRPO, Prompting Strategy=Implicit OP Prompting2026.02 | 54.6 | 88 | 49.1 | 84.7 | 47.9 | 79.7 | 45 | 78 | 43 | 73 | 42.3 | 73.3 | 45.1 | 75.5 | 46.7 | 78.9 | |
| Llama3.2-3B-Instruct (Implicit OP-GRPO)Backbone=Llama3.2-3B-Instruct, Training Algorithm=GRPO, Prompting Strategy=Implicit OP Prompting2026.02 | 53.2 | 88.3 | 47.5 | 79.6 | 47.6 | 81 | 42.8 | 71.3 | 42.9 | 71 | 41.1 | 69.3 | 44.4 | 75 | 45.6 | 76.5 | |
| Qwen2.5-1.5B-Instruct (Implicit OP-GRPO)Backbone=Qwen2.5-1.5B-Instruct, Training Algorithm=GRPO, Prompting Strategy=Implicit OP Prompting2026.02 | 52.4 | 85 | 45.8 | 78.7 | 45.4 | 80.6 | 42.4 | 72.7 | 40.1 | 68.3 | 39.3 | 62.7 | 41.6 | 66.5 | 43.9 | 73.5 | |
| Modular Pluralism (Qwen3-14B)Backbone=Mistral-7B, Summary Model=Qwen3-14B, Evaluation Protocol=Modular Pluralism2026.02 | 42.7 | 57.6 | 41.1 | 60.3 | 39.7 | 56 | 36.9 | 50.7 | 39.1 | 55 | 37.2 | 49 | 37.7 | 52.5 | 39.2 | 54.4 | |
| GPT-OSS (Explicit OP)Backbone=GPT-OSS, Prompting Strategy=Explicit OP Prompting2026.02 | 39.2 | 56.3 | 37.5 | 51 | 36 | 46.7 | 33.6 | 44 | 31.7 | 45.1 | 29.8 | 44.5 | 30.2 | 43.4 | 34 | 47.3 | |
| Qwen 3 – 8B (Explicit OP)Backbone=Qwen 3 – 8B, Prompting Strategy=Explicit OP Prompting2026.02 | 36.6 | 19 | 28.3 | 34 | 27.8 | 32.3 | 26.9 | 30.3 | 24.5 | 23.3 | 23 | 20.7 | 24.3 | 26.5 | 27.3 | 25.6 | |
| Qwen2.5-3B-Instruct (Explicit OP Prompting)Backbone=Qwen2.5-3B-Instruct, Prompting Strategy=Explicit OP Prompting2026.02 | 33.2 | 46.7 | 27.1 | 33.7 | 27.7 | 32.7 | 27.5 | 34 | 24.8 | 25.3 | 24.6 | 23.3 | 25.3 | 25 | 27.2 | 31.5 | |
| Qwen2.5-1.5B-Instruct (Explicit OP Prompting)Backbone=Qwen2.5-1.5B-Instruct, Prompting Strategy=Explicit OP Prompting2026.02 | 31.5 | 44.7 | 24.3 | 25.6 | 25.5 | 29 | 23.6 | 24 | 21.6 | 18.3 | 20.6 | 14.7 | 21.3 | 15 | 24.1 | 24.5 | |
| Qwen2.5-1.5B-Instruct (Implicit OP Prompting)Backbone=Qwen2.5-1.5B-Instruct, Prompting Strategy=Implicit OP Prompting2026.02 | 27.3 | 33 | 19.8 | 19 | 20.5 | 18.7 | 18.6 | 14.7 | 17.8 | 12.7 | 15.2 | 8 | 17 | 11 | 19.5 | 14.1 | |
| Llama3.2-3B-Instruct (Explicit OP Prompting)Backbone=Llama3.2-3B-Instruct, Prompting Strategy=Explicit OP Prompting2026.02 | 26.9 | 34 | 22.9 | 24 | 21.3 | 20 | 21.1 | 18.7 | 18.6 | 13.7 | 18 | 12.7 | 19.7 | 15 | 21.2 | 19.7 | |
| Qwen2.5-3B-Instruct (Implicit OP Prompting)Backbone=Qwen2.5-3B-Instruct, Prompting Strategy=Implicit OP Prompting2026.02 | 23.7 | 28 | 17.2 | 15 | 17.4 | 14.7 | 15 | 10.7 | 14.4 | 10.3 | 13.6 | 8.33 | 14.7 | 10 | 16.6 | 13.9 | |
| Llama3.2-3B-Instruct (Implicit OP Prompting)Backbone=Llama3.2-3B-Instruct, Prompting Strategy=Implicit OP Prompting2026.02 | 23.7 | 26.3 | 17.8 | 15.7 | 18.4 | 17.3 | 18.1 | 17.1 | 14.7 | 10.7 | 14.9 | 10.7 | 14.5 | 6.5 | 17.4 | 14.9 |