Rule-Inconsistency Task Classification on RIT (Rule-Inconsistency Tasks)
80.49WACLlama-8b
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Llama-8bEvaluation Protocol=Zero-Shot2026.01 | 80.49 | 51.11 | 50 | 16.67 | 75 | 0 | 64.83 | |
| Gemini-1.5-ProEvaluation Protocol=Zero-Shot2026.01 | 75.61 | 48.89 | 50 | 100 | 25 | 50 | 65.52 | |
| Deepseek-r1-7bEvaluation Protocol=One-Shot2026.01 | 75.61 | 26.67 | 0 | 16.67 | 25 | 16.67 | 53.1 | |
| Llama-70bEvaluation Protocol=Zero-Shot2026.01 | 67.07 | 31.11 | 0 | 33.33 | 50 | 0 | 50.34 | |
| Deepseek-r1-7bEvaluation Protocol=Two-Shot2026.01 | 64.63 | 6.67 | 50 | 0 | 0 | 0 | 39.31 | |
| Deepseek-r1-7bEvaluation Protocol=Zero-Shot2026.01 | 56.1 | 53.33 | 50 | 33.33 | 25 | 16.67 | 51.72 | |
| Llama-8bEvaluation Protocol=Two-Shot2026.01 | 54.88 | 15.56 | 100 | 33.33 | 25 | 0 | 39.31 | |
| Llama-8bEvaluation Protocol=One-Shot2026.01 | 53.66 | 15.56 | 0 | 0 | 50 | 0 | 36.55 | |
| Llama-70bEvaluation Protocol=One-Shot2026.01 | 52.44 | 2.22 | 0 | 16.67 | 25 | 0 | 31.72 | |
| Llama-70bEvaluation Protocol=Two-Shot2026.01 | 51.22 | 2.22 | 50 | 16.67 | 25 | 0 | 31.72 | |
| Llama-70bprompting_strategy=Zero-Shot, response_constraint=Single Response2026.01 | 51.22 | 0 | 0 | 50 | 50 | 0 | 32.41 | |
| Gemini-2.5-Proprompting_strategy=Zero-Shot, response_constraint=Single Response2026.01 | 51.22 | 51.11 | 50 | 100 | 25 | 50 | 52.41 | |
| Llama-70bprompting_strategy=One-Shot, response_constraint=Single Response2026.01 | 50 | 0 | 50 | 33.33 | 0 | 0 | 30.34 | |
| Gemini-1.5-ProEvaluation Protocol=Two-Shot2026.01 | 47.56 | 35.56 | 0 | 100 | 25 | 50 | 44.83 | |
| Gemini-2.5-Proprompting_strategy=One-Shot, response_constraint=Single Response2026.01 | 47.56 | 51.11 | 100 | 100 | 25 | 50 | 51.03 | |
| Gemini-2.5-Proprompting_strategy=Two-Shot, response_constraint=Single Response2026.01 | 46.34 | 53.33 | 0 | 100 | 25 | 50 | 49.66 | |
| Gemini-1.5-ProEvaluation Protocol=One-Shot2026.01 | 43.9 | 53.33 | 50 | 100 | 25 | 66.67 | 49.66 | |
| GPT-4oEvaluation Protocol=Zero-Shot2026.01 | 32.93 | 60 | 100 | 100 | 0 | 66.67 | 45.52 | |
| GPT-4oEvaluation Protocol=Two-Shot2026.01 | 29.27 | 40 | 50 | 100 | 75 | 66.67 | 60.16 | |
| GPT-4oprompting_strategy=Zero-Shot, response_constraint=Single Response2026.01 | 29.27 | 53.33 | 100 | 100 | 25 | 50 | 59.6 | |
| Llama-70bprompting_strategy=Two-Shot, response_constraint=Single Response2026.01 | 28.05 | 13.33 | 100 | 50 | 0 | 0 | 23.45 | |
| GPT-4oEvaluation Protocol=One-Shot2026.01 | 25.61 | 33.33 | 50 | 100 | 50 | 66.67 | 54.27 | |
| GPT-4oprompting_strategy=One-Shot, response_constraint=Single Response2026.01 | 25.61 | 31.11 | 50 | 100 | 100 | 66.67 | 62.23 | |
| Llama-8bprompting_strategy=Zero-Shot, response_constraint=Single Response2026.01 | 18.29 | 0 | 0 | 16.67 | 75 | 0 | 13.1 | |
| GPT-4oprompting_strategy=Two-Shot, response_constraint=Single Response2026.01 | 13.41 | 28.89 | 0 | 83.33 | 75 | 100 | 50.11 | |
| Llama-8bprompting_strategy=Two-Shot, response_constraint=Single Response2026.01 | 10.98 | 0 | 50 | 0 | 25 | 0 | 7.59 | |
| Llama-8bprompting_strategy=One-Shot, response_constraint=Single Response2026.01 | 4.88 | 2.22 | 0 | 16.67 | 25 | 0 | 4.83 | |
| Deepseek-r1-7bprompting_strategy=One-Shot, response_constraint=Single Response2026.01 | 4.88 | 0 | 0 | 66.67 | 0 | 0 | 5.52 | |
| Deepseek-r1-7bprompting_strategy=Two-Shot, response_constraint=Single Response2026.01 | 4.88 | 2.22 | 0 | 33.33 | 0 | 0 | 4.83 | |
| Deepseek-r1-7bprompting_strategy=Zero-Shot, response_constraint=Single Response2026.01 | 3.66 | 6.67 | 0 | 66.67 | 0 | 0 | 6.9 |