Grammatical Error Correction on CWEB-G (test)
56.1PrecisionGECToR (deberta-v3-large)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GECToR (deberta-v3-large)Parameters=0.4B, Training=Fine-tuned2026.05 | 56.1 | 28.3 | 46.9 | |
| GECToR (bert-base-cased)Parameters=0.1B, Training=Fine-tuned2026.05 | 45.6 | 28.9 | 40.8 | |
| T5 (t5-v1_1-large)Parameters=0.8B, Training=Fine-tuned2026.05 | 45 | 47.4 | 45.4 | |
| EPO (Llama2-7b-chat)Decoding=Proposal (edit-level majority voting), Training=Fine-tuned, Threshold (τ)=52026.05 | 44.7 | 38.6 | 43.4 | |
| EPO (Llama2-7b-chat)Decoding=Greedy, Training=Fine-tuned2026.05 | 42.8 | 47 | 43.6 | |
| Qwen3-8BDecoding=Proposal (edit-level majority voting), Training=Zero-shot (4-shot), Threshold (τ)=82026.05 | 42 | 46.8 | 42.9 | |
| gemma-2-9b-itDecoding=Proposal (edit-level majority voting), Training=Zero-shot (4-shot), Threshold (τ)=82026.05 | 40.8 | 52.1 | 42.7 | |
| Llama-3.1-8B-InstructDecoding=Proposal (edit-level majority voting), Training=Zero-shot (4-shot), Threshold (τ)=82026.05 | 37.9 | 33.1 | 36.9 | |
| EPO (Llama2-7b-chat)Decoding=MBR, Training=Fine-tuned2026.05 | 37.1 | 46.3 | 38.6 | |
| Qwen3-8BDecoding=MBR, Training=Zero-shot (4-shot)2026.05 | 36.7 | 53.1 | 39.1 | |
| Qwen3-8BDecoding=Greedy, Training=Zero-shot (4-shot)2026.05 | 36 | 53.5 | 38.5 | |
| gemma-2-9b-itDecoding=MBR, Training=Zero-shot (4-shot)2026.05 | 29.5 | 63.6 | 33.1 | |
| gemma-2-9b-itDecoding=Greedy, Training=Zero-shot (4-shot)2026.05 | 29.4 | 63.9 | 33 | |
| Llama-3.1-8B-InstructDecoding=Greedy, Training=Zero-shot (4-shot)2026.05 | 19.8 | 61.5 | 22.9 | |
| Llama-3.1-8B-InstructDecoding=MBR, Training=Zero-shot (4-shot)2026.05 | 19.7 | 60.7 | 22.8 |