Grammatical Error Correction on JFLEG (test)
64.9GLEULichtarge et al. (2020)
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Lichtarge et al. (2020)2022.04 | 64.9 | — | — | — | — | |
| Lichtarge et al. (2019)Ensemble=true2019.09 | 63.3 | — | — | — | — | |
| gemma-2-9b-itDecoding=MBR, Training=Zero-shot (4-shot)2026.05 | 62.9 | 70.3 | — | 71.7 | 65.2 | |
| gemma-2-9b-itDecoding=Proposal (edit-level majority voting), Training=Zero-shot (4-shot), Threshold (τ)=22026.05 | 62.8 | 68.5 | — | 69.1 | 66.1 | |
| gemma-2-9b-itDecoding=Greedy, Training=Zero-shot (4-shot)2026.05 | 62.6 | 70.8 | — | 72.5 | 65 | |
| Base + FB learning and inferencefluency boost learning=true, fluency boost inference=true2018.07 | 62.42 | — | — | — | — | |
| Baselinesetup=simple2022.04 | 62.1 | — | — | — | — | |
| Lichtarge et al. (2019)Ensemble=false2019.09 | 61.6 | — | — | — | — | |
| Lichtarge et al.ensemble=false2019.10 | 61.6 | — | — | — | — | |
| SMT-NMT hybrid2018.07 | 61.5 | — | — | — | — | |
| NMT SMT HybridYear=2018, ensemble_size=4, extra_language_model=true, Dict=bpe2019.03 | 61.5 | — | — | — | — | |
| Base + FB learningfluency boost learning=true2018.07 | 61.41 | — | — | — | — | |
| CNN + FB LearningYear=2018, ensemble_size=4, extra_language_model=true, Dict=bpe, data_type=Large Non-public Training Data2019.03 | 61.41 | — | — | — | — | |
| PRETLARGE+SSE+R2LEnsemble=true2019.09 | 61.4 | — | — | — | — | |
| Grundkiewicz et al. (2019)Ensemble=true2019.09 | 61.2 | — | — | — | — | |
| PRETLARGE+SSE+R2L+SEDEnsemble=true2019.09 | 61.2 | — | — | — | — | |
| Copy-augmented Model + DA + Multi-tasksensemble_size=4, Dict=word, re-ranked with extra LM=true2019.03 | 61 | — | — | — | — | |
| Zhao et al. (2019)Ensemble=true2019.09 | 61 | — | — | — | — | |
| Base convolutional seq2seq2018.07 | 60.87 | — | — | — | — | |
| EPO (Llama2-7b-chat)Decoding=MBR, Training=Fine-tuned2026.05 | 60.6 | 64.4 | — | 65 | 62.2 | |
| PIEensemble=false2019.10 | 60.3 | — | — | — | — | |
| Adapted-transformer2018.07 | 59.9 | — | — | — | — | |
| Transformer + MIMsYear=2018, ensemble_size=4, extra_language_model=true, Dict=bpe2019.03 | 59.9 | — | — | — | — | |
| Junczys-Dowmunt et al. (2018)Ensemble=true2019.09 | 59.9 | — | — | — | — | |
| Qwen3-8BDecoding=Proposal (edit-level majority voting), Training=Zero-shot (4-shot), Threshold (τ)=12026.05 | 59.8 | 70.6 | — | 73.6 | 60.5 | |
| PRETLARGEEnsemble=false2019.09 | 59.7 | — | — | — | — | |
| T5 (t5-v1_1-large)Parameters=0.8B, Training=Fine-tuned2026.05 | 59.6 | 70.7 | — | 73.9 | 60 | |
| Qwen3-8BDecoding=MBR, Training=Zero-shot (4-shot)2026.05 | 59.5 | 72.5 | — | 76.2 | 60.5 | |
| EPO (Llama2-7b-chat)Decoding=Proposal (edit-level majority voting), Training=Fine-tuned, Threshold (τ)=22026.05 | 59.5 | 61.6 | — | 61 | 64.1 | |
| Copy-augmented Modelensemble_size=4, Dict=word, re-ranked with extra LM=true2019.03 | 59.48 | — | — | — | — | |
| Qwen3-8BDecoding=Greedy, Training=Zero-shot (4-shot)2026.05 | 58.9 | 72.5 | — | 76.7 | 59.4 | |
| EPO (Llama2-7b-chat)Decoding=Greedy, Training=Fine-tuned2026.05 | 58.9 | 71.1 | — | 74.3 | 60.4 | |
| GECToR (deberta-v3-large)Parameters=0.4B, Training=Fine-tuned2026.05 | 58.9 | 68 | — | 70.8 | 58.6 | |
| Llama-3.1-8B-InstructDecoding=Proposal (edit-level majority voting), Training=Zero-shot (4-shot), Threshold (τ)=32026.05 | 58.5 | 65.4 | — | 67.2 | 59.2 | |
| Grundkiewicz and Junczys-Dowmunt (2018)Ensemble=false2019.09 | 57.9 | — | — | — | — | |
| Chollampatt and Ng (2018)Ensemble=true2019.09 | 57.5 | — | — | — | — | |
| MLConvembed (4 ens.) + EO + LM + SpellCheckword embeddings=pre-trained fastText, ensemble=4, edit operation features=true, web-scale LM=true, spell check=true2018.01 | 57.47 | 66.8 | — | — | — | |
| NUS182018.07 | 57.47 | — | — | — | — | |
| CNN + EOYear=2018, ensemble_size=4, extra_language_model=true, Dict=bpe2019.03 | 57.47 | — | — | — | — | |
| C&N (2017) + SpellCheckspell check=true2018.01 | 56.78 | 64.25 | — | — | — | |
| NUS172018.07 | 56.78 | — | — | — | — | |
| Back-CNN-seq2seq2018.07 | 56.6 | — | — | — | — | |
| MLConvembed (4 ens.) + EO + LMword embeddings=pre-trained fastText, ensemble=4, edit operation features=true, web-scale LM=true2018.01 | 55.99 | 65.87 | — | — | — | |
| GECToR (bert-base-cased)Parameters=0.1B, Training=Fine-tuned2026.05 | 55.3 | 62.5 | — | 65.9 | 52 | |
| Nested-RNN-seq2seq2018.07 | 53.41 | — | — | — | — | |
| Neural Hybird MTYear=2017, extra_language_model=true, Dict=char/word2019.03 | 53.41 | — | — | — | — | |
| MLConvembed (4 ens.) + EOword embeddings=pre-trained fastText, ensemble=4, edit operation features=true2018.01 | 53.38 | 62.15 | — | — | — | |
| C&N (2017)2018.01 | 53.18 | 60.95 | — | — | — | |
| Chollampatt and Ng (2018)Ensemble=false2019.09 | 53 | — | — | — | — | |
| Junczys-Dowmunt et al. (2018)Ensemble=false2019.09 | 53 | — | — | — | — | |
| CAMB162018.07 | 52.05 | — | — | — | — | |
| AMU162018.07 | 51.46 | — | — | — | — | |
| MLConvembedword embeddings=pre-trained fastText2018.01 | 51.34 | 58.82 | — | — | — | |
| Chollampatt and Ngensemble=false2019.10 | 51.3 | — | — | — | — | |
| MLConvembed (4 ens.)word embeddings=pre-trained fastText, ensemble=42018.01 | 51.06 | 59.21 | — | — | — | |
| NUS162018.07 | 50.13 | — | — | — | — | |
| Llama-3.1-8B-InstructDecoding=MBR, Training=Zero-shot (4-shot)2026.05 | 47.2 | 62.1 | — | 62.5 | 60.7 | |
| CAMB142018.07 | 46.04 | — | — | — | — | |
| No edit2018.07 | 40.54 | — | — | — | — | |
| Llama-3.1-8B-InstructDecoding=Greedy, Training=Zero-shot (4-shot)2026.05 | 36.2 | 64.3 | — | 65.6 | 59.8 | |
| BaselinePLM=false, Syntax=false2022.10 | — | — | 58.15 | — | — | |
| BaselinePLM=true, Syntax=false2022.10 | — | — | 61.53 | — | — | |
| SynGECPLM=false, Syntax=true2022.10 | — | — | 60.14 | — | — | |
| SynGECPLM=true, Syntax=true2022.10 | — | — | 62.15 | — | — |