Machine Translation on Virām 1.0 (test)
26.2BLEUIndicTrans2 en indic 200 M (Upper Performance Boundary)
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| IndicTrans2 en indic 200 M (Upper Performance Boundary)Type=Upper Performance boundary, Input=Original + Input as 'sentence meant'2025.12 | 26.2 | 0.8082 | 0.7606 | 61.15 | 57.41 | 0.9313 | 0.7915 | |
| Approach 1: Fine-tuned microsoft-mpnet + IndicTrans2Restoration Model=microsoft-mpnet, Translation Model=IndicTrans2 (Original), Paradigm=Restore Punctuation then Translate2025.12 | 25.12 | 0.7996 | 0.7597 | 60.56 | 56.79 | 0.921 | 0.7813 | |
| Approach 1: Fine-tuned t5-base + IndicTrans2Restoration Model=google-t5-base, Translation Model=IndicTrans2 (Original), Paradigm=Restore Punctuation then Translate2025.12 | 24.74 | 0.7977 | 0.7586 | 60.68 | 56.92 | 0.923 | 0.7838 | |
| Approach 2: Finetuned (w/o punct)Dataset Variant=Without Punctuation (x), Paradigm=Direct Translation2025.12 | 24.66 | 0.7774 | 0.7417 | 60.3 | 56.56 | 0.9122 | 0.783 | |
| Approach 2: Finetuned (alternate with and w/o punct)Dataset Variant=Combined x (alternate with and without punct), Paradigm=Direct Translation2025.12 | 24.28 | 0.7745 | 0.7433 | 60.21 | 56.52 | 0.9047 | 0.7761 | |
| Approach 2: Finetuned (with and w/o punct)Dataset Variant=Combined 2x (with and without punct), Paradigm=Direct Translation2025.12 | 24.27 | 0.7785 | 0.7443 | 60.61 | 56.83 | 0.912 | 0.7794 | |
| Approach 1: Fine-tuned bert-large-uncased + IndicTrans2Restoration Model=bert-large-uncased, Translation Model=IndicTrans2 (Original), Paradigm=Restore Punctuation then Translate2025.12 | 23.84 | 0.7955 | 0.7595 | 60.02 | 56.11 | 0.9199 | 0.7806 | |
| Approach 1: AI4Bharat’s cadence + IndicTrans2Restoration Model=AI4Bharat’s cadence, Translation Model=IndicTrans2 (Original), Paradigm=Restore Punctuation then Translate2025.12 | 23.44 | 0.798 | 0.7516 | 60.49 | 56.69 | 0.921 | 0.7809 | |
| DeepSeek-V3.2 (non-thinking)Type=LLM, Prompting=Zero-Shot, Translation Mode=Direct Translation2025.12 | 23.41 | 0.7858 | 0.759 | 58.48 | 54.82 | 0.9197 | 0.7765 | |
| IndicTrans2 en indic 200 MType=Baseline, Variant=Original Model2025.12 | 21.72 | 0.7916 | 0.7391 | 59.45 | 55.38 | 0.9126 | 0.7619 | |
| Approach 2: Finetuned (w/ punct)Dataset Variant=With Punctuation (x), Paradigm=Direct Translation2025.12 | 21.21 | 0.783 | 0.7426 | 58.9 | 54.72 | 0.9145 | 0.7685 | |
| GPT-4o-miniType=LLM, Prompting=Zero-Shot, Translation Mode=Direct Translation2025.12 | 18.69 | 0.7786 | 0.742 | 52.5 | 48.82 | 0.9096 | 0.7394 |