Tokenisation on Wikipedia/OpenWebText
99.94F1 ScoreOur (VnCoreNLP + BPE)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Our (VnCoreNLP + BPE)Tokenizer=VnCoreNLP + BPE2026.03 | 99.94 | 99.94 | 75,500 | |
| WordPiece (BERT)Model=BERT2026.03 | 99.93 | 99.93 | 18,870 | |
| SentencePiece (Unigram)Algorithm=Unigram2026.03 | 99.9 | 99.9 | 16,670 | |
| BPE (GPT-2)Model=GPT-22026.03 | 99.9 | 99.9 | 90,900 | |
| BlingFire2026.03 | 99.86 | 99.86 | 76,000 | |
| VnCoreNLP2026.03 | 96.5 | 96.5 | 72,400 | |
| Vi Word Segmentation (NlpHUST)Source=NlpHUST2026.03 | 96 | 96 | 68,200 | |
| RDRsegmenter2026.03 | 95.8 | 95.8 | 70,500 | |
| Underthesea2026.03 | 95.2 | 95.2 | 65,000 |