Token Compression on PangolinBench (Overall)
0.485Tok/Char RatioPangolinTokenizer
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PangolinTokenizerVocabulary size=114,8222026.07 | 0.485 | 1.795 | |
| TAIDEVocabulary size=318,0802026.07 | 0.486 | 0.647 | |
| Gemma 3/4Vocabulary size=262,1442026.07 | 0.526 | 0.725 | |
| Qwen 3.6Vocabulary size=248,0702026.07 | 0.541 | 0.745 | |
| LLaMA 3.1Vocabulary size=128,2562026.07 | 0.554 | 1.408 | |
| GPT-4oVocabulary size=200,0192026.07 | 0.558 | 0.896 | |
| Qwen 2.5/3Vocabulary size=151,6692026.07 | 0.559 | 1.179 | |
| VoxCPM2Vocabulary size=122,7532026.07 | 0.61 | 1.335 |