TITLE_HTML = "

πŸ† Multilingual Tokenizer Leaderboard

" LEADERBOARD_DESCRIPTION = """# How do tokenizers perform in your language? Welcome! This leaderboard ranks Hugging Face tokenizers based on their efficiency for multilingual text processing across various languages and text domains, including programming languages. Find out which tokenizers perform best in terms of compactness and vocabulary size. It was inspired by [MohamedRashad/arabic-tokenizers-leaderboard](https://huggingface.co/spaces/MohamedRashad/arabic-tokenizers-leaderboard) and [xu-song/tokenizer-arena](https://huggingface.co/spaces/xu-song/tokenizer-arena). ## Key Metrics Explained We evaluate tokenizers using the following core metrics: * **πŸͺΊ Fertility Score (Tokens/Word)**: The average number of tokens generated per word. **Lower scores are better**, indicating the tokenizer uses fewer tokens to represent words. The number of words is calculated by the libaries [NLTK (Punkt/Treebank)](https://www.nltk.org/) or [spaCy](https://spacy.io/) with the respective models for each natural language, for programming languages we use their Lexer from the [pygments](https://pygments.org/) libary. * **πŸ—ƒ Compression (Bytes/Token)**: The average number of UTF-8 bytes represented by a single token. **Higher scores are better**, indicating more compact representation. * **βš–οΈ Parity (vs. Reference Language)**: The ratio of tokens generated for a given language compared to a reference language (typically English) for the same text. **Closer to 1.0 is better.** A score of 1.0 means the tokenizer is equally efficient for both languages. A score **above 1.0** means the tokenizer uses more tokens for the target language than for English β€” penalising non-English speakers with higher inference costs. A score **below 1.0** means the tokenizer is actually more efficient for that language than for English. In the multilingual view this is averaged across all selected languages, so a value close to 1.0 indicates a well-balanced tokenizer with no strong language bias. (Introduced by Petrov et al., 2023, [https://arxiv.org/abs/2305.15425](https://arxiv.org/abs/2305.15425)). To calculate this metric we use the parallel corpus [FLORES+](https://huggingface.co/datasets/openlanguagedata/flores_plus). * **πŸ”„ Reversible (% docs)**: The percentage of documents that can be perfectly reconstructed (losslessly) after being tokenized and then decoded back to text (including spaces, tabs and newlines). **Higher scores are better**, indicating the tokenizer preserves the original information more accurately. * **❓ Unknown (% tokens)**: The percentage of tokens that are identified as unknown by the tokenizer. **Lower scores are better**, as they indicate the tokenizer has a better vocabulary coverage for the given text. \* Tokenizers with a high vocabulary size are more likely to have better metrics, but with a higher computational cost for the embedding layer in Tranformers Models. The tables also show the: * **βž• Total Number of Tokens**: The total number of tokens generated by the tokenizer in the selected datasets. * **πŸ“˜ Vocab Size**: The total number of unique tokens in the tokenizer's vocabulary. * **πŸ”‘ Tokenizer Hash (First 8 chars)**: A unique identifier for the specific tokenizer evaluated, is calculated from the tokenizer's vocabulary and some underlying configuration. Helps identify duplicated tokenizers submitted with different model names, duplicates are by default hidden. * **πŸ‘₯ Near Duplicates (% vocab)**: The percentage of tokens in the vocabulary that are considered "near duplicates" - tokens that are very similar to each other, differing only in minor ways like case, tokenizer prefixes ('##', 'Δ ', '▁'), or whitespace. **Lower scores are better**, as they indicate a more efficient vocabulary without redundant tokens. This concept is discussed by [StaniΔ‡ et al., 2023](https://arxiv.org/abs/2309.11197), [Deiseroth et al., 2024](https://arxiv.org/abs/2406.19223) and [SchΓ€fer et al., 2024](https://aclanthology.org/2024.findings-acl.571/). We consider the following types of near-duplicates: - Tokenizer prefix near-duplicates: "##ing" and "ing" (where ## indicates a continuation token of an WordPiece tokenizer) - Case near-duplicates: "Hello", "hello" and "HELLO" - Whitespace near-duplicates: "\\n\\n}" and " }" (where \\n is a newline character) - All of the above * **βš™οΈ Algorithm (Framework)**: The underlying algorithm (e.g., BPE, WordPiece, Unigram, CharLevel, WordLevel or Unknown) and its implementation framework, displayed as `Algorithm (Framework)`. The frameworks are `hf-f` for fast Hugging Face `tokenizers`, `hf-s` for the base `transformers` library (slow), `sp` for SentencePiece, `tik` for Tiktoken and `fbpe` for FastBPE. ## Explore More Beyond the main πŸ† Leaderboard, you can also: * **πŸ“€ Submit Tokenizers**: Add new tokenizers to be ranked in the leaderboard. As long as the tokenizer is publicly hosted on Hugging Face, it can be submitted. * **βš”οΈ Tokenizer Battle**: Compare two tokenizers side-by-side on a custom text. * **πŸ› οΈ Offline Evaluation Guide**: Learn how to run evaluations locally on your own machine with your own tokenizers. Happy tokenizing! """ MULTILINGUAL_DESCRIPTION = """ ΒΉ The multilingual metrics of πŸͺΊ Fertility, πŸ—ƒ Compression, βš–οΈ Parity, πŸ”„ Reversible and ❓ Unknown are calculated by averaging across all selected languages. This keeps the metrics balanced across languages regardeless of dataset size. """.strip() OFFLINE_DOCS = """ ## Offline Model Evaluation using `tokenizer_evaluate.py` This guide explains how to evaluate a Hugging Face tokenizer model offline using the `tokenizer_evaluate.py` script. ### Prerequisites 1. **Clone the Repository**: You need to have the `tokenizer_evaluate.py` script and its dependencies. Clone the repository: ```bash git clone https://huggingface.co/spaces/eduagarcia/multilingual-tokenizer-leaderboard cd multilingual-tokenizer-leaderboard ``` 2. **Python Environment**: Ensure you have a Python environment (e.g., Python 3.8+) with `pip` installed. 3. **Install Dependencies**: Install the required Python packages: ```bash pip install -r requirements.txt ``` ### Evaluation Steps The `tokenizer_evaluate.py` script is used to run the benchmark. Here's a basic command structure: ```bash python tokenizer_evaluate.py [OPTIONS] ``` **Arguments:** - ``: (Required) The Hugging Face model name (e.g., `google-bert/bert-base-multilingual-cased`) or a local path to a tokenizer. **Common Options:** - `--langs ...`: Specify a subset of languages to test (e.g., `en pt de`). If omitted, all available languages in the benchmark dataset will be tested. Language codes can be found in `dataset_meta.yaml`. - `--force`: Force re-evaluation of the model even if results for it already exist. Useful if the benchmark data or tokenizer has changed. - `--trust-remote-code`: Add this flag if the tokenizer you are evaluating requires executing custom code from its Hugging Face repository. Use with caution. - `--output-path `: Specify an alternative path to copy the final results JSONL file to. The directory will be created if it doesn't exist. **Example Usage:** 1. **Evaluate `google-bert/bert-base-multilingual-cased` on all languages and save results:** ```bash python tokenizer_evaluate.py google-bert/bert-base-multilingual-cased ``` 2. **Evaluate a local tokenizer located at `./my_custom_tokenizer` for English (`en`) and Portuguese (`pt`) only, and force re-evaluation:** ```bash python tokenizer_evaluate.py ./my_custom_tokenizer --langs en pt --force ``` 3. **Evaluate `mistralai/Mistral-7B-v0.1` and also copy the results to a specific folder:** ```bash python tokenizer_evaluate.py mistralai/Mistral-7B-v0.1 --output-path /mnt/mistral_results/ ``` ### Understanding the Output - The script will print progress information to the console. - A JSONL file named `results__.jsonl` will be created in `data/results/`. This file contains detailed metrics for each language and domain evaluated. - A CSV file named `lang_overalls_.csv` will be created in the same folder as the JSONL file. This file contains a summary of fertility, compression, and token/word/char counts for each language aggreated. - A summary of fertility, compression, and token/word/char counts will be printed to the console. ### Data Directory (`DATA_DIR`) - The script relies on a `DATA_DIR` environment variable or a default value (usually `./data`). - The benchmark datasets are expected to be under `DATA_DIR/benchmark/` and results are stored in `DATA_DIR/results/`. - The first time you run the script, it will download the benchmark data from the Hugging Face Hub (`HF_REPO_BENCHMARK`) and existing results (`HF_REPO_RESULTS`). This might take some time. By following these steps, you can evaluate tokenizers locally and contribute to understanding their performance characteristics. """ with open('docs/about.md', 'r') as f: ABOUT_MARKDOWN = f.read()