ForkTogether AITogether AIpublished Jul 27, 2026seen 1d

togethercomputer/gigatoken

forked from marcelroed/gigatoken

Open original ↗

Captured source

source ↗
published Jul 27, 2026seen 1dcaptured 1dhttp 200method plain

togethercomputer/gigatoken

Description: Language model tokenization at GB/s

License: MIT

Stars: 0

Forks: 0

Open issues: 0

Created: 2026-07-27T04:09:57Z

Pushed: 2026-07-27T04:10:04Z

Default branch: main

Fork: yes

Parent repository: marcelroed/gigatoken

Archived: no

README:

Gigatoken

What is Gigatoken?

Gigatoken is the fastest tokenizer for language modeling. It supports a wide range of CPU hardware, and nearly all commonly used tokenizers. See the [Benchmarks](#benchmarks) section for detailed throughput numbers across tokenizers and CPUs.

Installation

pip install gigatoken

Usage

Gigatoken can be used with its own API, or in compatibility mode with HuggingFace Tokenizers or Tiktoken.

Compatibility Mode (Easiest)

import gigatoken as gt

# Minimum change from existing HuggingFace tokenizers usage (compatibility mode)
hf_tokenizer = ...
tokenizer = gt.Tokenizer(hf_tokenizer).as_hf()

# tokenizer can be used in the same contexts as hf_tokenizer
tokens = tokenizer.encode_batch(["This is a test string", "And here is another"])

# OR with tiktoken
tiktokenizer = ...
tokenizer = gt.Tokenizer(tiktokenizer).as_tiktoken()

# Now works like existing tiktoken tokenizers
tokens = tokenizer.encode_batch(["This is a test string", "And here is another"])

A substantial amount of effort has been put into making sure the outputs match exactly with what you would get with HuggingFace Tokenizers in this setting, but this is at a non-negligible cost to performance. You can still expect way faster performance across the board, but not quite the 1000x you will get with the Gigatoken API.

Gigatoken API (Fastest)

import gigatoken as gt

tokenizer = gt.Tokenizer("Qwen/Qwen3-8B") # Accepts HF model names
file_source = gt.TextFileSource(["owt_train.txt"], separator=b"")
tokens = tokenizer.encode_files(file_source)

Using the Gigatoken API lets the Rust implementation read data directly, and skips as much overhead as possible while allowing for maximum parallelism. Keep in mind that passing Python data structures through this API still incurs the overhead of reading from Python.

Benchmarks

Encoding throughput on owt_train.txt (11.9 GB) — AMD EPYC 9565 72-Core Processor x 2 sockets (144 cores)

| Tokenizer | gigatoken | HF tokenizers | tiktoken | vs HF | vs tiktoken | |---|---:|---:|---:|---:|---:| | GPT-2 | 24.53 GB/s | 24.8 MB/s | 36.0 MB/s | 989× | 681× | | Phi-4 | 24.00 GB/s | 29.9 MB/s | — | 801× | — | | GPT-OSS | 23.96 GB/s | 49.7 MB/s | 42.8 MB/s | 482× | 560× | | OLMo 2 / 3 | 23.06 GB/s | 27.7 MB/s | — | 833× | — | | Nemotron 3 | 22.79 GB/s | 49.4 MB/s | — | 462× | — | | Qwen 3 | 22.16 GB/s | 34.2 MB/s | — | 648× | — | | Llama 3 / 3.1 / 3.2 | 22.15 GB/s | 48.5 MB/s | — | 457× | — | | GLM 5 | 20.97 GB/s | 74.8 MB/s | — | 280× | — | | Llama 3.3 | 20.82 GB/s | 48.3 MB/s | — | 431× | — | | Llama 4 | 20.77 GB/s | 72.7 MB/s | — | 286× | — | | GLM 4 | 20.61 GB/s | 72.3 MB/s | — | 285× | — | | Phi-4-mini | 20.05 GB/s | 27.6 MB/s | — | 726× | — | | DeepSeek V3 / R1 / V4 | 19.69 GB/s | 26.2 MB/s | — | 750× | — | | Qwen 2 / 2.5 | 19.12 GB/s | 27.7 MB/s | — | 691× | — | | Kimi K2 | 18.85 GB/s | — | — | — | — | | Qwen 3.5 / 3.6 | 15.49 GB/s | 27.7 MB/s | — | 558× | — | | Gemma 4 | 4.82 GB/s | 334.1 MB/s | — | 14× | — | | ModernBERT | 4.18 GB/s | 26.9 MB/s | — | 155× | — | | Mistral 7B v0.3 | 3.57 GB/s | 354.7 MB/s | — | 10× | — | | TinyLlama / Phi-3 (Llama 2) | 3.48 GB/s | 323.6 MB/s | — | 11× | — | | CodeLlama | 3.47 GB/s | 347.4 MB/s | — | 10.0× | — | | Gemma 3 | 3.43 GB/s | 357.2 MB/s | — | 9.6× | — | | Gemma 1 | 2.51 GB/s | 342.2 MB/s | — | 7.3× | — |

Encoding throughput on owt_train.txt (11.9 GB) — Apple M4 Max (16 cores)

| Tokenizer | gigatoken | HF tokenizers | tiktoken | vs HF | vs tiktoken | |---|---:|---:|---:|---:|---:| | GPT-2 | 8.79 GB/s | 6.9 MB/s | 62.8 MB/s | 1,268× | 140× | | Nemotron 3 | 7.82 GB/s | 10.9 MB/s | — | 715× | — | | Phi-4 | 7.76 GB/s | 7.7 MB/s | — | 1,012× | — | | Llama 3 / 3.1 / 3.2 | 7.60 GB/s | 11.2 MB/s | — | 676× | — | | OLMo 2 / 3 | 7.56 GB/s | 5.8 MB/s | — | 1,299× | — | | Llama 3.3 | 7.50 GB/s | 15.7 MB/s | — | 479× | — | | Phi-4-mini | 6.97 GB/s | 7.2 MB/s | — | 964× | — | | Kimi K2 | 6.88 GB/s | — | — | — | — | | Llama 4 | 6.81 GB/s | 11.6 MB/s | — | 590× | — | | Qwen 2 / 2.5 | 6.37 GB/s | 5.8 MB/s | — | 1,105× | — | | Qwen 3 | 6.36 GB/s | 6.9 MB/s | — | 918× | — | | Qwen 3.5 / 3.6 | 6.31 GB/s | 6.3 MB/s | — | 994× | — | | GPT-OSS | 6.20 GB/s | 20.2 MB/s | 87.2 MB/s | 306× | 71× | | GLM 4 | 6.17 GB/s | 15.8 MB/s | — | 392× | — | | DeepSeek V3 / R1 / V4 | 5.68 GB/s | 7.2 MB/s | — | 788× | — | | GLM 5 | 5.55 GB/s | 12.2 MB/s | — | 456× | — | | ModernBERT | 2.64 GB/s | 5.8 MB/s | — | 452× | — | | Mistral 7B v0.3 | 1.99 GB/s | 95.1 MB/s | — | 21× | — | | Gemma 4 | 1.82 GB/s | 85.2 MB/s | — | 21× | — | | CodeLlama | 1.73 GB/s | 80.2 MB/s | — | 22× | — | | TinyLlama / Phi-3 (Llama 2) | 1.69 GB/s | 80.1 MB/s | — | 21× | — | | Gemma 1 | 1.42 GB/s | 85.7 MB/s | — | 17× | — | | Gemma 3 | 1.38 GB/s | 82.2 MB/s | — | 17× | — |

Encoding throughput on owt_train.txt (11.9 GB) — AMD Ryzen 7 9800X3D 8-Core Processor (16 cores)

| Tokenizer | gigatoken | HF tokenizers | tiktoken | vs HF | vs tiktoken | |---|---:|---:|---:|---:|---:| | GPT-2 | 6.27 GB/s | 59.0 MB/s | 92.1 MB/s | 106× | 68× | | Phi-4 | 6.09 GB/s | 55.4 MB/s | — | 110× | — | | OLMo 2 / 3 | 6.06 GB/s | 55.4 MB/s | — | 109× | — | | Phi-4-mini | 5.80 GB/s | 54.6 MB/s | — | 106× | — | | GPT-OSS | 5.68 GB/s | 79.6 MB/s | 112.7 MB/s | 71× | 50× | | Qwen 3 | 5.34 GB/s | 54.4 MB/s | — | 98× | — | | Qwen 2 / 2.5 | 5.30 GB/s | 51.7 MB/s | — | 103× | — | | Llama 3.3 | 5.26 GB/s | 79.9 MB/s | — | 66× | — | | Llama 3 / 3.1 / 3.2 | 5.24 GB/s | 79.5 MB/s | — | 66× | — | | Kimi K2 | 5.23 GB/s | — | — | — | — | | Qwen 3.5 / 3.6 | 5.22 GB/s | 51.6 MB/s | — | 101× | — | | Nemotron 3 | 5.20 GB/s | 79.0 MB/s | — | 66× | — | | GLM 5 | 5.05 GB/s | 79.5 MB/s | — | 63× | — | | GLM 4 | 5.04 GB/s | 79.5 MB/s | — | 63× | — | | Llama 4 | 5.03 GB/s | 78.2 MB/s | — | 64× | — | | DeepSeek V3 / R1 / V4 | 4.21 GB/s | 51.6 MB/s | — | 82× | — | | ModernBERT | 2.84 GB/s | 52.1 MB/s | — | 54× | — | | Mistral 7B v0.3 | 1.47 GB/s | 91.6 MB/s | — | 16× | — | | Gemma 4 | 1.45 GB/s | 78.8 MB/s | — | 18× | — | | CodeLlama | 1.38 GB/s | 85.2 MB/s | — | 16× | — | | TinyLlama / Phi-3 (Llama 2) | 1.37 GB/s | 84.9 MB/s | — | 16× | — | | Gemma 1 | 1.14 GB/s | 84.9 MB/s | — | 13× | — | |...

Excerpt shown — open the source for the full document.