GigaToken 是一款高性能分词器工具,在语言模型处理中展现出显著的速度优势 [1]。与 HuggingFace tokenizers 相比,该工具的分词速度快约 1000 倍 [1],可达到 GB/s 级别的吞吐量 [1]。
具体性能表现方面,GigaToken 在 Apple M4 Max 上的分词速度比传统方案快 1353.13 倍,在 AMD EPYC 9565 上快 989.21 倍 [1]。在 AMD EPYC 9565 处理器上,其吞吐量可达到 24532.45 MB/s,相当于 5564.94 Mtok/s [1]。基于这一性能水平,该工具可在 6.5 小时内完成对整个 Common Crawl 数据集(130 万亿 tokens)的分词处理 [1]。
GigaToken 采用 Rust 语言实现,通过 SIMD 指令集、缓存优化和减少 Python 交互等技术手段实现了性能突破 [1]。该项目支持 HuggingFace Tokenizers 和 Tiktoken 的兼容模式 [1],可作为这两款工具的即插即用替代品 [1],兼容多种 CPU 硬件架构和常见分词器 [1]。
GigaToken is a high-performance tokenizer for language models that delivers approximately 1000x faster performance compared to HuggingFace tokenizers [1]. The tool achieves throughput at the gigabyte-per-second scale, processing tokens at speeds of up to 5564.94 million tokens per second on AMD EPYC 9565 hardware, corresponding to 24532.45 MB/s [1].
The implementation, written in Rust, leverages SIMD instructions and cache optimization techniques to reach these speeds [1]. Benchmarks show performance gains of 1353.13x faster on Apple M4 Max and 989.21x faster on AMD EPYC 9565 compared to existing solutions [1]. The efficiency improvements stem from SIMD acceleration, cache optimization, and reduced Python interaction overhead [1].
GigaToken supports multiple CPU architectures and works as a drop-in replacement for both HuggingFace Tokenizers and Tiktoken, maintaining compatibility with existing workflows [1]. The tool's speed enables processing of massive datasets in practical timeframes—the entire Common Crawl corpus of 1.3 quadrillion tokens can be tokenized in approximately 6.5 hours [1].