Tokenizer
9 items across the graph — tagged with Tokenizer.
From the graph · 9
Language model tokenization at GB/s
Syllable-aware BPE tokenizer for the Amharic language (አማርኛ) – fast, accurate, trainable.
a unix-like du command line tool to count token usage per files and directories
A high-performance tokenizer (BPE, WordPiece, SentencePiece) built with Rust with Python bindings, focused on speed, safety, and resource optimization.
TokEval: intrinsic quality metrics for tokenizers across natural language, code, and math
Small, independent TypeScript packages for LLM plumbing — token budgets, streaming JSON repair, cost accounting, retries, embedding caches. No provider SDKs.
A minimal, hackable Vision-Language Model built on Karpathy’s nanochat — add image understanding and multimodal chat for under $200 in compute.
Pure Rust GGUF tokenizer: 100% llama.cpp parity, zero C/C++ FFI, zero unsafe, WASM-ready. Free forever.
Dependency/tool package detected from repository manifests (github.com/tiktoken-go/tokenizer).
