repoGitHubTrust 82 · PrimaryPublished 4d agoLive · 3d ago
georg-jung/FastBertTokenizer
Fast and memory-efficient library for WordPiece tokenization as it is used by BERT.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 56%Token →
- PossiblePossibly related (embedding) · 48%Gigatoken: A new open source tokenizer ~100x faster than Tiktoken, -500-1000x faster than Huggingface →
- PossiblePossibly related (embedding) · 61%GigaToken: ~1000x faster Language model tokenization →
