repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago
cimeister/tokenizer-intrinsic-evals
TokEval: intrinsic quality metrics for tokenizers across natural language, code, and math
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages →
- PossiblePossibly related (embedding) · 51%MinGram: A Minimalist Unigram Tokenizer with High Compression and Competitive Morphological Alignment →
- PossiblePossibly related (embedding) · 49%Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement →
- PossiblePossibly related (embedding) · 49%How Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple Mitigation →
- PossiblePossibly related (embedding) · 47%Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking →
- PossiblePossibly related (embedding) · 46%BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech →
- PossiblePossibly related (embedding) · 47%The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs →
Implements
paperChallenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource LanguagespaperMinGram: A Minimalist Unigram Tokenizer with High Compression and Competitive Morphological AlignmentpaperAsk, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-ImprovementpaperHow Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple MitigationpaperClinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking
Implements (incoming)
Related across the graph
paperAsk, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-ImprovementpaperBlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching SpeechpaperThe Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMspaperChallenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource LanguagespaperClinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI BenchmarkingpaperHow Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple MitigationpaperMinGram: A Minimalist Unigram Tokenizer with High Compression and Competitive Morphological Alignment
