BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech
Off-the-shelf TTS systems are poorly adapted to Taiwanese Mandarin. Their accent defaults to other Mandarin variants, their tokenizers over-segment common Taiwanese text, and their pronunciation degrades at code-switching boundaries where Chinese and English alternate within one utterance. These problems share one root: the text side lacks adaptation to the Taiwanese context. We address the text side from the bottom up. PangolinTokenizer, a byte-level BPE tokenizer trained on Taiwan-context data, reaches the lowest token rate (0.485 tokens/character) with the smallest vocabulary among nine tok
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 51%PacificAI/langtest →
- PossiblePossibly related (embedding) · 49%Kyutai's Pocket TTS clones a voice from 5 seconds of audio, on CPU, under MIT. Benchmarked against Kokoro, Supertonic, and Inflect-Nano for Eng. TTS →
- PossiblePossibly related (embedding) · 47%manojmallick/sigmap →
- PossiblePossibly related (embedding) · 47%Apil-Shrestha/token-efficiency-lex →
- PossiblePossibly related (embedding) · 46%cimeister/tokenizer-intrinsic-evals →
- PossiblePossibly related (embedding) · 49%Xiaohao-Liu/Awesome-Multi-Token-Prediction →
- FuzzySimilar title/name (fuzzy) · 87%huggingface/speech-to-speech →
“Fuzzy title match (0.94): “BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model,” ≈ “huggingface/speech-to-speech””
- FuzzyOverlapping authors or contributors · 62%browser-use/browser-use →
“Shared author/contributor keys: lee”
