FreyaTTS Technical Report
We introduce Freya-TTS, a compact, tokenizer-free, Turkish-first text-to-speech model designed for highly reliable and efficient conversational synthesis. Freya-TTS is a 183.2M-parameter non-autoregressive conditional flow-matching Diffusion Transformer (DiT) that operates in the frozen continuous latent space of AudioVAE2 (16 kHz encode, 48 kHz decode), allowing the model to focus its capacity on text-to-latent mapping while inheriting high-quality 48 kHz reconstruction. We advance the framework along three key dimensions: (1) rule-free end-to-end modeling from a 92-symbol Turkish character v
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 57%Whisper-Lite →
- PossiblePossibly related (embedding) · 55%lgy1027/matrix-live-diarizer →
- PossiblePossibly related (embedding) · 54%Transformer →
- PossiblePossibly related (embedding) · 48%huggingface/speech-to-speech →
- PossiblePossibly related (embedding) · 48%jaswon/osu-dreamer →
- LinkedLinked via arxiv author · 85%Ahmet Erdem Pamuk →
“FreyaTTS Technical Report”
- LinkedLinked via arxiv author · 85%Ömer Yentür →
“FreyaTTS Technical Report”
- LinkedLinked via arxiv author · 85%Ahmet Tunga Bayrak →
“FreyaTTS Technical Report”
