AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing
We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To support this broad capability set, we construct approximately 3.03 billion instruction--audio instances and 1.95 million hours of effective supervision across five task families: speech generation, content editing, enhancement and separation, paralinguistic editing, and acoustic editing. AuK combines a multimodal large language model for semantic conditioning, an VAE jointly trained on speech, general audio, and music f
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 57%AI for editing existing Songs/Music? →
- PossiblePossibly related (embedding) · 56%STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation - Apple Machine Learning Research →
- PossiblePossibly related (embedding) · 54%Text-to-Speech AI Edits Single Words Mid-Recording: ViiTorVoice Goes Open Source - Tech Times →
- FuzzySimilar title/name (fuzzy) · 87%huggingface/speech-to-speech →
“Fuzzy title match (0.94): “AuK Technical Report: An Open-Source Foundational Model for ” ≈ “huggingface/speech-to-speech””
- FuzzyOverlapping authors or contributors · 62%affaan-m/ECC →
“Shared author/contributor keys: jiang”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%BerriAI/litellm →
“Shared author/contributor keys: jiang”
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
