Unified Audio Intelligence Without Regressing on Text Intelligence
Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text LLM built on Nemotron-Cascade-2-30B-A3B, a strong text-only MoE LLM. Audex adopts a simple unified design with a single Transformer decoder: audio inputs are encoded and projected into the text embedding space, while text tokens and quantized audio output tokens are treated uniformly during generation. This architecture enables strong audio-text fusion, seamless multimodal generation, and compatibility with standa
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 55%attevon-llc/OpenTranscribe →
- PossiblePossibly related (embedding) · 53%app-vox/vox →
- PossiblePossibly related (embedding) · 51%lgy1027/matrix-live-diarizer →
- PossiblePossibly related (embedding) · 49%morettt/my-neuro →
- PossiblePossibly related (embedding) · 48%jaswon/osu-dreamer →
- PossiblePossibly related (embedding) · 49%Intelligent transcription with Gemini 3.5 Transcribe →
- LinkedLinked via arxiv author · 85%Zhifeng Kong →
“Unified Audio Intelligence Without Regressing on Text Intelligence”
- LinkedLinked via arxiv author · 85%Sang-gil Lee →
“Unified Audio Intelligence Without Regressing on Text Intelligence”
