repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 27d ago
lgy1027/matrix-live-diarizer
Local-first real-time meeting transcription with speaker diarization, switchable ASR engines, and optional OpenAI-compatible LLM summaries.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 62%NAVER LABS Europe Submission to the Instruction-following 2026 Short Track →
- PossiblePossibly related (embedding) · 55%Adapting Foundation ASR Models to Dysarthric Speech: A Case Study →
- PossiblePossibly related (embedding) · 53%Whisper-Lite →
- PossiblePossibly related (embedding) · 52%AMR: Adaptive Modality Routing for Multimodal Polyglot Speaker Identification →
- PossiblePossibly related (embedding) · 50%LOPA: Enhancing Spoken Language Assessment via Latent Ordinal Prototype Alignment →
- PossiblePossibly related (embedding) · 51%Unified Audio Intelligence Without Regressing on Text Intelligence →
- PossiblePossibly related (embedding) · 51%Streaming Neural Speech Codecs through Time-Invariant Representations →
- PossiblePossibly related (embedding) · 49%REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing →
Implements
paperNAVER LABS Europe Submission to the Instruction-following 2026 Short TrackpaperAdapting Foundation ASR Models to Dysarthric Speech: A Case StudypaperAMR: Adaptive Modality Routing for Multimodal Polyglot Speaker IdentificationpaperLOPA: Enhancing Spoken Language Assessment via Latent Ordinal Prototype Alignment
Related to
Implements (incoming)
paperUnified Audio Intelligence Without Regressing on Text IntelligencepaperStreaming Neural Speech Codecs through Time-Invariant RepresentationspaperREDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution EditingpaperWordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTSpaperWhen Synthetic Speech Is All You Have: Better Call GRPOpaperFreyaTTS Technical ReportpaperVoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice ConversionpaperEncoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models
Covers (incoming)
newsKyutai's Pocket TTS clones a voice from 5 seconds of audio, on CPU, under MIT. Benchmarked against Kokoro, Supertonic, and Inflect-Nano for Eng. TTSnewsEdge AI ASL Recognition on Raspberry Pi 5 – Looking for Feedback on My System Design [P]newsBest Dictation Software for Windows 2026, Wispr Flow vs Dragon vs Free Open Source Honest Comparison - YouTube
Related across the graph
newsKyutai's Pocket TTS clones a voice from 5 seconds of audio, on CPU, under MIT. Benchmarked against Kokoro, Supertonic, and Inflect-Nano for Eng. TTSmodelWhisper-LitepaperEncoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language ModelspaperFreyaTTS Technical ReportpaperAMR: Adaptive Modality Routing for Multimodal Polyglot Speaker IdentificationpaperWordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTSpaperREDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution EditingnewsEdge AI ASL Recognition on Raspberry Pi 5 – Looking for Feedback on My System Design [P]paperStreaming Neural Speech Codecs through Time-Invariant RepresentationspaperVoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice ConversionpaperAdapting Foundation ASR Models to Dysarthric Speech: A Case StudynewsBest Dictation Software for Windows 2026, Wispr Flow vs Dragon vs Free Open Source Honest Comparison - YouTubepaperNAVER LABS Europe Submission to the Instruction-following 2026 Short TrackpaperLOPA: Enhancing Spoken Language Assessment via Latent Ordinal Prototype AlignmentpaperWhen Synthetic Speech Is All You Have: Better Call GRPOpaperUnified Audio Intelligence Without Regressing on Text Intelligence
