Skip to main content
Angestrom home
SearchPapersModelsLive AIIntelligence
Search⌕⌘K
EnterprisePricingSign in

Stay Ahead in the AI Revolution

Weekly digest — EPI pulse, top intelligence, fresh lineage. Free, no account.

Follow Angestrom
Global source network
Synced every 5 minutes

Continuous sync from primary AI sources — indexed, enriched, and queryable in real time.

arXivHugging FaceGitHubOpenAIAnthropicDeepMindReutersBBC TechHacker NewsReddit MLVerified feedsFunding
ANGESTROM

The Intelligence Layer of Humanity. Everything AI. All in One Place.

Angestrom connects every piece of the AI ecosystem — data, models, research, companies, tools, and people.

info@angestrom.comwww.angestrom.comLucknow, Uttar Pradesh, India

Product

  • AI Search
  • AI Models
  • Research Papers
  • Companies
  • News & Events
  • GitHub Explorer
  • APIs & Tools
  • Datasets
  • Benchmarks
  • Model lifecycle
  • Funding graph
  • Contributors
  • AI Agents

Resources

  • Weekly digest
  • Documentation
  • Tutorials
  • Guides
  • News
  • Help / Start
  • Community

Company

  • About
  • Contact
  • Privacy Policy
  • Terms of Service
  • Acceptable Use

Enterprise

  • Pricing
  • Workspace
  • Contact Sales

Developer

  • Developer Hub
  • API docs
  • GitHub

Learn

  • Learning Academy
  • Roadmaps
  • Glossary
  • AI for Beginners

Popular Topics

Loading topics…
View All Topics →
© 2026 Angestrom Intelligence Private Limited. All rights reserved.
English
Theme
Angestrom home
SearchPapersModelsLive AIIntelligence
Search⌕⌘K
EnterprisePricingSign in
  1. Home
  2. /Repositories
  3. /huggingface/speech-to-speech
Read original ↗
repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 8h ago

huggingface/speech-to-speech

Build local voice agents with open-source models

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 61%Looking for open-source AI meeting note-takers (like Fathom, Fireflies, Notion AI) →
  • PossiblePossibly related (embedding) · 54%How Loka Built a Natural, Low-Latency Voice Agent with Amazon Nova 2 Sonic →
  • FuzzySimilar title/name (fuzzy) · 87%Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation →

    “Fuzzy title match (0.94): “Auditing Protocol-Level Shortcuts in Large Audio Language Mo” ≈ “huggingface/speech-to-speech””

  • FuzzySimilar title/name (fuzzy) · 87%Self-supervised Speech Comparison for L2 Phone, Rhythm, and Intonation Scoring →

    “Fuzzy title match (0.94): “Self-supervised Speech Comparison for L2 Phone, Rhythm, and ” ≈ “huggingface/speech-to-speech””

  • FuzzySimilar title/name (fuzzy) · 87%BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech →

    “Fuzzy title match (0.94): “BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model,” ≈ “huggingface/speech-to-speech””

  • FuzzySimilar title/name (fuzzy) · 84%SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models →

    “Fuzzy title match (0.92): “SPEARBench: A Benchmark for Naturalness Evaluation in Stream” ≈ “huggingface/speech-to-speech””

  • FuzzySimilar title/name (fuzzy) · 87%From Black-Box to Clinical Insight: A Multi-Stage Explainable Framework for Speech-Based Cognitive Impairment Detection →

    “Fuzzy title match (0.94): “From Black-Box to Clinical Insight: A Multi-Stage Explainabl” ≈ “huggingface/speech-to-speech””

  • FuzzySimilar title/name (fuzzy) · 87%Comparing Human and Automatic Recognition of Dutch Dysarthric Continuous Speech: A Case Study →

    “Fuzzy title match (0.94): “Comparing Human and Automatic Recognition of Dutch Dysarthri” ≈ “huggingface/speech-to-speech””

Covers

newsLooking for open-source AI meeting note-takers (like Fathom, Fireflies, Notion AI)newsHow Loka Built a Natural, Low-Latency Voice Agent with Amazon Nova 2 Sonic

Implements

paperAuditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech EvaluationpaperSelf-supervised Speech Comparison for L2 Phone, Rhythm, and Intonation ScoringpaperBlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching SpeechpaperSPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language ModelspaperFrom Black-Box to Clinical Insight: A Multi-Stage Explainable Framework for Speech-Based Cognitive Impairment DetectionpaperComparing Human and Automatic Recognition of Dutch Dysarthric Continuous Speech: A Case StudypaperLuxEmo: Expressive Text-to-Speech Corpus for LuxembourgishpaperUnlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction TuningpaperVoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice ConversionpaperFrom Sinhala to Dhivehi: Cross-Lingual Transfer Learning for Low-Resource Speech RecognitionpaperRABBiT: Rapidly adaptive BOLD foundation model via brain-tuning for accurate zero-shot and few-shot prediction of speech-elicited responses in the brainpaperStreaming Neural Speech Codecs through Time-Invariant RepresentationspaperAudio-Native Speech Recognition with a Frozen Discrete-Diffusion Language ModelpaperContextual Semantic Relevance Tracks fMRI BOLD Responses During Naturalistic Speech ComprehensionpaperBenchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial ResultspaperContent is What Remains: Invariant Speech Tokenization from Parallel UtterancespaperImproving multichannel speech enhancement through accurate room-acoustic simulationspaperToward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language ModelspaperDONDO: Open w2v-BERT Speech-Recognition Base Models for African LanguagespaperMotor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease Detection

Covers (incoming)

newsVoice agents, demystified: STT+TTS and 4 demo agents you can talk to in the browser + build yours with RAG and ToolsnewsAs promised, here is the GitHub link for my 100% local voice-to-voice assistantnewsIntroducing Real World VoiceEQ: Measuring the human quality of voice AInewsAnthropic updates Claude voice mode with more capable modelsnewsOpenTune is a FREE open-source AI pitch correction tool for vocals - Bedroom Producers BlognewsLaunch HN: Speko (YC S26) – OpenRouter for Voice AI

Implements (incoming)

paperWordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTSpaperHierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMspaperFreyaTTS Technical Report

Related to (incoming)

paperFacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational AvatarsmodelWhisper-Lite

Related across the graph

paperSPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language ModelspaperFrom Black-Box to Clinical Insight: A Multi-Stage Explainable Framework for Speech-Based Cognitive Impairment DetectionmodelWhisper-LitepaperBenchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial ResultspaperContextual Semantic Relevance Tracks fMRI BOLD Responses During Naturalistic Speech ComprehensionpaperBlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching SpeechpaperHierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMsnewsLooking for open-source AI meeting note-takers (like Fathom, Fireflies, Notion AI)paperComparing Human and Automatic Recognition of Dutch Dysarthric Continuous Speech: A Case StudypaperFreyaTTS Technical ReportpaperLuxEmo: Expressive Text-to-Speech Corpus for LuxembourgishpaperWordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTSpaperImproving multichannel speech enhancement through accurate room-acoustic simulationspaperRABBiT: Rapidly adaptive BOLD foundation model via brain-tuning for accurate zero-shot and few-shot prediction of speech-elicited responses in the brainnewsHow Loka Built a Natural, Low-Latency Voice Agent with Amazon Nova 2 SonicpaperSelf-supervised Speech Comparison for L2 Phone, Rhythm, and Intonation ScoringpaperUnlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction TuningpaperMotor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease DetectionnewsOpenTune is a FREE open-source AI pitch correction tool for vocals - Bedroom Producers BlogpaperFrom Sinhala to Dhivehi: Cross-Lingual Transfer Learning for Low-Resource Speech RecognitionpaperAudio-Native Speech Recognition with a Frozen Discrete-Diffusion Language ModelpaperStreaming Neural Speech Codecs through Time-Invariant RepresentationsnewsLaunch HN: Speko (YC S26) – OpenRouter for Voice AInewsVoice agents, demystified: STT+TTS and 4 demo agents you can talk to in the browser + build yours with RAG and ToolspaperVoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice ConversionnewsAnthropic updates Claude voice mode with more capable modelsnewsIntroducing Real World VoiceEQ: Measuring the human quality of voice AIpaperToward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language ModelspaperFacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational AvatarspaperContent is What Remains: Invariant Speech Tokenization from Parallel UtterancespaperAuditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech EvaluationnewsAs promised, here is the GitHub link for my 100% local voice-to-voice assistantpaperDONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages
Knowledge path·PSPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models→PFrom Black-Box to Clinical Insight: A Multi-Stage Explainable Framework for Speech-Based Cognitive Impairment Detection→MWhisper-Lite→Rhuggingface/speech-to-speech

Topics

aiassistantlanguage-modelmachine-learningpythonspeechspeech-synthesisspeech-to-textspeech-translation

Explore

Search similar →Knowledge graph →All repos →Full intelligence feed →
Graph trust82Primary
Graph score12597