Whisper-Lite
A compact speech-to-text model for on-device use.
News6
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1up07mk/kyutais_pocket_tts_clones_a_voice_from_5_
newsText-to-Speech AI Edits Single Words Mid-Recording: ViiTorVoice Goes Open Source - Tech Times<a href="https://news.google.com/rss/articles/CBMiywFBVV95cUxQWUpsazBhRGxnT0dLMlhDb2NEWlh4QnlMVHN3VW10dlhfU3pDVGpiMHBFTD
newsShow HN: Orate – On-device neural text-to-speech queue for Mac<p>Hey friends! Brad here - long time HN user, founder, engineer (Coinbase, Soothe, Jeevz, Release, Hired, GrowthX, etc)
newsApple's new SpeechAnalyzer API, benchmarked against Whisper and its predecessor<p>Article URL: <a href="https://get-inscribe.com/blog/apple-speech-api-benchmark.html">https://get-inscribe.com/blog/ap
newsVocalinux 0.14 Beta Released For Offline Voice Dictation / Speech-To-Text On Linux - Phoronix<a href="https://news.google.com/rss/articles/CBMiXkFVX3lxTE8xSDZaeEprM2liWG9VXzhlelpLblR3bmhhRmtscC1LUk05bFk4cUhvQmhsN1
newsSeven on-device large language models for mobile phones have received regulatory approval! Apple Intelligence is among them, with support from Alibaba and Baidu. - Moomoo<a href="https://news.google.com/rss/articles/CBMipAFBVV95cUxQMER2MWo4dXVNWjRnWFRUWHJSYWtDR2NVcDZfRm8xeVhzT1Q3RXg5VXlQN0
Repos13
Machine learning powered Karaoke app (with scores!)
repoattevon-llc/OpenTranscribeSelf-hosted AI-powered transcription platform with speaker diarization, search, and collaboration features. Built with S
repohuggingface/speech-to-speechBuild local voice agents with open-source models
repoOpen-Less/openlessHold a key, speak, release — AI-polished text appears at your cursor in any app. Open-source voice input for macOS & Win
repoRYOITABASHI/ShellyAI-powered chat-first terminal IDE for Android. Built entirely on a phone, by someone who can't write code.
repomirkobozzetto/flowflowVoice notes for iPhone and macOS - 100% Rust, Dioxus, local-first (SQLite + LanceDB + RIG)
repoamplitudesoldierheed/AI-Voice-Changer-Real-Time-Desktoprepotover0314-w/opentypelessOpen-source AI voice typing for macOS, Windows, and Linux. Press a hotkey, speak naturally, get polished text in any app
Papers11
While Large Multimodal Models excel in comprehension, high-throughput inference engines lack native support for multimod
paperComparing Human and Automatic Recognition of Dutch Dysarthric Continuous Speech: A Case StudyIn our goal to develop personalised dysarthric speech recognition (DSR) models, this study compared the recognition perf
paperFreyaTTS Technical ReportWe introduce Freya-TTS, a compact, tokenizer-free, Turkish-first text-to-speech model designed for highly reliable and e
paperLuxEmo: Expressive Text-to-Speech Corpus for LuxembourgishState-of-the-art speech datasets predominantly focus on widely spoken languages, often overlooking low-resource language
paperWordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTSWhile recent Large Language Model (LLM)-based Text-to-Speech (TTS) systems have achieved remarkable naturalness, they pr
paperUnlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction TuningInstruction tuning for speech language models (SLMs) is substantially more challenging than for text-based large languag
paperAudio-Native Speech Recognition with a Frozen Discrete-Diffusion Language ModelAutomatic speech recognition is dominated by autoregressive decoders that emit one token at a time. We ask whether a dis
paperSIMAX: A Scalable and Interpretable Framework for Multi-Fidelity and Annotated Clinician-Patient Dialogue SimulationBackground. The widespread deployment of ambient digital scribes is driving large-scale capture of clinician-patient dia
