Audio AI tools
- VoiceStudio — VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video…
- ElevenLabs — Lifelike text-to-speech and voice cloning.
- VoxCPM — VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life…
- sherpa-onnx — Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using…
- GPT-SoVITS — 1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
- index-tts — An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
- edge-tts — Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or…
- pyvideotrans — Translate the video from one language to another and embed dubbing & subtitles.
- OpenVoice — Instant voice cloning by MIT and MyShell. Audio foundation model.
- CosyVoice — Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
- TTS — 🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
- voice-pro — Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning…
- Voice Cursor — Real-time Voice Agent works across apps on Mac
- piper — A fast, local neural text to speech system
- Gradium — Giving voice agents a voice that adapts to context
- myned-ai/pocket-tts-greek · Hugging Face — Greek and Cypriot Greek TTS that runs faster than real time on CPU
- ProgressCove — Calm and smart to-do app with Home Assistant integration
- EDA Benchmark Leaderboard — EDA Benchmark Leaderboard
- Hemory — Keep listening. Searchable memory for your AI agents.
- Lisen — Free Read Aloud with Cartesia Voices
- Kairn — Turn any recorded convo into actionables and follow ups
- Zeph — A $25 DIY alternative to $159 AI voice recorders – BYOK or local
- Speechka — Real-time voice translation that sounds like you
- Outbound benchmark calculators — No signup, no tracking
- Hola AI — An AI voicemail assistant that answers calls when you can’t.
- Robot Voice Bridge — Run ElevenLabs like a voiceover session, right on your Mac
- CC — AI agent that helps keep your family in sync
- Stardrift — A 3D galaxy of all music, to explore and learn
- SpeechText.AI — Play DSS and DS2 dictation files online without uploading your audio
- VoiceChanger.Live — Sound like anime girl in realtime on your next call
- TypeDash — Voice shortcuts and fast dictation for your desktop
- VoiceCap — The AI notetaker for meetings in your language
- Dictation API by AssemblyAI — Add fast, accurate dictation with a single line API call
- MosMos — Voice writing that works before, during, and after meetings
- Duvi — Make an AI for any business just using a URL
- In-Browser Sanskrit ASR Model — In-Browser Sanskrit ASR Model
- Airy TTS — 2K characters/second, $2 per 1M characters
- Threshyr — An offline automatic time tracker with on-device AI
- Ringee — Open-source human and AI calling
- CleverCrow — Get paid to work your backlog
- Famulor — White-label AI voice and chat agent platform for resellers
- Chronicle — Local-first, open-source authoring software for writers
- Jinfer — AI inference engine for the JVM. AI in a jar
- Voiskey — AI voice typing that sounds right in every app
- Kinapod Technology — Uses existing AirPods to capture motion data for sports tracking
- Understand AI — How LLMs work, explained through music, football or cricket analogies
- GhostWriter by MyHandler — Two taps and it's already written
- Epilude Notetaker — 100% private meeting notes
- Visiby — Track and grow your visibility across AI search
- CertArena — Free AI certification simulator (zero deps, native Node 24)
- Knockin' — Turns your static bio into an AI business card that replies
- Sparrow-2 — Noise cancellation isn't designed for conversational AI
- Airuncode — Run multiple local coding agents on your machine
- MacTap — Use knocks on your MacBook as shortcuts
- supertonic — Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.
- MockingBird — 🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
- Amphion — Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support…
- dia — A TTS model capable of generating ultra-realistic dialogue in one pass.
- Udio — Generate studio-quality music with vocals.
- Suno — Create full songs from a text prompt.
Opening Liz…