Best Audio & Speech AI Tools

AI audio tools for voice synthesis, music generation, and speech recognition

9 tools
Chat & AssistantsFoundation ModelsAI AgentsCode & DevelopmentImage GenerationVideo & AnimationAI AvatarsAudio & SpeechMusic GenerationWriting & ContentDesign & CreativeSearch & ResearchWebsite BuildersAutomation & WorkflowsAPI PlatformsOpen Source
ElevenLabs

ElevenLabs

Leading AI voice synthesis and cloning platform with multi-language support

FreemiumAPI
Whisper

Whisper

OpenAI's open-source speech recognition model for multi-language speech-to-text

FreeOpen Source
Fish Audio

Fish Audio

Voice generation platform from the team behind the open-source TTS star Fish Speech. Clone a voice from just 10-30 seconds of audio; the S1/S2 models deliver natural, expressive speech with commercial use and pay-as-you-go API

FreemiumOpen SourceAPI
MiniMax Audio

MiniMax Audio

MiniMax's voice generation platform. The Speech model family delivers hyper-realistic TTS in 40+ languages with 10-second voice cloning, controllable emotion and sound-effect tags, and a free web trial

FreemiumAPI
Cartesia

Cartesia

Real-time voice AI platform: Sonic 3.5 text-to-speech frequently ranks #1 for low latency, paired with Ink-2 transcription — built for voice agents and AI phone calls on Mamba/SSM research, with a free tier of ~20k credits a month

FreemiumAPI
Murf AI

Murf AI

An enterprise-grade AI voiceover and TTS platform: the Gen2 model delivers broadcast-quality voices in 20+ languages with fine-grained control over tone, pauses and emphasis, plus 200+ voices, voice cloning, multilingual dubbing and a developer API — free trial, Creator at $19/month billed annually

FreemiumAPI
Krisp

Krisp

The de-facto standard in AI noise cancellation: removes background noise and echo both ways across 800+ calling apps, takes AI meeting notes without a bot joining, and its 2025 flagship real-time Accent Conversion targets call centers, with an SDK for developers — free tier of 60 minutes a day, Pro around $8/month billed annually

FreemiumAPI
Voicebox

Voicebox

Open-source AI voice studio running locally — voice cloning, TTS across 7 engines, dictation, and agent voice in one app. A free alternative to ElevenLabs and WisprFlow, supporting 23 languages.

FreeOpen SourceAPI
VoxCPM

VoxCPM

OpenBMB's tokenizer-free Text-to-Speech system: directly generates continuous speech representations via diffusion autoregressive architecture. 2B parameter model, 30 languages, Voice Design, controllable voice cloning, 48kHz studio-quality audio.

Open SourceFree
Browse All