Skip to content

Audio AI apps

AI audio generation and processing — sound design, enhancement, dubbing, and speech-to-text engines.

30 apps · researched & kept current by Claude Code

Filter & search these 30 apps
  • View CassetteAI details
    MusicFREEMIUM

    CassetteAI

    Pixl Technologies, Inc.

    Real-time generative audio API for music, sound effects, and speech.

    CassetteAI is a generative audio platform that produces music, sound effects, and speech from text prompts through a single API. Its latent-diffusion models render a 30-second sample in roughly two seconds and full multi-minute tracks at 44.1 kHz stereo, aimed at creators and developers who need audio on demand. It is offered pay-per-use with a free monthly tier.

    Sub-second generation latency
    Newer, smaller catalog than Suno/Udio
    • music-generation
    • sound-effects
    • text-to-audio
    • api
    • +1
  • View Podcastle details
    AudioFREEMIUM

    Podcastle

    Podcastle

    AI-powered studio for recording, editing, and producing audio and video.

    Podcastle is a browser-based content-creation platform for podcasters and video creators: multi-track remote recording, AI text-to-speech and voice cloning, transcription, and one-click audio and video editing with noise removal and leveling. It bundles studio capture, AI voices, and editing into a single workflow so creators don't need a separate DAW or video editor.

    Recording + AI voices + editing in one app
    Cloud-only; needs a connection
    • podcasting
    • text-to-speech
    • voice-cloning
    • transcription
    • +1
  • View Rev details
    AudioFREEMIUM

    Rev

    Rev

    AI and human transcription, captions, and speech-to-text.

    Transcription and captioning service offering fast AI transcription (~96% accuracy) plus on-demand human transcription at 99%+ accuracy, captions, and subtitles in dozens of languages. Adds AI note-taking and clip generation, and exposes its in-house ASR to developers via the Rev AI speech-to-text API.

    Human-reviewed tier for top accuracy
    Human transcription costs add up per minute
    • transcription
    • captions
    • speech-to-text
    • subtitles
  • View LOVO AI details
    VoiceFREEMIUM

    LOVO AI

    LOVO, Inc.

    AI voice generation studio with voice cloning and a built-in video editor.

    LOVO AI's Genny platform generates text-to-speech voiceovers in 500+ voices across 100+ languages, with voice cloning from short samples and directable, expressive delivery. The browser-based studio bundles an online video editor, auto subtitles, an AI script writer and a developer API, targeting creators producing YouTube videos, e-learning, podcasts and ads.

    Large voice library across many languages
    Credit limits constrain heavy use
    • text-to-speech
    • voice-cloning
    • voiceover
    • video-editing
    • +1
  • View Alitu details
    AudioPAID

    Alitu

    The Podcast Host

    All-in-one podcast maker: record, auto-clean, edit, and host.

    Alitu is a podcast production platform that guides creators through recording, automatic audio cleanup, transcript-based editing, and publishing in one app. It automates the technical work — noise removal, leveling, and processing — so non-technical podcasters can record and publish the same day, with built-in hosting and one-click distribution to Apple Podcasts and Spotify. Made by The Podcast Host team.

    Beginner-friendly guided workflow
    Subscription with no free tier
    • podcast
    • audio-editing
    • podcast-hosting
    • noise-removal
    • +1
  • View Podsqueeze details
    AudioFREEMIUM

    Podsqueeze

    Podsqueeze

    Turn podcast episodes into show notes, clips, and posts in one click.

    Podsqueeze is an AI assistant for podcasters that takes an episode and generates show notes, timestamps, titles, transcripts, newsletters, blog posts, and social clips automatically. It targets the post-production and promotion busywork that follows recording, repurposing a single upload into a full set of distribution assets. A bootstrapped indie product used by tens of thousands of podcasters and agencies.

    Many asset types from one episode
    Outputs need human editing and polish
    • podcast
    • show-notes
    • transcription
    • content-repurposing
    • +1
  • View pyannoteAI details
    AudioFREEMIUMOpen core

    pyannoteAI

    pyannoteAI

    Speaker intelligence — diarization that tells who spoke when.

    pyannoteAI turns conversational audio into speaker-attributed transcripts: it identifies speakers, separates overlapping voices, and provides speaker metadata. Built on the widely used open-source pyannote.audio library, it adds a premium REST API and Python SDK with higher accuracy and near real-time speed.

    State-of-the-art diarization accuracy
    Diarization only, not transcription
    • audio
    • speaker-diarization
    • speech
    • open-source
    • +1
  • View Songscription details
    AudioFREEMIUM

    Songscription

    Songscription

    Turn any audio into sheet music, MIDI, and guitar tabs with AI.

    Songscription transcribes audio recordings into readable notation — sheet music, MIDI, MusicXML, and GuitarPro tabs — across instruments including piano, guitar, bass, strings, horns, drums, and vocals. Often described as a 'Shazam for sheet music', it automates a task that traditionally took hours by ear. A free tier covers unlimited 30-second clips, with paid plans for longer recordings and more export formats.

    Sheet music, MIDI, MusicXML, tabs
    Free tier capped at 30-second clips
    • music-transcription
    • sheet-music
    • midi
  • View Happy Scribe details
    AudioFREEMIUM

    Happy Scribe

    Happy Scribe

    Transcription, subtitles, and AI meeting notes.

    A transcription and subtitling platform that turns calls, interviews, and recordings into accurate, searchable text. It offers instant AI transcription, a professional human transcriber network for 99%+ accuracy, automatic subtitles and translation across 150+ languages, and an AI notetaker that joins Google Meet, Microsoft Teams, and Zoom calls.

    Optional human transcriber network
    Volume-priced by audio minute
    • transcription
    • subtitles
    • translation
    • notetaker
    • +1
  • View Gaudio Studio details
    AudioFREEMIUM

    Gaudio Studio

    Gaudio Lab

    AI stem separation that splits any song into vocals and instruments.

    Web and mobile tool that separates a track into up to six stems — vocals, drums, bass, piano, guitar, and other instruments — using Gaudio Lab's GSEP source-separation model. Accepts wav, flac, m4a, and mp3 uploads or a YouTube URL, and is aimed at remixers, producers, and musicians who want clean isolated parts for practice or rework.

    Up to six separate stems per track
    Paid use is credit-based, not flat-rate
    • stem-separation
    • vocal-remover
    • music-production
    • remix
  • View Snipd details
    AudioFREEMIUM

    Snipd

    Snipd

    AI podcast player with chapters, episode summaries, and shareable highlights.

    Snipd is an AI-powered podcast app that transcribes episodes, breaks them into chapters, and generates summaries and key takeaways you can read before listening. While playing, it detects highlight-worthy moments and lets you save them as 'snips' — short clips with auto-generated transcripts and notes — and you can chat with an episode's content. Built by a Zurich-based team, it positions podcast listening as active learning.

    AI chapters and episode summaries
    Mobile-first; web app is newer
    • podcasts
    • summarization
    • transcription
    • note-taking
    • +1
  • View Fish Audio details
    AudioFREEMIUM

    Fish Audio

    Fish Audio

    Expressive, emotionally controllable text-to-speech, voice cloning, and voice agents.

    Fish Audio is a voice AI platform for real-time text-to-speech with emotion tags, voice cloning from clips as short as 15 seconds, speech-to-text, and end-to-end voice agents. Its flagship S2 model targets natural, expressive, multilingual narration, and the company maintains the open-source Fish Speech (OpenAudio) models on GitHub. A free tier covers personal use, with paid plans and an API for commercial use.

    Expressive, emotion-controllable TTS
    Hosted platform itself is proprietary
    • text-to-speech
    • voice-cloning
    • speech-to-text
    • voice-agents
    • +1
  • View Gladia details
    VoiceFREEMIUM

    Gladia

    Gladia

    Real-time speech-to-text and audio intelligence through a single API.

    End-to-end audio infrastructure to record, transcribe, and enrich speech via one API — real-time streaming under ~300ms latency, batch transcription, diarization, translation, and summarization across 100+ languages. Reengineered Whisper for production before shipping its own Solaria models, with EU data residency for compliance-bound teams.

    Low-latency real-time streaming
    API-only, no end-user app
    • stt
    • transcription
    • streaming
    • diarization
    • +1
  • View MacWhisper details
    AudioFREEMIUM

    MacWhisper

    Good Snooze

    Private, on-device audio and video transcription for Mac.

    MacWhisper transcribes audio and video files entirely on your Mac — no internet connection and no data leaving the machine. Drop in MP3, WAV, MOV, or MP4 and get a transcript with speaker labels, timestamps, and 50+ export formats. It runs OpenAI's open-source Whisper (and Nvidia's Parakeet) locally, hitting roughly 30x realtime on Apple Silicon via Metal acceleration.

    One-time lifetime license option
    macOS only (no Windows/Linux/web)
    • transcription
    • on-device
    • whisper
    • privacy
    • +1
  • View Good Tape details
    AudioFREEMIUM

    Good Tape

    Good Tape

    Automated AI transcription for audio and video.

    Good Tape is an AI transcription tool that turns audio and video into accurate, searchable text across 100+ languages with automatic language detection. It adds speaker identification, AI summaries, and smart search, and is built around a privacy-first posture: GDPR compliance, ISO 27001 certification, EU-based servers, AES-256 encryption, and a commitment not to use uploads to train AI. It runs on the web plus iOS and Android apps.

    Accurate speech-to-text in 100+ languages
    Free tier capped at a few short files
    • transcription
    • speech-to-text
    • privacy
    • multilingual
  • View Async details
    AudioFREEMIUM

    Async

    Async

    AI studio to record, edit, and produce podcasts and video.

    Async is a browser-based creative studio for recording, editing, and producing podcasts and video, with studio-quality remote recording, text-based editing, AI enhancement, and multilingual dubbing. Its Revoice feature clones a voice from a short sample to fix mistakes without re-recording, and a developer Voice API exposes its 1,000+ AI voices. It was known as Podcastle until a 2026 rebrand.

    All-in-one record, edit, and publish workflow
    No Android app — web only on Android
    • podcasting
    • voice-clone
    • audio-editing
    • dubbing
  • View Cleanvoice AI details
    AudioPAID

    Cleanvoice AI

    Cleanvoice

    Automatically edit podcasts and audio in minutes.

    An AI audio/video editor that removes filler words, background noise, mouth sounds, stutters, and dead silences from recordings, then re-levels the audio. Aimed at podcasters and creators who want a clean cut without manual waveform editing.

    Removes filler words and silences
    No perpetual free tier (trial only)
    • podcast
    • audio-editing
    • filler-words
    • noise-removal
  • View Riverside details
    AudioFREEMIUM

    Riverside

    RiversideFM

    Record studio-quality podcasts and video remotely, then edit with AI.

    Riverside is a browser-based studio that records each participant locally in separate tracks — up to 4K video and uncompressed audio — so quality isn't tied to anyone's internet connection. Its AI layer handles text-based editing, filler-word and background-noise removal, automatic clip generation, show notes, captions, and translation/dubbing into 30+ languages. A 2025 chat-based editor lets creators cut and repurpose footage by talking to an agent.

    Separate uncompressed track per guest
    Local upload can be slow on weak hardware
    • podcast
    • recording
    • video
    • transcription
    • +1
  • View LANDR details
    AudioFREEMIUM

    LANDR

    LANDR Audio

    AI mastering, distribution, samples, and plugins for independent musicians.

    A cloud platform for musicians built around the original AI mastering engine, launched in 2014. Upload a track and get a streaming-ready master in minutes, then distribute to Spotify, Apple Music, and other stores from the same account. The subscription also covers a million-plus sample library and plugin deals; a free account includes a limited number of masters per month.

    Masters ready in minutes
    Can over-compress dynamics
    • mastering
    • music-production
    • distribution
    • samples
  • View Auphonic details
    AudioFREEMIUM

    Auphonic

    Auphonic GmbH

    Your AI sound engineer: automatic leveling, denoising, and loudness mastering for podcasts.

    An AI audio post-production service that masters podcast and video sound in one pass — intelligent leveling, AI noise and reverb reduction, AutoEQ, and loudness normalization to targets like EBU R128, -16 LUFS podcast norms, and ACX audiobook specs. Automatic cutting removes silence, filler words, and coughs; multitrack productions get ducking and mic-bleed removal. Self-hosted Whisper transcription feeds AI shownotes and chapters, with output publishing straight to podcast hosts.

    One-pass leveling, denoise, and AutoEQ
    A mastering pipeline, not a full editing timeline
    • podcasting
    • mastering
    • loudness
    • noise-reduction
    • +1
  • View Wondercraft details
    AudioFREEMIUM

    Wondercraft

    Wondercraft

    AI audio studio that turns ideas into produced podcasts and audiobooks.

    Wondercraft is a browser-based AI audio studio that turns ideas, documents, or URLs into fully produced audio — podcasts, audiobooks, and ads — complete with scripts, natural-sounding voices, music, and sound effects. It offers hundreds of AI voices across dozens of languages plus voice cloning, with a timeline editor for assembling and revising episodes. Free for individuals, with paid creator and business tiers.

    End-to-end podcast/audiobook production
    Output can sound templated
    • podcast
    • audio-generation
    • text-to-speech
    • voice-cloning
  • View Adobe Podcast details
    AudioFREEMIUM

    Adobe Podcast

    Adobe

    Web-based AI audio recording, editing, and speech enhancement.

    Adobe's browser-based suite for recording and cleaning up spoken audio. Its flagship Enhance Speech filter uses AI to remove background noise and echo and rebuild a voice to sound as if recorded in a soundproofed studio, working on audio and video files. A Studio environment (beta) lets you record, edit, and enhance entirely in the browser, with a free tier and an Adobe Podcast Premium plan for higher limits.

    Removes background noise and echo
    Free tier caps daily processing hours
    • audio-enhancement
    • podcast
    • noise-removal
    • speech
  • View Moises details
    AudioFREEMIUM

    Moises

    Moises Systems

    The musician's app: AI stem separation, chord detection, and practice tools.

    An AI music suite that splits any track into isolated stems (vocals, drums, bass, and more), removes vocals, and detects chords and key. It adds practice-focused tools — pitch and tempo control, a smart metronome, and recording — across web, desktop, and native mobile apps.

    All-in-one practice suite, not just stems
    Separation quality trails pro-grade splitters
    • stem-separation
    • music
    • vocal-remover
    • practice
  • View Stable Audio details
    AudioFREEMIUM

    Stable Audio

    Stability AI

    Generative AI for music and sound effects from a text prompt.

    Stability AI's text-to-audio tool: describe a track or sound effect and it generates studio-quality stereo audio, with structured full-length songs in later versions. A web studio plus a generation API on the Stability platform. Subscriptions add longer outputs, more monthly generations, and commercial licensing.

    Music + sound-effects generation
    Vocals weaker than song-first rivals
    • text-to-music
    • sound-effects
    • music-generation
    • licensed-data
  • View Kits AI details
    AudioFREEMIUM

    Kits AI

    Kits AI

    Studio-quality AI voice cloning and music tools for artists and producers.

    Kits AI lets musicians clone or use studio-quality singing and speaking voices, convert vocals between voices, and run a suite of audio tools (stem splitting, mastering, vocal cleanup). Voice models are ethically licensed from artists with revenue-sharing rather than scraped. A free plan is available, with paid Starter, Producer, and Professional tiers.

    Ethically licensed artist voices with payouts
    Cloned-voice quality reviews are mixed
    • voice-cloning
    • music
    • vocals
    • stem-separation
    • +1
  • View AudioShake details
    AudioFREEMIUM

    AudioShake

    AudioShake

    AI audio separation that splits any track into clean instrument stems.

    AudioShake uses AI to separate finished recordings into clean, performance-quality stems — vocals, drums, bass, guitar, and more — plus dialogue/music/effects splits and word-synced lyric transcription. It is offered as the Indie web app for artists, Live for studios, and a REST API/SDK for developers. Major labels and studios use it for remixing, dubbing, and remastering.

    High-quality stem separation
    Closed source
    • stem-separation
    • music
    • audio-processing
    • remixing
  • View LALAL.AI details
    AudioFREEMIUM

    LALAL.AI

    LALAL.AI

    AI vocal remover and stem separation, built for pro-level quality.

    An AI audio service that removes vocals and splits any track into clean stems — vocals, drums, bass, guitars, piano, and more — using transformer-based separation models. It adds voice cleaning, echo and reverb removal, and lead/backing vocal separation. Available on the web, desktop, native iOS and Android apps, a DAW plugin, and a developer API.

    Top-tier isolated vocal cleanliness
    Max-quality burns 2x+ the audio length in minutes
    • stem-separation
    • vocal-remover
    • music
    • audio-cleanup
  • View Krisp details
    AudioFREEMIUM

    Krisp

    Krisp

    On-device AI noise cancellation, transcription, and meeting notes.

    Voice AI platform that removes background noise, transcribes calls, and generates meeting notes. It installs as a virtual microphone/speaker, so noise cancellation works across Zoom, Teams, Meet, and 800+ other apps without joining as a bot. Also offers accent conversion and a call-center product on the same engine.

    Strong real-time noise removal
    Transcription covers ~16 languages only
    • noise-cancellation
    • meeting-notes
    • transcription
    • voice
    • +1
  • View Hume AI details
    VoiceFREEMIUM

    Hume AI

    Hume AI

    Empathic Voice Interface — speech-to-speech AI that hears tone.

    A voice AI toolkit built around the Empathic Voice Interface (EVI), a speech-to-speech model that infers emotion and prosody from a user's voice and modulates its replies accordingly. Exposed as an API for building expressive voice agents and assistants. From a research lab focused on emotional intelligence in AI.

    Emotion/prosody-aware voice interface
    Emotion inference accuracy is contested
    • voice
    • speech-to-speech
    • emotion
    • api
  • View Descript details
    AudioFREEMIUM

    Descript

    Descript

    Podcast + audio editing where the transcript is the timeline.

    Audio and video editing built around an editable transcript — cut words, get cut audio. Add AI cleanup, overdub voice clones, and screen-recording for podcasts and tutorials in one tool.

    Transcript-as-timeline editing
    Transcription accuracy varies by audio
    • editing
    • podcasts
    • transcript
    • voice-clone