Skip to content

VoiceWispr

Wispr Flow

AI voice dictation that types for you across every app, on desktop and mobile.

Category
Voice
Pricing
FREEMIUM
Hosting
Cloud
Platforms
macOSWindowsiOSAndroid
Verified
Jun 7, 2026

A dictation tool that turns speech into clean, formatted text in any app — removing filler words and applying context-aware edits as you talk. One subscription works across macOS, Windows, iOS, and Android, syncing your custom vocabulary and snippets between devices.

Capabilities 3

What it actually does — grouped by capability family.

  • Dictation (primary capability)
  • Transcription (STT) (primary capability)
  • Editing / grammar (secondary capability)

Pros & cons

  • Types into any app system-wide
  • Auto-removes filler words
  • Custom vocabulary syncs across devices
  • Context-aware formatting
  • Subscription required for full use
  • Cloud processing, not local
  • Accuracy varies by accent/noise
  • No free unlimited tier

Tags

View all Voice
  • View TurboScribe details
    VoiceFREEMIUM

    TurboScribe

    TurboScribe

    Unlimited audio and video transcription powered by Whisper.

    TurboScribe converts audio and video files into text using OpenAI's Whisper speech-to-text model. It supports 98+ languages, automatic speaker recognition and translation of transcripts or subtitles into 134+ languages, with exports to formats like DOCX, PDF, SRT and VTT. A free tier allows three transcriptions a day, while the paid Unlimited plan removes usage caps for a flat monthly fee.

    Exports to DOCX, PDF, SRT, and VTT
    Free tier capped at 3 files/day, 30 minutes each
    • transcription
    • speech-to-text
    • whisper
    • subtitles
    • +1
  • View Phonic details
    VoicePAID

    Phonic

    Phonic

    Speech-to-speech platform for reliable voice agents.

    Phonic is a platform for building production voice agents on its own end-to-end speech-to-speech models, rather than chaining separate speech-to-text, LLM, and text-to-speech stages. It targets sub-300ms latency for natural turn-taking and reliable tool calling, and bundles evaluation, session records, and real-time observability to surface failure points. Aimed at enterprises, it offers cloud API access plus containerized deployment in your own environment.

    Reliable tool calling for voice agents
    Enterprise-focused, no public free tier
    • voice-agents
    • speech-to-speech
    • conversational-ai
    • low-latency
    • +1
  • View Regal details
    VoicePAID

    Regal

    Regal

    Voice AI agent platform for contact centers.

    A platform to build, deploy, and manage AI voice agents that handle inbound and outbound customer interactions across phone, SMS, chat, and WebRTC. Agents learn from past conversations, and a 2026 Copilot lets teams stand up a working voice agent within a day using existing business processes.

    Phone, SMS, chat, and WebRTC in one platform
    No public free tier — demo/sales-led
    • voice-agents
    • contact-center
    • cx
    • outbound
  • View SoundHound AI details
    VoicePAID

    SoundHound AI

    SoundHound AI

    Voice-native conversational AI platform for enterprise agents.

    SoundHound AI builds voice-native conversational AI used to deploy autonomous agents for customer interactions across automotive, restaurants, financial services, healthcare, and smart devices. Its full-stack speech technology pairs Speech-to-Meaning understanding with the Amelia enterprise agent platform and newer OASYS stack, automating phone answering, drive-thru ordering, and IT service management. It powers billions of conversations a year and is offered to developers and enterprises via APIs and embedded SDKs.

    Full-stack proprietary speech tech
    Enterprise focus, custom pricing
    • voice-ai
    • conversational-ai
    • customer-service
    • enterprise
    • +1
  • View LOVO AI details
    VoiceFREEMIUM

    LOVO AI

    LOVO, Inc.

    AI voice generation studio with voice cloning and a built-in video editor.

    LOVO AI's Genny platform generates text-to-speech voiceovers in 500+ voices across 100+ languages, with voice cloning from short samples and directable, expressive delivery. The browser-based studio bundles an online video editor, auto subtitles, an AI script writer and a developer API, targeting creators producing YouTube videos, e-learning, podcasts and ads.

    Large voice library across many languages
    Credit limits constrain heavy use
    • text-to-speech
    • voice-cloning
    • voiceover
    • video-editing
    • +1
  • View LMNT details
    VoiceFREEMIUM

    LMNT

    LMNT

    Streaming text-to-speech with voice cloning for real-time apps.

    LMNT is an AI text-to-speech platform that turns text into natural speech with ultra-low latency, built for conversational agents, games, and real-time apps. It supports instant voice cloning from a short sample and multilingual synthesis, and is exposed as a developer API plus a web playground. It is offered as a built-in voice provider across major voice-agent frameworks.

    Multilingual streaming synthesis
    Smaller voice library than ElevenLabs
    • text-to-speech
    • voice-cloning
    • low-latency
    • tts-api
  • View Neuphonic details
    VoiceFREEMIUMOpen core

    Neuphonic

    Neuphonic

    Ultra-low-latency text-to-speech that runs on-device.

    Neuphonic is a voice-AI company building text-to-speech and voice cloning that run locally with very low latency. Its cloud API targets real-time voice agents, and in October 2025 it open-sourced NeuTTS Air, a 748M-parameter speech language model that runs on CPU via llama.cpp and clones a voice from a few seconds of audio. Aimed at private, offline, and voice-agent use cases.

    On-device, CPU-only synthesis
    Cloud API pricing not clearly published
    • text-to-speech
    • voice-cloning
    • on-device
    • open-source
  • View WellSaid details
    VoiceFREEMIUM

    WellSaid

    WellSaid Labs

    Enterprise AI text-to-speech with voices licensed from real voice actors.

    An enterprise-grade AI voice generator that produces realistic voiceovers from scripts. It offers 120+ voices across languages and accents — modeled on licensed recordings by real voice actors — plus a studio for script import and audio tuning, team workspaces, pronunciation libraries, Adobe integrations, and an API for products, LMS platforms, and IVRs.

    120+ voices across languages and accents
    Voiceover-focused, not conversational/agent TTS
    • text-to-speech
    • voiceover
    • tts
    • enterprise
    • +1
  • View Smallest.ai details
    VoiceFREEMIUM

    Smallest.ai

    Smallest.ai

    Real-time voice AI: fast TTS and production phone agents.

    An enterprise voice-AI platform built on deliberately small, fast speech models. Waves handles text-to-speech, voice cloning, and conversion in 30+ languages, while Atoms is a real-time voice-agent platform that plugs into business systems for support, lead qualification, and outbound calls. The company says it can generate 10 seconds of speech in about 100 milliseconds for sub-second voicebot responsiveness.

    Very low TTS latency
    Younger, smaller company
    • tts
    • voice-agents
    • voice-cloning
    • low-latency
  • View Soniox details
    VoiceFREEMIUM

    Soniox

    Soniox

    One speech AI API for real-time transcription, TTS, and translation.

    A multilingual speech platform built on Soniox's own universal recognition model: real-time and async speech-to-text, text-to-speech, and any-to-any speech translation across 60+ languages from a single API. It returns token-level results within milliseconds and keeps transcribing through crosstalk, speaker overlap, and mid-sentence language switches. A consumer app (web and iOS) wraps the same engine for recording, transcription, and notes.

    60+ languages, mid-sentence switching
    Smaller brand than incumbents
    • stt
    • transcription
    • speech-translation
    • real-time
    • +1