Skip to content

VoiceDeepgram

Deepgram

Production speech-to-text. The STT default for many companies.

Category
Voice
Pricing
FREEMIUM
Hosting
Cloud
Platforms
API
Models
Single model (proprietary)
Verified
Jul 5, 2026

End-to-end speech recognition platform — real-time streaming, batch transcription, speaker diarization, and language detection. Strong on accented speech, telephony audio, and long-form recordings.

Capabilities 5

What it actually does — grouped by capability family.

  • Voice agent (secondary capability)
  • Transcription (STT) (primary capability)
  • Speaker diarization (secondary capability)
  • Speech synthesis (TTS) (secondary capability)
  • Summarization (secondary capability)

Pros & cons

  • Strong on accented/telephony audio
  • Real-time streaming + batch
  • Diarization and language detection
  • Low latency
  • API-only, no end-user app
  • Proprietary Nova models
  • English strongest, other langs vary

Tags

View all Voice
  • View AssemblyAI details
    VoiceFREEMIUM

    AssemblyAI

    AssemblyAI

    Production speech-to-text + audio intelligence API.

    Speech recognition API with batch and real-time streaming transcription, speaker diarization, and language detection. Its Universal models pair with optional Speech Understanding features (summarization, sentiment, redaction) so a single API can build conversation-intelligence products. Starts with a free credit and pay-as-you-go, per-second billing.

    High transcription accuracy
    Cloud-only, no self-host
    • stt
    • transcription
    • streaming
    • audio-intelligence
  • View ElevenLabs details
    VoiceFREEMIUM

    ElevenLabs

    ElevenLabs

    Text-to-speech, voice cloning, and multilingual dubbing.

    Hosted speech synthesis at near-human quality — TTS, voice cloning, multilingual dubbing, and conversational voice agents. Default choice when you need a voice that sounds like a person, not a robot.

    Best-in-class voice realism
    Pricier than commodity TTS at scale
    • tts
    • voice-cloning
    • dubbing
    • multilingual
  • View Cartesia details
    VoiceFREEMIUM

    Cartesia

    Cartesia

    Low-latency streaming text-to-speech for real-time voice.

    Streaming-first speech synthesis built around the Sonic family of state-space models. Aims at real-time agent voices where latency between turns is the product. Strong choice for sub-200ms voice loops.

    Streaming over WebSocket for fast first audio
    Long-form expressive texture trails ElevenLabs
    • tts
    • streaming
    • low-latency
    • real-time
  • View Speechify details
    VoiceFREEMIUM

    Speechify

    Speechify

    AI text-to-speech that reads any document, PDF, or page aloud.

    Speechify is an AI text-to-speech app that turns articles, PDFs, emails, and books into natural-sounding audio with high-definition voices, adjustable speed, and OCR for scanned text. It runs on iOS, Android, web, a browser extension, and desktop, and offers a separate Studio product plus a text-to-speech API for developers.

    OCR reads scanned text and PDFs
    Best features behind paywall
    • text-to-speech
    • read-aloud
    • accessibility
    • voice-cloning