Skip to content

VoiceSpeechmatics

Speechmatics

Enterprise speech APIs — real-time STT, TTS, and voice agents.

Category
Voice
Pricing
FREEMIUM
Hosting
Hybrid
Platforms
API
Models
Self-contained (on-device)
Verified
Jun 10, 2026

Speechmatics provides speech-to-text (batch and sub-second real-time), text-to-speech, and a Flow API for building voice agents, with accuracy that holds up across accents and dialects in 55+ languages. The same engine deploys in cloud, container, on-prem, or fully on-device, and it powers products from Adobe, LiveKit, and Ubisoft.

Capabilities 5

What it actually does — grouped by capability family.

  • Voice agent (secondary capability)
  • Transcription (STT) (primary capability)
  • Speech synthesis (TTS) (secondary capability)
  • Speaker diarization (secondary capability)
  • Speech translation (secondary capability)

Pros & cons

  • STT, TTS, and voice agents in one API
  • 55+ languages supported
  • Free tier (8 hrs of STT per month)
  • ISO 27001, SOC 2, HIPAA compliant
  • Pricier than budget STT rivals
  • TTS newer than its core STT
  • Enterprise-leaning packaging

Tags

Further reading

View all Voice
  • View Deepgram details
    VoiceFREEMIUM

    Deepgram

    Deepgram

    Production speech-to-text. The STT default for many companies.

    End-to-end speech recognition platform — real-time streaming, batch transcription, speaker diarization, and language detection. Strong on accented speech, telephony audio, and long-form recordings.

    Strong on accented/telephony audio
    API-only, no end-user app
    • stt
    • transcription
    • streaming
    • diarization
  • View AssemblyAI details
    VoiceFREEMIUM

    AssemblyAI

    AssemblyAI

    Production speech-to-text + audio intelligence API.

    Speech recognition API with batch and real-time streaming transcription, speaker diarization, and language detection. Its Universal models pair with optional Speech Understanding features (summarization, sentiment, redaction) so a single API can build conversation-intelligence products. Starts with a free credit and pay-as-you-go, per-second billing.

    High transcription accuracy
    Cloud-only, no self-host
    • stt
    • transcription
    • streaming
    • audio-intelligence
  • View ElevenLabs details
    VoiceFREEMIUM

    ElevenLabs

    ElevenLabs

    Text-to-speech, voice cloning, and multilingual dubbing.

    Hosted speech synthesis at near-human quality — TTS, voice cloning, multilingual dubbing, and conversational voice agents. Default choice when you need a voice that sounds like a person, not a robot.

    Best-in-class voice realism
    Pricier than commodity TTS at scale
    • tts
    • voice-cloning
    • dubbing
    • multilingual