Skip to content

VoiceHume AI

Hume AI

Empathic Voice Interface — speech-to-speech AI that hears tone.

Categories
VoiceAudio
Pricing
FREEMIUM
Hosting
Cloud
Platforms
WebAPI
Models
Multi-model
Verified
Jun 7, 2026

A voice AI toolkit built around the Empathic Voice Interface (EVI), a speech-to-speech model that infers emotion and prosody from a user's voice and modulates its replies accordingly. Exposed as an API for building expressive voice agents and assistants. From a research lab focused on emotional intelligence in AI.

Capabilities 3

What it actually does — grouped by capability family.

  • Voice agent (primary capability)
  • Speech synthesis (TTS) (secondary capability)
  • Voice cloning (secondary capability)

Pros & cons

  • Emotion/prosody-aware voice interface
  • Speech-to-speech, low-latency replies
  • Pairs with a configurable LLM
  • Research-grade emotion models
  • Emotion inference accuracy is contested
  • Narrower than full TTS/STT suites
  • Usage-metered pricing
  • Smaller ecosystem than ElevenLabs

Tags

Further reading

View all Voice
  • View ElevenLabs details
    VoiceFREEMIUM

    ElevenLabs

    ElevenLabs

    Text-to-speech, voice cloning, and multilingual dubbing.

    Hosted speech synthesis at near-human quality — TTS, voice cloning, multilingual dubbing, and conversational voice agents. Default choice when you need a voice that sounds like a person, not a robot.

    Best-in-class voice realism
    Pricier than commodity TTS at scale
    • tts
    • voice-cloning
    • dubbing
    • multilingual
  • View Cartesia details
    VoiceFREEMIUM

    Cartesia

    Cartesia

    Low-latency streaming text-to-speech for real-time voice.

    Streaming-first speech synthesis built around the Sonic family of state-space models. Aims at real-time agent voices where latency between turns is the product. Strong choice for sub-200ms voice loops.

    Streaming over WebSocket for fast first audio
    Long-form expressive texture trails ElevenLabs
    • tts
    • streaming
    • low-latency
    • real-time
  • View Vapi details
    VoiceFREEMIUM

    Vapi

    Vapi

    Voice agent infrastructure. Build a phone-agent in a weekend.

    Production voice-agent platform — telephony, STT, LLM, TTS, and interrupt handling stitched together so you call an endpoint and get a working phone agent. Pluggable models at every layer.

    Telephony and interrupts handled
    Per-minute costs stack across layers
    • voice-agents
    • telephony
    • phone
    • real-time
  • View Deepgram details
    VoiceFREEMIUM

    Deepgram

    Deepgram

    Production speech-to-text. The STT default for many companies.

    End-to-end speech recognition platform — real-time streaming, batch transcription, speaker diarization, and language detection. Strong on accented speech, telephony audio, and long-form recordings.

    Strong on accented/telephony audio
    API-only, no end-user app
    • stt
    • transcription
    • streaming
    • diarization