Skip to content

VoiceSesame

Sesame

Conversational voice companion chasing "voice presence."

Category
Voice
Pricing
FREE
Hosting
Cloud
Platforms
Web
Models
Self-contained (on-device)
Verified
Jun 7, 2026

A conversational-speech company building lifelike voice companions — Maya and Miles — that interrupt, self-correct, and use natural pacing. The web demo lets you talk to them in real time, and Sesame has open-sourced its underlying CSM (Conversational Speech Model) base model. Co-founded by Oculus co-creator Brendan Iribe.

Capabilities 3

What it actually does — grouped by capability family.

  • Voice agent (primary capability)
  • Speech synthesis (TTS) (secondary capability)
  • Companion chat (secondary capability)

Pros & cons

  • Open Apache-2.0 CSM-1B base model
  • Lifelike, natural conversational pacing
  • Free real-time web demo
  • Founder pedigree (Oculus co-creator)
  • Demo only; no production API yet
  • Companions not self-hostable
  • Early-stage product

Tags

View all Voice
  • View Cartesia details
    VoiceFREEMIUM

    Cartesia

    Cartesia

    Low-latency streaming text-to-speech for real-time voice.

    Streaming-first speech synthesis built around the Sonic family of state-space models. Aims at real-time agent voices where latency between turns is the product. Strong choice for sub-200ms voice loops.

    Streaming over WebSocket for fast first audio
    Long-form expressive texture trails ElevenLabs
    • tts
    • streaming
    • low-latency
    • real-time
  • View ElevenLabs details
    VoiceFREEMIUM

    ElevenLabs

    ElevenLabs

    Text-to-speech, voice cloning, and multilingual dubbing.

    Hosted speech synthesis at near-human quality — TTS, voice cloning, multilingual dubbing, and conversational voice agents. Default choice when you need a voice that sounds like a person, not a robot.

    Best-in-class voice realism
    Pricier than commodity TTS at scale
    • tts
    • voice-cloning
    • dubbing
    • multilingual
  • View Deepgram details
    VoiceFREEMIUM

    Deepgram

    Deepgram

    Production speech-to-text. The STT default for many companies.

    End-to-end speech recognition platform — real-time streaming, batch transcription, speaker diarization, and language detection. Strong on accented speech, telephony audio, and long-form recordings.

    Strong on accented/telephony audio
    API-only, no end-user app
    • stt
    • transcription
    • streaming
    • diarization
  • View Retell AI details
    VoiceFREEMIUM

    Retell AI

    Retell AI

    Build, test, and deploy AI voice agents for phone calls.

    A no-code platform for humanlike voice agents that handle inbound and outbound phone calls — receptionists, IVR, and outbound campaigns. It bundles telephony (SIP / Twilio), a proprietary turn-taking model for low-latency conversations, prompts, tools, and call analytics. Pay-as-you-go pricing with free starter credits.

    Inbound and outbound call handling
    Per-minute costs stack with LLM/TTS
    • voice-agents
    • telephony
    • call-automation
    • no-code