Ear3 for developers

Build voice interviews into your product. Embed a component, receive a signed webhook, ship. Or stress-test your interview with AI personas before it ever reaches a real user.

The Ear3 ecosystem

How the voice pipeline works

Every live interview flows through the same six stages. Ear3 orchestrates them via Pipecat so latency stays under a conversational threshold end-to-end.

   Respondent (browser)

        │  WebRTC audio, ~200 ms glass-to-glass

   Daily transport ──► Deepgram STT ──► LLM (OpenAI / Gemini / Anthropic)


                              Cartesia TTS ──► WebRTC ──► Respondent
  • TransportDaily WebRTC rooms for sub-300 ms audio in and out.
  • STTDeepgram Nova-3 for streaming transcription with word timings.
  • LLM — swappable per-interview. OpenAI GPT-4-class, Anthropic Claude, or Google Gemini — configured per workspace.
  • TTSCartesia Sonic for low-latency neural voices; ElevenLabs available for expressive fallback.
  • VAD — Silero for endpoint detection so the model knows when the respondent is done speaking.
  • OrchestrationPipecat 1.3 connects the pieces and streams frames between them.

Two ways to run it

Pick your path

I want to…Start here
Ship a voice interview in my app tomorrowSDK → Quickstart
Understand the mental model (keys, sessions)SDK → Concepts
Run the voice pipeline on my own infraServer → Self-host
Call the API without an npm dependencyAPI Reference

Community

Questions, feedback, feature requests welcome. We read everything.

  • GitHubgithub.com/ear3-ai — issues, PRs, discussions
  • Discord — real-time support + roadmap conversation
  • Blogear3.ai/blog — launches + engineering notes

Built by Ear3 — voice interviews for any app.
⌘/