The caller speaks → speech becomes text → the agent decides what to say or do → the answer is spoken back in your chosen voice.
Three coordinated layers
Every production voice agent needs to listen, think, and speak in real time. ElevenLabs unifies these so you design conversations instead of stitching separate speech-to-text, language-model, and text-to-speech vendors together.
Listen
Real-time speech recognition (STT) transcribes the caller quickly and accurately so the agent can respond without awkward lag.
Think
Agent logic routes the transcript through instructions, tools, and RAG-backed knowledge to decide what to say or do.
Speak
Low-latency TTS generates natural conversational audio in a voice chosen for the agent's role, brand tone, use case, and language.
Start blank, or from a template
Blank agent
For custom workflows, unique logic, and full control over behavior and tools.
- Best for bespoke business flows
- When the assistant needs special rules
- More control, more setup
Quick-start template
For a fast start — a personal assistant or business agent you configure instead of building from zero.
- Best for prototypes and demos
- Common assistant patterns
- Configure, don't code
For a TeleTalker virtual receptionist, a business-agent template is usually the fastest route; go blank when your call flow is unusual.
Define the purpose and boundaries
The use case and industry shape the agent's persona, tone, vocabulary, and what it should prioritize. Write it in plain language: who it is, who it serves, what it must never do, and when it hands off to a human.
You are a [role] for [business or team]. Your goal is to help callers with [main tasks]. Use a [tone] style. Only answer using the approved knowledge base when accuracy matters. If asked, clearly confirm you are an AI assistant — never claim to be human. Escalate to a human when [handoff conditions].
Pick a voice that matches the job
Filter thousands of voices by language, accent, tone, and use case. A support agent may sound calm and informative; a coaching agent may sound warm and encouraging. Match the voice to your brand and the caller's expectations.
| Context | Voice character |
|---|---|
| Customer support | Calm, clear, patient |
| Sales / bookings | Warm, confident, upbeat |
| Professional services | Measured, trustworthy |
| Hospitality | Friendly, welcoming |
Add knowledge and tools
A useful agent needs both facts and actions. Knowledge grounds it; tools let it do real work.
Knowledge base (RAG)
Upload URLs, files, or text so answers come from approved information, which reduces hallucinations.
- Product facts & pricing
- Policies & hours
- Support articles & scripts
Tools & integrations
Add webhooks, client tools, or integrations so the agent acts beyond talk.
- Book or reschedule a meeting
- Fetch order or booking status
- Create a ticket
- Route to a human
action: book_meeting when to call: caller asks to schedule required fields: name, phone, preferred_time, topic success: confirm the booking was captured fallback: route to human if details are missing
Name knowledge sources clearly, keep them current, and remove outdated ones when policies or prices change — including the corrected $0.05/min rate.
Connect the agent to TeleTalker
TeleTalker is the integration surface: it receives caller audio, sends it to the agent, gets a spoken reply, and can trigger actions. The flow mirrors the listen–think–speak loop over a real GSM call.
- InboundTeleTalker captures caller audio and sends it to the STT endpoint.
- ProcessThe transcript feeds the configured agent (LLM + RAG knowledge).
- OutboundThe agent's response is spoken back through low-latency TTS.
- ActionWebhooks fire asynchronous events or hand off to a human.
| Area | What to prepare |
|---|---|
| Authentication | Use scoped API keys limited to only the endpoints the agent needs. |
| Streaming | Plan low-latency audio buffering so conversations feel natural. |
| Safety | Set credit quotas and monitor usage to avoid surprise costs. |
| Actions | Connect only the tools the agent needs, and test each with realistic calls. |
Monitor, audit, and improve
After launch, track whether the agent is useful, affordable, and responsive — then use real conversations to refine it.
Watch these metrics
- Active calls & average duration
- Total cost & credit usage
- Success rate (workflow completed)
- Response latency
Audit conversations
- Did the caller get the right answer?
- Did the agent ask needless questions?
- Did a tool fail?
- Was a response too slow?
Review metrics weekly during early launch, then settle into a steady cadence once the agent is stable. Configure a human-handoff path for repeated misunderstandings or failed tool actions.
Voice, consent & disclosure
A natural voice is a responsibility. Treat privacy and consent as part of the build, not an afterthought.
- Disclose the AI in the opening greeting, and configure the agent to confirm it's an AI if asked.
- Notice of recording plays automatically — calls are recorded and transcribed, and ElevenLabs processes call audio and transcripts.
- Consent for cloned voices — only with the owner's permission.
- Data handling — secure consent when working with recorded voices or personal data; set retention and support deletion.
"Hi, this call is handled by an AI assistant and may be recorded for service quality. How can I help you?"
Launch checklist
- Clear use case, role, tone, and audience defined for the agent.
- Boundaries set for what the agent should not answer or do.
- Voice matches the brand, task, language, and caller context.
- Consent secured for any cloned or recorded voice.
- Knowledge base current and approved; outdated sources removed.
- Webhooks & tools tested with realistic scenarios, including failure paths.
- API keys scoped to only the endpoints needed; quotas enabled.
- Latency acceptable for real-time voice.
- Human handoff configured for uncertainty, errors, and sensitive cases.
- Disclosure greeting live — AI + recording notice at the start of every call.
- Conversation review scheduled to keep improving after launch.