Build multi-modal AI applications using our new open-source Vision AI SDK.

ElevenLabs with Vision Agents: Add Text-to-Speech in a Few Lines of Code

2 min read

Create a voice agent in Python that users can interect with like a real person.

Stefan B.
Stefan B.
Published February 24, 2026

ElevenLabs delivers some of the most lifelike and expressive text-to-speech experiences in many of the consumer and enterprise speech-enabled applications.

Its natural intonation, emotion, and multilingual support helps developers build voice apps the give human-level responses.

Let's build a sample demo agent that welcomes new participants the moment they join a live video conferencing/meeting, cracks programming jokes, and keeps the human-agent to-and-fro conversation going seamlessly.

In This Guide:

  • You will create an agent that speaks with ultra-realistic, emotionally nuanced voices from ElevenLabs.
  • You will discover the voice model's customization options.
  • Teams will find out how text-to-speech fits into existing agentic workflows.

Voice Pipeline Components

  • ElevenLabs: Voice generation AI model.
  • An LLM from any AI provider OpenAI, Gemini, or xAI.
  • A dictation model from Deepgram, ElevenLabs or any open-sorce alternative.
  • Smart-Turn: An open-source turn detection solution.
  • Stream Video: For realtime audio and video communication
  • Vision Agents: An open-source project for building agents.

Store Your Credentials

Add the following API credentials to your Python project's .env.

bash
1
2
3
4
5
ELEVENLABS_API_KEY=... STREAM_API_KEY=... STREAM_API_SECRET=... DEEPGRAM_API_KEY=... OPENAI_API_KEY=...

1: Add Project Dependencies

In your Python project, run these commands to integrate the framework and its dependencies for the context of the demo.

bash
1
2
uv add vision-agents uv add "vision-agents[getstream, elevenlabs, gemini, deepgram, smart-turn]"

2: Build the Voice Agent that Joins a Live Video Call

python
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
from typing import Any, Dict from dotenv import load_dotenv from vision_agents.core import Agent, AgentLauncher, Runner, User from vision_agents.core.utils.examples import get_weather_by_location from vision_agents.plugins import ( deepgram, elevenlabs, gemini, getstream, ) async def create_agent(**kwargs) -> Agent: llm = gemini.LLM("gemini-2.5-flash-lite") tts = elevenlabs.TTS(api_key=os.getenv("ELEVENLABS_API_KEY")) deepgram.STT() agent = Agent( edge=getstream.Edge(), agent_user=User(name="My happy AI friend", id="agent"), instructions="You are a friendly, witty assistant. Tell jokes when asked and greet warmly.", llm=llm, tts=tts, stt=stt ) @agent.subscribe("participant_joined") async def greet_new_participant(event): await agent.say( "Hello! Thanks for joining the call. " "I'm glad you're here—what's on your mind today?" ) return agent async def join_call(agent: Agent, call_type: str, call_id: str, **kwargs) -> None: call = await agent.create_call(call_type, call_id) # Have the agent join the call/room async with agent.join(call): # Use agent.simple response or... await agent.simple_response("tell me something interesting in a short sentence") # run till the call ends await agent.finish() if __name__ == "__main__": Runner(AgentLauncher( create_agent=create_agent, join_call=join_call )).cli()

Once you execute this script, you can listen as the agent speaks and responds to you through the ElevenLabs TTS model.

Additional Resources

Scaling WebRTC Video to 100,000 Participants
View Stream's latest Video API benchmark and the architecture that powers performance at scale.