Build multi-modal AI applications using our new open-source Vision AI SDK.

Build a Voice-Controlled GitHub Agent in Python (MCP + Vision Agents)

4 min read
Amos G.
Amos G.
Published December 23, 2025

Build a voice assistant that can answer questions about GitHub branches, open issues, create pull requests, list contributors, etc.

Watch this YouTube Short to see how it works.

Overview of the Voice-Controlled GitHub Agent

Voice-Controlled GitHub Agent in Python diagram
  • An assistant that can read and act on any public or private repository.
  • Supports queries like finding the number of branches in a specific repo.
  • Secure GitHub access via personal access token and MCP.

You will build the AI assistant with the OpenAI Realtime API, GitHub MCP server, Straem Video calling, and the Vision Agents project.

What You Need

How To Build the GitHub Agent

Our demo uses the OpenAI Realtime API. For this reason, we do not need to create a voice pipeline for speech-to-text, LLM, text-to-speech, and turn detection. The single OpenAI model handles all these aspects. Let's initialize a new Python project and add the Vision Agents AI framework.

bash
1
2
3
4
5
uv init github-voice-agent cd github-voice-agent # Install Vision Agents and the OpenAI Python plugin uv add "vision-agents[getstream, openai]"

The Complete Sript (main.py)

python
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
import logging import os from dotenv import load_dotenv from vision_agents.core.agents import Agent, AgentLauncher from vision_agents.core import cli from vision_agents.core.mcp import MCPServerRemote from vision_agents.plugins.openai.openai_realtime import Realtime from vision_agents.plugins import getstream from vision_agents.core.events import CallSessionParticipantJoinedEvent from vision_agents.core.edge.types import User # Load environment variables from .env file load_dotenv() # Set up logging logging.basicConfig(level=logging.INFO) logger = logging.getLogger(__name__) async def create_agent(**kwargs) -> Agent: """Demonstrate OpenAI Realtime with GitHub MCP server integration.""" # Get GitHub PAT from environment github_pat = os.getenv("GITHUB_PAT") if not github_pat: logger.error("GITHUB_PAT environment variable not found!") logger.error("Please set GITHUB_PAT in your .env file or environment") raise ValueError("GITHUB_PAT environment variable not found") # Check OpenAI API key from environment openai_api_key = os.getenv("OPENAI_API_KEY") if not openai_api_key: logger.error("OPENAI_API_KEY environment variable not found!") logger.error("Please set OPENAI_API_KEY in your .env file or environment") raise ValueError("OPENAI_API_KEY environment variable not found") # Create GitHub MCP server github_server = MCPServerRemote( url="https://api.githubcopilot.com/mcp/", headers={"Authorization": f"Bearer {github_pat}"}, timeout=10.0, # Shorter connection timeout session_timeout=300.0, ) # Create OpenAI Realtime LLM (uses OPENAI_API_KEY from environment) llm = Realtime(model="gpt-4o-realtime-preview-2024-12-17") # Create real edge transport and agent user edge = getstream.Edge() agent_user = User(name="GitHub AI Assistant", id="github-agent") # Create agent with GitHub MCP server and Gemini Realtime LLM agent = Agent( edge=edge, llm=llm, agent_user=agent_user, instructions="You are a helpful AI assistant with access to GitHub via MCP server. You can help with GitHub operations like creating issues, managing pull requests, searching repositories, and more. Keep responses conversational and helpful. When you need to perform GitHub operations, use the available MCP tools.", processors=[], mcp_servers=[github_server], ) logger.info("Agent created with OpenAI Realtime and GitHub MCP server") logger.info(f"GitHub server: {github_server}") return agent async def join_call(agent: Agent, call_type: str, call_id: str, **kwargs) -> None: try: # Set up event handler for when participants join @agent.subscribe async def on_participant_joined(event: CallSessionParticipantJoinedEvent): # Check MCP tools after connection available_functions = agent.llm.get_available_functions() mcp_functions = [ f for f in available_functions if f["name"].startswith("mcp_") ] logger.info( f"✅ Found {len(mcp_functions)} MCP tools available for function calling" ) await agent.say( f"Hello {event.participant.user.name}! I'm your GitHub AI assistant powered by OpenAI Realtime. I have access to {len(mcp_functions)} GitHub tools and can help you with repositories, issues, pull requests, and more through voice commands!" ) # ensure the agent user is created await agent.create_user() # Create a call call = await agent.create_call(call_type, call_id) # Have the agent join the call/room logger.info("🎤 Agent joining call...") with await agent.join(call): logger.info( "✅ Agent is now live with OpenAI Realtime! You can talk to it in the browser." ) logger.info("Try asking:") logger.info(" - 'What repositories do I have?'") logger.info(" - 'Create a new issue in my repository'") logger.info(" - 'Search for issues with the label bug'") logger.info(" - 'Show me recent pull requests'") logger.info("") logger.info( "The agent will use OpenAI Realtime's real-time function calling to interact with GitHub!" ) # Run until the call ends await agent.finish() except Exception as e: logger.error(f"Error with OpenAI Realtime GitHub MCP demo: {e}") logger.error("Make sure your GITHUB_PAT and OPENAI_API_KEY are valid") import traceback traceback.print_exc() # Clean up await agent.close() logger.info("Demo completed!") if __name__ == "__main__": cli(AgentLauncher(create_agent=create_agent, join_call=join_call))

The script opens a browser window with a video call UI to start interacting with the GitHub voice assistant as demonstrated below.

Scaling WebRTC Video to 100,000 Participants
View Stream's latest Video API benchmark and the architecture that powers performance at scale.