Intended Audience
- Developers
License
- OSI Approved :: Apache Software License
Programming Language
- Python :: 3
- Python :: 3 :: Only
- Python :: 3.10
Topic
- Multimedia :: Sound/Audio
- Multimedia :: Video
- Scientific/Engineering :: Artificial Intelligence
LiveKit Plugins: Phonic
Realtime voice AI integration for Phonic with LiveKit Agents.
Installation
uv add livekit-plugins-phonic
Usage
import asyncio
import logging
from dotenv import load_dotenv
from livekit.agents import (
Agent,
AgentServer,
AgentSession,
JobContext,
cli,
function_tool,
)
from livekit.plugins.phonic.realtime import RealtimeModel
logger = logging.getLogger("phonic-agent")
load_dotenv()
class MyAgent(Agent):
def __init__(self) -> None:
super().__init__(
instructions="You are a helpful voice AI assistant named Sabrina.",
llm=RealtimeModel(
voice="sabrina",
audio_speed=1.2,
),
)
@function_tool(
description="Toggle a light on or off. Available lights are A05, A06, A07, and A08."
)
async def toggle_light(self, light_id: str, state: str) -> str:
"""Called when the user asks to toggle a light on or off.
Args:
light_id: The ID of the light to toggle
state: Whether to turn the light on or off, e.g., 'on', 'off'
"""
logger.info(f"Turning {state} light {light_id}")
await asyncio.sleep(1.0)
return f"Light {light_id} turned {state}"
server = AgentServer()
@server.rtc_session()
async def entrypoint(ctx: JobContext):
session = AgentSession()
await session.start(agent=MyAgent(), room=ctx.room)
await session.generate_reply(
instructions="Greet the user, asking about their day.",
)
if __name__ == "__main__":
cli.run_app(server)
Reusing tools with Phonic Responses
Convert an existing LiveKit ToolContext into the schema-only definitions
accepted by Phonic's Responses API:
from livekit.plugins.phonic.realtime import to_phonic_tool_definitions
tool_definitions = to_phonic_tool_definitions(tool_context)
The executable functions remain in the ToolContext; only their names,
descriptions, and parameter schemas are returned.
Configuration
Set the PHONIC_API_KEY environment variable, or pass api_key directly to RealtimeModel. All other options are optional.
| Option | Type | Description |
|---|---|---|
api_key |
str |
Phonic API key. Falls back to PHONIC_API_KEY environment variable |
phonic_agent |
str |
Phonic agent name. Options set explicitly here override agent settings |
voice |
str |
Voice ID — see available voices |
welcome_message |
str |
Message the agent says when the conversation starts. Ignored when generate_welcome_message is True |
generate_welcome_message |
bool |
Auto-generate the welcome message (ignores welcome_message) |
project |
str |
Project name (default: main) |
default_language |
str |
ISO 639-1 default language for recognition and speech |
additional_languages |
list[str] |
Further ISO 639-1 codes (must not repeat default_language) |
multilingual_mode |
"auto" | "request" |
Per-utterance language detection vs. change on user request (recommended: request) |
audio_speed |
float |
Audio playback speed |
phonic_tools |
list[str] |
Names of Phonic-side tools available to the assistant: Webhook tools and built-in tools (choose_not_to_respond, keypad_input, natural_conversation_ending) |
boosted_keywords |
list[str] |
Keywords to boost in speech recognition |
min_words_to_interrupt |
int |
Minimum number of user words required to interrupt the assistant |
generate_no_input_poke_text |
bool |
Auto-generate poke text when user is silent |
no_input_poke_sec |
float |
Seconds of silence before sending poke message |
no_input_poke_text |
str |
Poke message text (ignored when generate_no_input_poke_text is True) |
no_input_end_conversation_sec |
float |
Seconds of silence before ending conversation |
websocket_timeout_sec |
int |
Seconds of inactivity before the Phonic websocket is closed |
intelligence_level |
"standard" | "high" |
Model intelligence level |
phonic_model |
"phonic_v0_5" | "phonic_v1" | "phonic_v1_1" |
Phonic model version to use |
is_welcome_message_interruptible |
bool |
When False, the welcome message cannot be interrupted |
vad_prebuffer_duration_ms |
int |
Voice-activity-detection prebuffer duration (ms) |
vad_min_speech_duration_ms |
int |
Minimum speech duration for VAD (ms) |
vad_min_silence_duration_ms |
int |
Minimum silence duration for VAD (ms) |
vad_threshold |
float |
Voice-activity-detection threshold |
enable_assistant_backchannel |
bool |
When True, the assistant backchannels (e.g. "mm-hmm") while the user speaks |
assistant_backchannel_aggressiveness |
float |
How aggressively the assistant backchannels (needs enable_assistant_backchannel) |
pronunciation_dictionary |
list[PronunciationEntry] |
{ word, pronunciation } entries; words must be unique |
template_variables |
dict[str, str] |
Variables substituted into the system prompt and welcome message |
enable_redaction |
bool |
Redact PII/PHI from transcripts and bleep it from audio after the conversation |
enable_watermarking |
bool |
Embed an inaudible provenance watermark in generated audio. Adds a very small amount of latency |
mcp_servers |
list[str] |
Names of pre-configured MCP servers to make available (must be unique) |
observability_integrations |
list["braintrust"] |
Observability integrations to forward traces to |
configuration_endpoint |
ConfigurationEndpoint | None |
Endpoint the agent calls to fetch per-conversation configuration |
additional_params |
dict[str, Any] |
Additional runtime parameters forwarded to Phonic |
configs_for_tools |
list[PhonicToolConfig] |
Per-tool behavior overrides (see Per-tool configuration) |
Per-tool configuration
configs_for_tools takes one entry per tool you want to customize. Each entry is keyed by the tool name; every other field is optional and falls back to the plugin default when omitted. Tools with no entry keep the defaults.
RealtimeModel(
configs_for_tools=[
{"name": "transfer_call", "forbid_speech_after_tool_call": True},
{"name": "submit_form", "forbid_tool_call_after_speech": True},
],
)
| Field | Type | Default | Description |
|---|---|---|---|
name |
str |
— | Tool this config applies to (required) |
require_speech_before_tool_call |
bool |
False |
Require the agent to speak before the tool can be called |
forbid_speech_after_tool_call |
bool |
False |
Suppress the auto-generated spoken reply after the tool. Use for tools that always hand off to another agent (a non-handoff tool set here would leave the agent silent) |
forbid_tool_call_after_speech |
bool |
False |
Drop the tool call if the agent already spoke this turn |
allow_tool_chaining |
bool |
False |
Allow another tool call immediately after this tool's output |
respond_after_sec |
float |
— | choose_not_to_respond only. Seconds to wait after the tool fires; if the user stays silent, the agent speaks a follow-up. Omit to keep the default (stay silent). |
speech_before_tool_call |
str |
— | keypad_input / natural_conversation_ending only. required | optional | suppressed. |
The plugin always sends tool calls with wait_for_speech_before_tool_call on; this is not configurable per tool.
Deprecated: the top-level
forbid_speech_after_tool_call: list[str]option still works but is deprecated — it now folds each listed tool intoconfigs_for_toolsasforbid_speech_after_tool_call=True(an explicitconfigs_for_toolsentry wins) and logs a warning. Preferconfigs_for_tools.
Built-in tools
Phonic's built-in tools — choose_not_to_respond, keypad_input, natural_conversation_ending — are enabled by listing their names in phonic_tools, alongside any Webhook tools. To configure one, add a configs_for_tools entry keyed by the same name (respond_after_sec for choose_not_to_respond; speech_before_tool_call for the other two). A built-in listed without a config uses its Phonic-side defaults.
RealtimeModel(
phonic_tools=["choose_not_to_respond", "keypad_input"],
configs_for_tools=[
{"name": "choose_not_to_respond", "respond_after_sec": 5},
],
)
If you already have an agent set up on the Phonic platform, you can use the phonic_agent option to specify the agent name. As a note, configuration options you set in the LiveKit Agents SDK will override the agent settings set on the Phonic platform. This means the system prompt you have set on the Phonic platform will be ignored in favor of the instructions field set on the LiveKit Agent. Likewise, options explicitly set in the RealtimeModel constructor will override the Phonic agent's settings.
If you have Webhook tools set up on the Phonic platform, you can use phonic_tools to make them available to your agent, together with Phonic's built-in tools. Custom function tools you define on the LiveKit Agent are also supported and run over the websocket.