Overview
mod_realtime_ai is a native FreeSWITCH module for realtime speech-to-speech AI integrations. It connects FreeSWITCH calls directly to supported realtime AI providers while handling bidirectional audio streaming, playback and audio resampling inside the module.
Provider connections are selected through named profiles. A profile defines the provider, model and provider-native configuration file, allowing different AI configurations to be selected on a per-call basis.
Installation
Install the package built for your FreeSWITCH platform, then make sure
mod_realtime_ai is loaded by FreeSWITCH.
Configuration
Create the FreeSWITCH configuration file at
conf/autoload_configs/realtime_ai.conf.xml.
Profiles are defined inside the <profiles> section.
<configuration name="realtime_ai.conf" description="Realtime AI">
<profiles>
<profile name="sales">
<param name="provider" value="openai"/>
<param name="api-key" value=""/>
<param name="config" value="sales.json"/>
<param name="model" value="gpt-realtime-2.1"/>
</profile>
</profiles>
</configuration>
The Free Edition requires no additional license configuration. After installation, you can configure your provider profiles and start using mod_realtime_ai with the Free Edition channel limit.
Trial and licensed editions require an additional
<settings> section in
realtime_ai.conf.xml. The license information is provided
together with your trial or commercial license.
<settings>
<param name="type" value="trial"/>
<param name="issued_to" value="ACME"/>
<param name="expires_at" value="1793093900"/>
<param name="sig" value="<license-key>"/>
</settings>
Set type to trial for an evaluation license
or license for a commercial license. The remaining values,
including the license signature, are supplied with the license and should
be used exactly as provided.
Profiles
Profiles provide reusable provider configurations. Each profile has a unique name that is passed to the FreeSWITCH API when starting a realtime AI session.
The provider-specific session configuration is stored separately as JSON. The JSON uses the selected provider's native configuration format rather than a proprietary mod_realtime_ai session format.
start
Starts a realtime AI session on an existing FreeSWITCH channel using a configured profile.
The profile must exist in realtime_ai.conf.xml.
The module loads the provider, credentials, model and provider-native
configuration associated with that profile.
stop
Stops the active realtime AI session and releases its media and provider connection resources.
pause / resume
Temporarily pauses audio streaming without destroying the realtime
session. Use resume to continue streaming.
break
Interrupts the current AI playback without stopping the realtime session. This is useful for barge-in and application-controlled interruption.
update
Sends a provider-native JSON update to the active realtime session. The JSON is passed to the selected provider implementation.
This allows session configuration to be changed while the connection is active without restarting the FreeSWITCH media session.
send_text
Sends UTF-8 text to the active realtime AI provider.
The provider implementation converts the text into the appropriate provider-native realtime message.
pipe
Exposes realtime playback audio through a local named pipe (FIFO) while normal caller playback continues.
The sample rate is optional. When specified, pipe audio is resampled independently to 8, 16, 24, 32 or 48 kHz.
license
Display the current license status or generate the information required for license activation.
Displays the current license status and expiration date.
Returns the unique server identifier required for license activation. Send this identifier to AMSOFTSWITCH to request a license for the current FreeSWITCH installation.
FreeSWITCH Events
mod_realtime_ai exposes realtime session activity through FreeSWITCH custom events. Applications can subscribe to these events through ESL and react to connection state, provider readiness, errors, transcripts and audio playback.
Standard FreeSWITCH channel information is attached to every event. Event-specific data is provided in the event body.
connect
Fired when the connection to the configured realtime AI provider has been established.
{
"status": "connected"
}
ready
Fired when the provider reports that the realtime session is ready for interaction.
{
"status": "ready"
}
connect indicates that the transport connection exists.
ready indicates that the provider session itself is ready.
disconnect
Fired when the realtime provider connection closes.
Provider connection closed (1000): Normal closure
The event body contains the provider connection close code and close reason.
error
Fired when a provider or provider protocol error occurs during an active realtime session.
Provider error message
The body contains the error message reported by the provider or provider protocol implementation.
json
Exposes provider-native realtime events to the FreeSWITCH event system.
{
"type": "provider-specific-event"
}
The body is not converted into a mod_realtime_ai-specific schema. It contains the provider-native event received by the active provider implementation.
playback
Reports progress of AI audio being consumed by the FreeSWITCH playback queue.
{
"event": "chunk_played",
"seq": 1,
"size": 320,
"remaining": 4
}
{
"event": "queue_completed",
"total_chunks": 5
}
chunk_played reports playback queue progress.
queue_completed is emitted when the queued AI audio has
been consumed.
transcripts
User and assistant transcripts are exposed as separate FreeSWITCH events, making it easy for an ESL application to distinguish the two sides of the realtime conversation.
Hello, how are you?
I'm doing well, thank you. How can I help?
The event body contains the transcript text produced by the active provider.
OpenAI
Configure an OpenAI profile with your API key, realtime model and a provider-native JSON configuration file.
<profile name="sales">
<param name="provider" value="openai"/>
<param name="api-key" value="YOUR_API_KEY"/>
<param name="config" value="sales.json"/>
<param name="model" value="gpt-realtime-2.1"/>
</profile>
config is stored in
conf/realtime_ai/ and uses OpenAI's native session configuration format.
{
"type": "session.update",
"session": {
"type": "realtime",
"instructions": "You are a helpful voice assistant.",
"output_modalities": ["audio"],
"audio": {
"input": {
"format": {
"type": "audio/pcma"
},
"turn_detection": {
"type": "server_vad",
"threshold": 0.7,
"prefix_padding_ms": 300,
"silence_duration_ms": 500,
"create_response": true,
"interrupt_response": true
}
},
"output": {
"format": {
"type": "audio/pcma"
},
"voice": "marin"
}
}
}
}
Google Gemini
Gemini profiles use an API key and a provider-native JSON configuration file.
<profile name="support">
<param name="provider" value="gemini"/>
<param name="api-key" value="YOUR_API_KEY"/>
<param name="config" value="support.json"/>
</profile>
config is stored in
conf/realtime_ai/ and contains the Gemini session configuration.
{
"setup": {
"model": "models/gemini-3.8-live",
"generationConfig": {
"responseModalities": ["AUDIO"]
}
}
}
setup object when required.
ElevenLabs
ElevenLabs uses an existing Agent configuration. Specify the ElevenLabs Agent ID
through the profile's model parameter.
<profile name="elevenlabs">
<param name="provider" value="elevenlabs"/>
<param name="api-key" value="YOUR_API_KEY"/>
<param name="model" value="YOUR_AGENT_ID"/>
</profile>
model parameter contains the
ElevenLabs Agent ID. The agent itself is configured in ElevenLabs.
Azure Voice Live
Azure Voice Live requires an Azure endpoint in addition to the API key, model and provider-native JSON session configuration.
<profile name="azure">
<param name="provider" value="azure"/>
<param name="api-key" value="YOUR_API_KEY"/>
<param name="config" value="azure.json"/>
<param name="model" value="gpt-realtime"/>
<param name="endpoint" value="https://your-resource.services.ai.azure.com"/>
</profile>
endpoint identifies the Azure resource used by
the Voice Live connection.
{
"type": "session.update",
"session": {
"modalities": ["text", "audio"],
"voice": {
"type": "openai",
"name": "alloy"
},
"instructions": "You are a helpful voice assistant. Respond briefly and naturally.",
"input_audio_format": "pcm16",
"output_audio_format": "pcm16",
"input_audio_sampling_rate": 24000,
"input_audio_transcription": {
"model": "gpt-4o-mini-transcribe"
},
"turn_detection": {
"type": "azure_semantic_vad",
"threshold": 0.5,
"prefix_padding_ms": 420,
"silence_duration_ms": 500
}
}
}
Audio Formats
mod_realtime_ai handles the audio format required by each realtime AI provider independently from the codec used by the FreeSWITCH channel.
Provider audio can use PCM16, PCMU or PCMA. The provider implementation defines the required input and output format and sample rate, while the module handles the conversion required for FreeSWITCH media.
Audio settings that belong to the provider remain in the provider-native configuration. There is no separate proprietary audio configuration format introduced by mod_realtime_ai.
Resampling
Audio is automatically resampled when the FreeSWITCH channel sample rate differs from the sample rate required by the realtime AI provider.
The same applies in the opposite direction. Provider audio is decoded when necessary, converted to PCM16 internally and resampled to the FreeSWITCH channel rate before playback.
No additional resampling configuration is required by the application.
Playback
Audio returned by the realtime AI provider is processed continuously and written back to the FreeSWITCH channel while caller audio continues to be streamed to the provider.
Incoming provider audio is decoded and resampled when required, then placed into the playback queue and consumed by the channel playback path.
Playback progress is also exposed through the
mod_realtime_ai::playback FreeSWITCH event.
Interruption
Active AI playback can be interrupted without stopping the realtime session.
Interruption can originate from the provider during barge-in or can be
requested explicitly through the FreeSWITCH break API.
Queued playback audio is cleared immediately while the realtime session
remains active.
For providers that support response interruption, mod_realtime_ai also reports the amount of audio already played so the provider can keep its conversation state synchronized with what the caller actually heard.
Troubleshooting
Most integration issues can be diagnosed quickly by checking the FreeSWITCH logs together with the mod_realtime_ai custom events.
Make sure the channel has reached pre-answer or answer state before starting the realtime session. Verify that the requested profile exists and contains the required provider configuration.
A successful transport connection does not necessarily mean that the
provider has accepted the realtime session configuration. Subscribe to
mod_realtime_ai::ready,
mod_realtime_ai::error and
mod_realtime_ai::json to inspect the provider lifecycle
and provider-native events.
Check the provider-native audio configuration, including the requested input and output formats. Provider audio is decoded and resampled by mod_realtime_ai when required, but the provider configuration itself must use values supported by that provider.
Subscribe to mod_realtime_ai::error and
mod_realtime_ai::json. Provider protocol errors and
provider-native realtime events are exposed through the FreeSWITCH
event system and can also be correlated with the FreeSWITCH log.
After changing the realtime AI profile configuration, reload the FreeSWITCH XML configuration:
mod_realtime_ai reloads its profiles when FreeSWITCH processes
reloadxml. If the new profile configuration cannot be
loaded, the existing configuration remains active.