DOCUMENTATION

mod_realtime_ai

Installation, configuration, FreeSWITCH API reference and provider integration guides.

GETTING STARTED

Overview

mod_realtime_ai is a native FreeSWITCH module for realtime speech-to-speech AI integrations. It connects FreeSWITCH calls directly to supported realtime AI providers while handling bidirectional audio streaming, playback and audio resampling inside the module.

Provider connections are selected through named profiles. A profile defines the provider, model and provider-native configuration file, allowing different AI configurations to be selected on a per-call basis.

FreeSWITCH mod_realtime_ai Realtime AI Provider
Supported providers: OpenAI, Google Gemini, ElevenLabs and Azure Voice Live.
GETTING STARTED

Installation

Install the package built for your FreeSWITCH platform, then make sure mod_realtime_ai is loaded by FreeSWITCH.

Prebuilt packages are available for supported Debian and EL9 platforms. Free, 30-day Trial and Commercial builds are provided separately. Detailed package names and installation commands are provided with each release. View Downloads →
GETTING STARTED

Configuration

Create the FreeSWITCH configuration file at conf/autoload_configs/realtime_ai.conf.xml. Profiles are defined inside the <profiles> section.

XML conf/autoload_configs/realtime_ai.conf.xml
<configuration name="realtime_ai.conf" description="Realtime AI">
    <profiles>
        <profile name="sales">
            <param name="provider" value="openai"/>
            <param name="api-key" value=""/>
            <param name="config" value="sales.json"/>
            <param name="model" value="gpt-realtime-2.1"/>
        </profile>
    </profiles>
</configuration>
Free, Trial and Licensed Editions

The Free Edition requires no additional license configuration. After installation, you can configure your provider profiles and start using mod_realtime_ai with the Free Edition channel limit.

Trial and licensed editions require an additional <settings> section in realtime_ai.conf.xml. The license information is provided together with your trial or commercial license.

XML conf/autoload_configs/realtime_ai.conf.xml
<settings>
  <param name="type" value="trial"/>
  <param name="issued_to" value="ACME"/>
  <param name="expires_at" value="1793093900"/>
  <param name="sig" value="<license-key>"/>
</settings>

Set type to trial for an evaluation license or license for a commercial license. The remaining values, including the license signature, are supplied with the license and should be used exactly as provided.

GETTING STARTED

Profiles

Profiles provide reusable provider configurations. Each profile has a unique name that is passed to the FreeSWITCH API when starting a realtime AI session.

The provider-specific session configuration is stored separately as JSON. The JSON uses the selected provider's native configuration format rather than a proprietary mod_realtime_ai session format.

$ uuid_realtime_ai <uuid> start sales
The same FreeSWITCH API is used regardless of which supported realtime AI provider is selected by the profile.
FREESWITCH API

start

Starts a realtime AI session on an existing FreeSWITCH channel using a configured profile.

$ uuid_realtime_ai <uuid> start <profile>
+OK Success

The profile must exist in realtime_ai.conf.xml. The module loads the provider, credentials, model and provider-native configuration associated with that profile.

FREESWITCH API

stop

Stops the active realtime AI session and releases its media and provider connection resources.

$ uuid_realtime_ai <uuid> stop
+OK Success
FREESWITCH API

pause / resume

Temporarily pauses audio streaming without destroying the realtime session. Use resume to continue streaming.

$ uuid_realtime_ai <uuid> pause
+OK Success
$ uuid_realtime_ai <uuid> resume
+OK Success
FREESWITCH API

break

Interrupts the current AI playback without stopping the realtime session. This is useful for barge-in and application-controlled interruption.

$ uuid_realtime_ai <uuid> break
+OK Success
FREESWITCH API

update

Sends a provider-native JSON update to the active realtime session. The JSON is passed to the selected provider implementation.

$ uuid_realtime_ai <uuid> update '<json>'

This allows session configuration to be changed while the connection is active without restarting the FreeSWITCH media session.

FREESWITCH API

send_text

Sends UTF-8 text to the active realtime AI provider.

$ uuid_realtime_ai <uuid> send_text '<text>'

The provider implementation converts the text into the appropriate provider-native realtime message.

FREESWITCH API

pipe

Exposes realtime playback audio through a local named pipe (FIFO) while normal caller playback continues.

$ uuid_realtime_ai <uuid> pipe start <path> [8000|16000|24000|32000|48000]
$ uuid_realtime_ai <uuid> pipe stop

The sample rate is optional. When specified, pipe audio is resampled independently to 8, 16, 24, 32 or 48 kHz.

FREESWITCH API

license

Display the current license status or generate the information required for license activation.

LICENSE STATUS FreeSWITCH CLI
$ uuid_realtime_ai license status
+OK license OK. expires_at=2026-10-27 09:38:20 UTC

Displays the current license status and expiration date.

LICENSE REQUEST FreeSWITCH CLI
$ uuid_realtime_ai license request
+OK server_id=<server-id>

Returns the unique server identifier required for license activation. Send this identifier to AMSOFTSWITCH to request a license for the current FreeSWITCH installation.

EVENTS

FreeSWITCH Events

mod_realtime_ai exposes realtime session activity through FreeSWITCH custom events. Applications can subscribe to these events through ESL and react to connection state, provider readiness, errors, transcripts and audio playback.

$ event plain CUSTOM mod_realtime_ai::connect mod_realtime_ai::ready mod_realtime_ai::disconnect mod_realtime_ai::error mod_realtime_ai::json mod_realtime_ai::playback mod_realtime_ai::transcript_user mod_realtime_ai::transcript_assistant

Standard FreeSWITCH channel information is attached to every event. Event-specific data is provided in the event body.

EVENT

connect

Fired when the connection to the configured realtime AI provider has been established.

EVENT mod_realtime_ai::connect
{
  "status": "connected"
}
EVENT

ready

Fired when the provider reports that the realtime session is ready for interaction.

EVENT mod_realtime_ai::ready
{
  "status": "ready"
}

connect indicates that the transport connection exists. ready indicates that the provider session itself is ready.

EVENT

disconnect

Fired when the realtime provider connection closes.

EVENT mod_realtime_ai::disconnect
Provider connection closed (1000): Normal closure

The event body contains the provider connection close code and close reason.

EVENT

error

Fired when a provider or provider protocol error occurs during an active realtime session.

EVENT mod_realtime_ai::error
Provider error message

The body contains the error message reported by the provider or provider protocol implementation.

EVENT

json

Exposes provider-native realtime events to the FreeSWITCH event system.

EVENT mod_realtime_ai::json
{
  "type": "provider-specific-event"
}

The body is not converted into a mod_realtime_ai-specific schema. It contains the provider-native event received by the active provider implementation.

EVENT

playback

Reports progress of AI audio being consumed by the FreeSWITCH playback queue.

CHUNK PLAYED mod_realtime_ai::playback
{
  "event": "chunk_played",
  "seq": 1,
  "size": 320,
  "remaining": 4
}
QUEUE COMPLETED mod_realtime_ai::playback
{
  "event": "queue_completed",
  "total_chunks": 5
}

chunk_played reports playback queue progress. queue_completed is emitted when the queued AI audio has been consumed.

EVENTS

transcripts

User and assistant transcripts are exposed as separate FreeSWITCH events, making it easy for an ESL application to distinguish the two sides of the realtime conversation.

USER mod_realtime_ai::transcript_user
Hello, how are you?
ASSISTANT mod_realtime_ai::transcript_assistant
I'm doing well, thank you. How can I help?

The event body contains the transcript text produced by the active provider.

PROVIDERS

OpenAI

Configure an OpenAI profile with your API key, realtime model and a provider-native JSON configuration file.

XML OpenAI profile
<profile name="sales">
    <param name="provider" value="openai"/>
    <param name="api-key" value="YOUR_API_KEY"/>
    <param name="config" value="sales.json"/>
    <param name="model" value="gpt-realtime-2.1"/>
</profile>
The file referenced by config is stored in conf/realtime_ai/ and uses OpenAI's native session configuration format.
JSON conf/realtime_ai/sales.json
{
  "type": "session.update",
  "session": {
    "type": "realtime",
    "instructions": "You are a helpful voice assistant.",
    "output_modalities": ["audio"],
    "audio": {
      "input": {
        "format": {
          "type": "audio/pcma"
        },
        "turn_detection": {
          "type": "server_vad",
          "threshold": 0.7,
          "prefix_padding_ms": 300,
          "silence_duration_ms": 500,
          "create_response": true,
          "interrupt_response": true
        }
      },
      "output": {
        "format": {
          "type": "audio/pcma"
        },
        "voice": "marin"
      }
    }
  }
}
PROVIDERS

Google Gemini

Gemini profiles use an API key and a provider-native JSON configuration file.

XML Google Gemini profile
<profile name="support">
    <param name="provider" value="gemini"/>
    <param name="api-key" value="YOUR_API_KEY"/>
    <param name="config" value="support.json"/>
</profile>
The file referenced by config is stored in conf/realtime_ai/ and contains the Gemini session configuration.
JSON conf/realtime_ai/support.json
{
  "setup": {
    "model": "models/gemini-3.8-live",
    "generationConfig": {
      "responseModalities": ["AUDIO"]
    }
  }
}
Additional Gemini-native options, including tools and function declarations, can be added to the same setup object when required.
PROVIDERS

ElevenLabs

ElevenLabs uses an existing Agent configuration. Specify the ElevenLabs Agent ID through the profile's model parameter.

XML ElevenLabs profile
<profile name="elevenlabs">
    <param name="provider" value="elevenlabs"/>
    <param name="api-key" value="YOUR_API_KEY"/>
    <param name="model" value="YOUR_AGENT_ID"/>
</profile>
ElevenLabs: the model parameter contains the ElevenLabs Agent ID. The agent itself is configured in ElevenLabs.
PROVIDERS

Azure Voice Live

Azure Voice Live requires an Azure endpoint in addition to the API key, model and provider-native JSON session configuration.

XML Azure Voice Live profile
<profile name="azure">
    <param name="provider" value="azure"/>
    <param name="api-key" value="YOUR_API_KEY"/>
    <param name="config" value="azure.json"/>
    <param name="model" value="gpt-realtime"/>
    <param name="endpoint" value="https://your-resource.services.ai.azure.com"/>
</profile>
Azure: endpoint identifies the Azure resource used by the Voice Live connection.
JSON conf/realtime_ai/azure.json
{
  "type": "session.update",
  "session": {
    "modalities": ["text", "audio"],
    "voice": {
      "type": "openai",
      "name": "alloy"
    },
    "instructions": "You are a helpful voice assistant. Respond briefly and naturally.",
    "input_audio_format": "pcm16",
    "output_audio_format": "pcm16",
    "input_audio_sampling_rate": 24000,
    "input_audio_transcription": {
      "model": "gpt-4o-mini-transcribe"
    },
    "turn_detection": {
      "type": "azure_semantic_vad",
      "threshold": 0.5,
      "prefix_padding_ms": 420,
      "silence_duration_ms": 500
    }
  }
}
This example uses 24 kHz PCM audio. The tested Azure configuration can also use provider-native G.711/8 kHz settings when required by the deployment.
AUDIO

Audio Formats

mod_realtime_ai handles the audio format required by each realtime AI provider independently from the codec used by the FreeSWITCH channel.

Provider audio can use PCM16, PCMU or PCMA. The provider implementation defines the required input and output format and sample rate, while the module handles the conversion required for FreeSWITCH media.

FreeSWITCH Channel mod_realtime_ai Provider Audio

Audio settings that belong to the provider remain in the provider-native configuration. There is no separate proprietary audio configuration format introduced by mod_realtime_ai.

AUDIO

Resampling

Audio is automatically resampled when the FreeSWITCH channel sample rate differs from the sample rate required by the realtime AI provider.

The same applies in the opposite direction. Provider audio is decoded when necessary, converted to PCM16 internally and resampled to the FreeSWITCH channel rate before playback.

Channel Rate Automatic Resampling Provider Rate

No additional resampling configuration is required by the application.

AUDIO

Playback

Audio returned by the realtime AI provider is processed continuously and written back to the FreeSWITCH channel while caller audio continues to be streamed to the provider.

Incoming provider audio is decoded and resampled when required, then placed into the playback queue and consumed by the channel playback path.

AI Provider Decode / Resample Playback Queue Caller

Playback progress is also exposed through the mod_realtime_ai::playback FreeSWITCH event.

AUDIO

Interruption

Active AI playback can be interrupted without stopping the realtime session.

Interruption can originate from the provider during barge-in or can be requested explicitly through the FreeSWITCH break API. Queued playback audio is cleared immediately while the realtime session remains active.

$ uuid_realtime_ai <uuid> break
+OK Success

For providers that support response interruption, mod_realtime_ai also reports the amount of audio already played so the provider can keep its conversation state synchronized with what the caller actually heard.

HELP

Troubleshooting

Most integration issues can be diagnosed quickly by checking the FreeSWITCH logs together with the mod_realtime_ai custom events.

Session does not start

Make sure the channel has reached pre-answer or answer state before starting the realtime session. Verify that the requested profile exists and contains the required provider configuration.

$ uuid_realtime_ai <uuid> start <profile>
Provider connects but the session is not ready

A successful transport connection does not necessarily mean that the provider has accepted the realtime session configuration. Subscribe to mod_realtime_ai::ready, mod_realtime_ai::error and mod_realtime_ai::json to inspect the provider lifecycle and provider-native events.

No audio from the AI provider

Check the provider-native audio configuration, including the requested input and output formats. Provider audio is decoded and resampled by mod_realtime_ai when required, but the provider configuration itself must use values supported by that provider.

Provider returns an error

Subscribe to mod_realtime_ai::error and mod_realtime_ai::json. Provider protocol errors and provider-native realtime events are exposed through the FreeSWITCH event system and can also be correlated with the FreeSWITCH log.

Profile changes are not applied

After changing the realtime AI profile configuration, reload the FreeSWITCH XML configuration:

$ reloadxml

mod_realtime_ai reloads its profiles when FreeSWITCH processes reloadxml. If the new profile configuration cannot be loaded, the existing configuration remains active.