Labsco
google-gemini logo

gemini-live-api-dev

โœ“ Officialโ˜… 3,787

by google-gemini ยท part of google-gemini/gemini-skills

Use this skill when building real-time, bidirectional streaming applications with the Gemini Live API. Covers WebSocket-based audio/video/text streaming, voice activity detection (VAD), native audio features, function calling, session management, ephemeral tokens for client-side auth, live translation, and all Live API configuration options. SDKs covered - google-genai (Python), @google/genai (JavaScript/TypeScript).

๐Ÿ”ฅ๐Ÿ”ฅ๐Ÿ”ฅ๐Ÿ”ฅโœ“ VerifiedFreeQuick setup
๐Ÿงฉ One of 3 skills in the google-gemini/gemini-skills package โ€” works on its own, and pairs well with its siblings.

This is the playbook your agent receives when the skill activates โ€” you don't need to read it to use the skill, but it's here to audit before installing.

Gemini Live API Development Skill

Overview

The Live API enables low-latency, real-time voice and video interactions with Gemini over WebSockets. It processes continuous streams of audio, video, or text to deliver immediate, human-like spoken responses and background reasoning.

Key capabilities:

  • Bidirectional audio streaming โ€” real-time mic-to-speaker conversations
  • Background reasoning (extended thinking) โ€” multi-step background reasoning with spoken conversational fillers
  • Live streaming transcription โ€” real-time speech-to-text with interim and finalized streams
  • Video streaming โ€” send camera/screen frames alongside audio
  • Text input/output โ€” send and receive text within a live session
  • Audio transcriptions โ€” get text transcripts of both input and output audio
  • Voice Activity Detection (VAD) โ€” automatic server VAD, client-side Hybrid VAD, and manual Push-to-Talk
  • Asynchronous function calling โ€” non-blocking tool execution while audio continues streaming
  • Full-session client content โ€” inject and update conversation turns mid-stream
  • Session management โ€” context compression, session resumption, GoAway signals
  • Ephemeral tokens โ€” secure client-side authentication

[!NOTE] The Live API connects directly via WebSockets. For WebRTC support or simplified integration, use a partner integration.

Models

Current Models (Use These)

  • gemini-3.8-live โ€” Default option for most low-latency voice agent experiences and real-time dialogue without reasoning delays. Supports interleaved reasoning, asynchronous function calling by default (behavior: NON_BLOCKING), and full-session client content updates.
  • gemini-3.8-live-extended-thinking โ€” High-reasoning audio-to-audio model recommended when higher background reasoning is required during live interactions. Processes background reasoning and async tool calls (behavior: NON_BLOCKING required) while streaming continuous spoken conversational fillers; lifecycle managed via interaction_status (IN_PROGRESS vs IDLE).
  • gemini-3.5-transcribe-live โ€” Real-time streaming speech-to-text with interim hypotheses, finalized transcripts, smart formatting, and Hybrid VAD.
  • gemini-3.5-live-translate-preview โ€” Real-time speech-to-speech streaming translation across 70+ languages.

[!WARNING] Legacy Models (gemini-3.1-flash-live-preview, gemini-2.5-flash-native-audio-*, gemini-live-2.5-flash-preview, gemini-2.0-flash-live-001): Read references/migration.md for breaking protocol changes (behavior: "NON_BLOCKING", thinking_level, interaction_status, send_client_content).

SDKs

  • Python: google-genai >= 2.3.0 โ€” pip install -U google-genai
  • JavaScript/TypeScript: @google/genai >= 2.3.0 โ€” npm install @google/genai

[!WARNING] Legacy SDKs google-generativeai (Python) and @google/generative-ai (JS) are deprecated. Never use them.

Partner Integrations

To streamline real-time audio/video app development, use a third-party integration supporting the Gemini Live API over WebRTC or WebSockets:

  • LiveKit โ€” Use the Gemini Live API with LiveKit Agents.
  • Pipecat by Daily โ€” Create a real-time AI chatbot using Gemini Live and Pipecat.
  • Fishjam by Software Mansion โ€” Create live video and audio streaming applications with Fishjam.
  • Vision Agents by Stream โ€” Build real-time voice and video AI applications with Vision Agents.
  • Voximplant โ€” Connect inbound and outbound calls to Live API with Voximplant.
  • Firebase AI SDK โ€” Get started with the Gemini Live API using Firebase AI Logic.

Audio Formats

  • Input: Raw PCM, little-endian, 16-bit, mono. 16kHz native (will resample others). MIME type: audio/pcm;rate=16000
  • Output: Raw PCM, little-endian, 16-bit, mono. 24kHz sample rate.

[!IMPORTANT] Use send_realtime_input / sendRealtimeInput for all real-time streaming user input (audio, video, and text). On Gemini 3.8 models, send_client_content / sendClientContent is supported across the full session lifecycle with explicit roles (user or model) to inject conversation context (turn_complete=true unconditionally interrupts active generation).

[!WARNING] Do not use media in sendRealtimeInput. Use the specific keys: audio for audio data, video for images/video frames, and text for text input.


Background Reasoning (Extended Thinking)

Use gemini-3.8-live-extended-thinking when your voice agent must evaluate complex data, plan multiple steps, or handle long-running tools. The model speaks natural conversational fillers (e.g. "Checking flight options now...") while executing asynchronous tools in the background.

Key requirements:

  • Thinking config: Set thinking_config=types.ThinkingConfig(thinking_level="low") ("minimal" | "low" | "medium" | "high").
  • Non-blocking tools: All function declarations must set behavior="NON_BLOCKING". Synchronous blocking mode is not supported and returns an error.
  • Lifecycle tracking (interaction_status): Do not rely on turn_complete=True alone to detect turn completion. Monitor message.interaction_status (Python) / message.interactionStatus (JS):
    • "IN_PROGRESS": Server is reasoning, speaking conversational fillers, or waiting for async tool responses.
    • "IDLE": Server has completed all background reasoning and tool calls; session is ready for user input.

See references/migration.md and the Thinking in Live API Guide for complete Python and JavaScript implementation examples.


Live Translation (Gemini Live Translate)

The Live API supports real-time, low-latency streaming translation of speech (audio) across 70+ languages. For full details on options and capabilities, see the Live Translate Guide.

Model

  • gemini-3.5-live-translate-preview โ€” The recommended translation model for all Live Translate use cases.

Configuration (TranslationConfig)

To enable translation, specify a TranslationConfig object inside your live session setup:

  • Python SDK: Configure the connection using translation_config on LiveConnectConfig:
    config = types.LiveConnectConfig(
        response_modalities=[types.Modality.AUDIO],
        translation_config=types.TranslationConfig(
            target_language_code="es",  # Target language code (e.g. es, fr, pl)
            echo_target_language=True,
        ),
        input_audio_transcription=types.AudioTranscriptionConfig(),
        output_audio_transcription=types.AudioTranscriptionConfig(),
    )
  • Raw WebSockets: Place translationConfig inside generationConfig:
    {
      "setup": {
        "model": "models/gemini-3.5-live-translate-preview",
        "generationConfig": {
          "responseModalities": ["AUDIO"],
          "translationConfig": {
            "targetLanguageCode": "es",
            "echoTargetLanguage": true
          }
        }
      }
    }

Live Streaming Transcription (Gemini Live Transcribe)

The Live API supports real-time streaming speech-to-text over WebSockets with low-latency interim hypotheses, finalized transcripts, and Hybrid VAD. For full details, see the Live Transcription Guide and Colab Cookbook.

Model

  • gemini-3.5-transcribe-live

Modes

  • smart: cleans up filler words, resolves inline self-corrections, and structures formatting.
  • verbatim (default): exact word-for-word transcript.

Python

config = types.LiveConnectConfig(
    response_modalities=["TEXT"],
    input_audio_transcription=types.AudioTranscriptionConfig(),
)

async with client.aio.live.connect(model="gemini-3.5-transcribe-live", config=config) as session:
    # Stream audio
    await session.send_realtime_input(audio=types.Blob(data=chunk, mime_type="audio/pcm;rate=16000"))
    # Hybrid VAD: notify turn end on client-detected silence for zero latency
    await session.send_realtime_input(audio_stream_end=True)

JavaScript

const session = await ai.live.connect({
  model: 'gemini-3.5-transcribe-live',
  config: {
    responseModalities: ['text'],
    inputAudioTranscription: { mode: 'smart' }
  },
  callbacks: {
    onmessage: (msg) => {
      if (msg.serverContent?.interimInputTranscription) {
        console.log('Interim:', msg.serverContent.interimInputTranscription.text);
      }
      if (msg.serverContent?.inputTranscription) {
        console.log('Final:', msg.serverContent.inputTranscription.text);
      }
    }
  }
});

session.sendRealtimeInput({ audio: { data: chunkBase64, mimeType: 'audio/pcm;rate=16000' } });
session.sendRealtimeInput({ audioStreamEnd: true }); // Hybrid VAD

Raw WebSockets

{
  "setup": {
    "model": "models/gemini-3.5-transcribe-live",
    "generationConfig": {
      "responseModalities": ["TEXT"],
      "speechConfig": {
        "voiceConfig": {}
      }
    },
    "inputAudioTranscription": {
      "mode": "smart"
    }
  }
}

Upgrading & Migration

For step-by-step migration checklists and protocol deltas when upgrading from gemini-3.1-flash-live-preview, gemini-2.5-flash-native-audio-*, or gemini-2.0-flash-live-001 to Gemini 3.8 Live or Gemini 3.8 Live Extended Thinking, read references/migration.md.

Best Practices

  1. Use headphones when testing mic audio to prevent echo/self-interruption
  2. Enable context window compression for sessions longer than 15 minutes
  3. Implement session resumption to handle connection resets gracefully
  4. Use ephemeral tokens for client-side deployments โ€” never expose API keys in browsers
  5. Use send_realtime_input for real-time user input (audio, video, text). Use send_client_content with explicit user/model roles to inject context turns mid-stream
  6. Send audioStreamEnd / audio_stream_end (Hybrid VAD) when the mic is paused or user finishes speaking
  7. Clear audio playback queues on interruption signals (interrupted: true)
  8. Process all parts in each server event โ€” events can contain multiple content parts
  9. Monitor interaction_status (IN_PROGRESS vs IDLE) when using gemini-3.8-live-extended-thinking rather than relying on turn_complete alone

Documentation Lookup

When MCP is Installed (Preferred)

If the search_docs tool (from the Google MCP server) is available, use it as your only documentation source:

  1. Call search_docs with your query
  2. Read the returned documentation
  3. Trust MCP results as source of truth for API details โ€” they are always up-to-date.

[!IMPORTANT] When MCP tools are present, never fetch URLs manually. MCP provides up-to-date, indexed documentation that is more accurate and token-efficient than URL fetching.

When MCP is NOT Installed (Fallback Only)

If no MCP documentation tools are available, fetch from the official docs index:

llms.txt URL: https://ai.google.dev/gemini-api/docs/llms.txt

This index contains links to all documentation pages in .md.txt format. Use web fetch tools to:

  1. Fetch llms.txt to discover available documentation pages
  2. Fetch specific pages (e.g., https://ai.google.dev/gemini-api/docs/live-session.md.txt)

Key Documentation Pages

[!IMPORTANT] Those are not all the documentation pages. Use the llms.txt index to discover available documentation pages

Supported Languages

The Live API supports 70 languages including: English, Spanish, French, German, Italian, Portuguese, Chinese, Japanese, Korean, Hindi, Arabic, Russian, and many more. Native audio models automatically detect and switch languages.