Labsco
deepgram logo

api

★ 13

by deepgram · part of deepgram/skills

Deepgram API reference for speech-to-text, text-to-speech, voice agents, audio intelligence, and account management. Use whenever building with Deepgram APIs — REST or WebSocket. Covers authentication, all endpoints, query parameters, request/response schemas, and WebSocket message formats. Reference files are organized by domain: listen (STT), speak (TTS), agent (voice agents), read (text/audio intelligence), models, projects, auth, and self-hosted.

🔥🔥🔥✓ VerifiedFreeAdvanced setup
🧩 One of 7 skills in the deepgram/skills package — works on its own, and pairs well with its siblings.

This is the playbook your agent receives when the skill activates — you don't need to read it to use the skill, but it's here to audit before installing.

Deepgram API

Build with Deepgram's speech-to-text, text-to-speech, voice agent, and audio intelligence APIs.

"Flux" names two separate products. Flux STT is conversational speech-to-text on /v2/listen (model=flux-general-en). Flux TTS is turn-based speech synthesis on /v2/speak (model=flux-{voice}-{language}). They share a name and a design philosophy — turn-aware, built for voice agents — but they are different endpoints with different models, params, and messages. When a request just says "Flux", check whether it is about transcribing audio or producing it.

How Deepgram's APIs Fit Together

                   ┌──────────────────────────────┐
                   │       api.deepgram.com        │
                   └──────────────────────────────┘
                                │
  ┌───────────┬───────────┬─────┴─────┬───────────┬───────────┐
  ▼           ▼           ▼           ▼           ▼           ▼
  /v1/listen  /v2/listen  /v1/speak   /v2/speak   /v1/read    /v1/projects/*
  Nova — STT  Flux — STT  Aura — TTS  Flux — TTS  Text AI     Management
  REST + WSS  WSS only    REST + WSS  REST + WSS  REST only   REST only

                   ┌──────────────────────────────┐
                   │      agent.deepgram.com       │
                   └──────────────────────────────┘
                                │
                                ▼
                   /v1/agent/converse
                   WebSocket only
                   audio ──▶ STT ──▶ LLM ──▶ TTS ──▶ audio
                   (Deepgram orchestrates the full pipeline)

Which API Should I Use?

Audio → text (transcription)?
├─ General-purpose transcription (captions, batch, call logs, live streams with custom turn logic)
│  └─ Nova models via /v1/listen
│     ├─ Pre-recorded file    →  REST  POST https://api.deepgram.com/v1/listen?model=nova-3
│     └─ Live stream          →  WSS   wss://api.deepgram.com/v1/listen?model=nova-3
│
└─ Conversational audio / voice-agent-style turn detection
   └─ Flux STT models via /v2/listen
      └─ Live stream          →  WSS   wss://api.deepgram.com/v2/listen?model=flux-general-en

Text → audio (speech synthesis)?
├─ General-purpose TTS (broadest voice catalog, compressed/containerized audio)
│  └─ Aura models via /v1/speak
│     ├─ One-shot             →  REST  POST https://api.deepgram.com/v1/speak?model=aura-2-thalia-en
│     └─ Low-latency stream   →  WSS   wss://api.deepgram.com/v1/speak?model=aura-2-thalia-en
│
└─ Voice-agent TTS (turn-based lifecycle, barge-in, cross-turn consistency)
   └─ Flux TTS models via /v2/speak  — model is REQUIRED, and must be flux-*
      ├─ Pre-render a block   →  REST  POST https://api.deepgram.com/v2/speak?model=flux-alexis-en
      └─ Live conversation    →  WSS   wss://api.deepgram.com/v2/speak?model=flux-alexis-en

Full conversational voice agent (audio in, audio out)?
└─ WSS wss://agent.deepgram.com/v1/agent/converse
   Deepgram handles STT + your configured LLM + TTS internally

Analyze text for insights?
└─ REST POST /v1/read
   (summaries, sentiment, topics, intents)

Speech-to-Text: Nova (/v1/listen) vs Flux STT (/v2/listen)

Both model families are actively maintained and industry-leading. They solve different problems — pick the one that matches your use case.

Nova (/v1/listen)Flux STT (/v2/listen)
Endpoint/v1/listen/v2/listen
Available modelsnova-3, nova-2, nova, enhanced, baseflux-general-en
Best forGeneral transcription — captions, subtitles, call logs, batchConversational audio — voice agents, interactive assistants, turn-taking UIs
OutputContinuous transcript streamStructured turn events + transcripts (built-in turn state machine)
Turn detectionManual (utterance_end_ms, VAD events)Built-in (EOT, eager-EOT, turn_index)
TransportsREST + WebSocketWebSocket only
Intelligence overlaysYes — summarize, sentiment, topics, intents, diarize, redact, etc.No — smaller focused param set; no smart_format / diarize / punctuate
Mid-session reconfigNo (reconnect to change)Yes (Configure message updates EOT thresholds + keyterms live)

Pick Nova (/v1/listen, model=nova-3) when:

  • Generating captions, subtitles, or transcripts for recorded media
  • Running batch transcription over files (REST)
  • You need analytics overlays (summarize, sentiment, topics, intents, diarize, redact)
  • You want WebSocket streaming with your own turn-detection logic

Pick Flux STT (/v2/listen, model=flux-general-en) when:

  • Building an interactive voice agent or assistant
  • You want end-of-turn detection handled for you
  • You need low-latency turn signals and barge-in support
  • You want to update EOT thresholds or keyterms mid-session without reconnecting

Migrating from Nova 3 to Flux STT? See the official Nova 3 → Flux migration guide.

Text-to-Speech: Aura (/v1/speak) vs Flux TTS (/v2/speak)

Both TTS families are actively maintained. /v2/speak is a new endpoint, not a replacement — /v1/speak is unchanged, and there is no aliasing, redirect, or deprecation. The families do not overlap: Aura voices are served only on /v1/speak, Flux TTS voices only on /v2/speak.

Aura (/v1/speak)Flux TTS (/v2/speak)
Endpoint/v1/speak/v2/speak
Modelsaura-2-* (en, es, de, nl, fr, it, ja), aura-*flux-{voice}-{language}, e.g. flux-alexis-en — English at launch
model paramOptional (defaults to aura-asteria-en)Required; an aura-* string is rejected
Best forBroadest voice catalog, multilingual, compressed audio, one-shot synthesisVoice agents — streaming LLM output, barge-in, multi-turn conversations
Mental modelText buffer → audio streamStreaming-first, turn-based conversation
Turn lifecycleNoneSpeechStarted → audio → Flushed → SpeechMetadata per turn (server-assigned speech_id)
Cross-turn contextNone (reconnect to reset)Prosody persists across turns automatically — no API surface
TransportsREST + WebSocketREST (batch) + WebSocket (streaming)
Streaming encodingslinear16, mulaw, alawlinear16, mulaw, alaw — raw audio only
Batch encodingsmp3, opus, flac, aac, linear16, mulaw, alaw + container / bit_rateSame — but batch-only; the socket rejects them
InterruptionClear discards the buffer, no feedbackInterrupt → SpeechInterrupted with text_spoken / text_remaining
Mid-stream reconfigNo (fixed at connection)Yes — Configure updates speed only
speed0.7–1.5 — Aura-2, English and Spanish onlySeven values, 0.85–1.15 in 0.05 steps
expressivityNot supported-2…2, default 0 (beta; fixed for the connection)
Voice Agent provider.versionv1 (the default when a provider is specified)v2 (required)

Pick Aura (/v1/speak) when:

  • You need a language other than English, or a specific Aura voice
  • You want compressed or containerized output (mp3, opus, flac, aac) from a stream
  • You're doing one-shot synthesis and don't need a turn lifecycle
  • You're already on Aura and nothing in Flux TTS is pulling you over — v1 is unchanged

Pick Flux TTS (/v2/speak) when:

  • Building a voice agent, phone assistant, or customer-service bot
  • You're streaming LLM tokens to a speaker in real time and want the lowest time-to-first-audio
  • The user may barge in mid-response and you need to know what they actually heard
  • You want tone to carry across turns without managing state yourself
  • You're pre-rendering fixed audio (IVR prompts, notifications) with a Flux TTS voice — use the batch REST transport

Migrating from Aura? See the official Migrating from Aura to Flux TTS guide and Batch vs Streaming.

API Domains

DomainRESTWebSocketReference
Listen v1 — STT, Nova modelsPOST /v1/listenwss://api.deepgram.com/v1/listenlisten.md
Listen v2 — STT, Flux STT (conversational)—wss://api.deepgram.com/v2/listenlisten.md
Speak v1 — TTS, Aura modelsPOST /v1/speakwss://api.deepgram.com/v1/speakspeak.md
Speak v2 — TTS, Flux TTS (turn-based)POST /v2/speakwss://api.deepgram.com/v2/speakspeak.md
Voice AgentGET /v1/agent/settings/think/modelswss://agent.deepgram.com/v1/agent/converseagent.md
Read (Intelligence)POST /v1/read—read.md
ModelsGET /v1/models—models.md
Projects/v1/projects/*—projects.md
AuthPOST /v1/auth/grant—auth.md
Self-Hosted/v1/projects/*/selfhosted/*—self-hosted.md

SDK-Specific Skills

This api skill covers the product contracts (endpoints, query params, message shapes) that are identical across SDKs. For language-idiomatic code — imports, async patterns, builder APIs, common errors — install the SDK-specific skills. Each Deepgram SDK publishes 7 product skills named deepgram-{lang}-{product} (e.g. deepgram-python-speech-to-text, deepgram-js-voice-agent) plus a maintainer skill deepgram-{lang}-maintaining-sdk. The deepgram-{lang}- prefix avoids collisions when you install skills from multiple SDKs.

# Install all skills from a specific SDK
npx skills add deepgram/deepgram-python-sdk     # Python
npx skills add deepgram/deepgram-js-sdk         # JavaScript / TypeScript
npx skills add deepgram/deepgram-java-sdk       # Java
npx skills add deepgram/deepgram-go-sdk         # Go
npx skills add deepgram/deepgram-rust-sdk       # Rust
npx skills add deepgram/deepgram-swift-sdk      # Swift
npx skills add deepgram/deepgram-kotlin-sdk     # Kotlin
npx skills add deepgram/deepgram-dotnet-sdk     # C# / .NET
npx skills add deepgram/deepgram-browser-sdk    # Browser TypeScript

# Or install a specific product skill from one SDK (note the deepgram-{lang}- prefix)
npx skills add deepgram/deepgram-python-sdk --skill deepgram-python-speech-to-text
npx skills add deepgram/deepgram-js-sdk     --skill deepgram-js-voice-agent
SkillPurpose
recipesMinimal runnable snippets per feature per language
examplesFull integration examples with third-party platforms (Twilio, LiveKit, etc.)
startersRunnable starter apps (framework × feature matrix)
docsNavigate Deepgram documentation
setup-mcpInstall the Deepgram MCP server

Documentation