Labsco
firecrawl logo

openai-whisper-api

β˜… 3

by firecrawl Β· part of firecrawl/openclaw

Transcribes an audio file to text through OpenAI's cloud Whisper API, using your own OpenAI API key.

πŸ”₯πŸ”₯πŸ”₯βœ“ VerifiedAccount requiredNeeds API keys
πŸ”Œ Needs an OpenAI API key and curl. Set that up first β€” this skill doesn't include it. One of 59 skills in the firecrawl/openclaw package.

WHEN YOUR AGENT SHOULD USE IT

A QUICK BOUNDARY

USE FOR

  • Transcribe an audio file into a plain-text or JSON transcript.
  • Guide a transcription with a prompt, such as listing names it might otherwise mishear.
  • Force transcription in a specific spoken language instead of relying on auto-detection.

This is the playbook your agent receives when the skill activates β€” you don't need to read it to use the skill, but it's here to audit before installing.

OpenAI Whisper API (curl)

Transcribe an audio file via OpenAI’s /v1/audio/transcriptions endpoint.

Quick start

{baseDir}/scripts/transcribe.sh /path/to/audio.m4a

Defaults:

  • Model: whisper-1
  • Output: <input>.txt

Useful flags

{baseDir}/scripts/transcribe.sh /path/to/audio.ogg --model whisper-1 --out /tmp/transcript.txt
{baseDir}/scripts/transcribe.sh /path/to/audio.m4a --language en
{baseDir}/scripts/transcribe.sh /path/to/audio.m4a --prompt "Speaker names: Peter, Daniel"
{baseDir}/scripts/transcribe.sh /path/to/audio.m4a --json --out /tmp/transcript.json

API key

Set OPENAI_API_KEY, or configure it in ~/.openclaw/openclaw.json:

{
  skills: {
    "openai-whisper-api": {
      apiKey: "OPENAI_KEY_HERE",
    },
  },
}