Agent Skills
Instruction packs that give your AI agent know-how — some work anywhere, some only with the tool they came with.
✦ Standalone skills3,373
Self-contained. Install one into any project and it works on its own — no other software needed.
🧰 Tool add-ons952
Come bundled with a specific tool and only work together with it — they teach your agent how to operate that tool.
Voice & Video
3 standalone skillsgemini-live-api-dev
✓★ 3,787by google-gemini
Use this skill when building real-time, bidirectional streaming applications with the Gemini Live API. Covers WebSocket-based audio/video/text streaming, voice activity detection (VAD), native audio features, function calling, session management, ephemeral tokens for client-side auth, live translation, and all Live API configuration options. SDKs covered - google-genai (Python), @google/genai (JavaScript/TypeScript).
azure-speech-to-text-rest-py
✓★ 2,676by microsoft
Azure Speech to Text REST API for short audio (Python). Use for simple speech recognition of audio files up to 60 seconds without the Speech SDK. Triggers: "speech to text REST", "short audio transcription", "speech recognition REST API", "STT REST", "recognize speech REST". DO NOT USE FOR: Long audio (>60 seconds), real-time streaming, batch transcription, custom speech models, speech translation. Use Speech SDK or Batch Transcription API instead.
whisper
★ 11by firecrawl
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.