One of 43 skills in the datadog-labs/agent-skills package — works on its own, and pairs well with its siblings.
WHEN YOUR AGENT SHOULD USE IT
USE FOR
- Start monitoring an AI app or agent in Datadog.
- Add session tracking so one user's conversation can be followed across traces.
- Check and repair an existing LLM Observability setup.
DO NOT USE FOR
- Browser monitoring, APM or other Datadog products: it stays within LLM Observability.
Documents
This is the playbook your agent receives when the skill activates — you don't need to read it to use the skill, but it's here to audit before installing.
Datadog LLM Observability Instrumentation
This skill assesses the current project (backend runtime, LLM/agent framework, existing instrumentation) and routes you to the right reference file under references/ for the actual setup steps. It is fully self-contained — it does not call any Datadog MCP server. Credential provisioning uses the local CreateApiKey tool, and all code changes are made by you, directly, using your own file-editing tools.
Stay within LLM Observability scope. Do not add RUM, APM application instrumentation, or unrelated Datadog products. RUM and a Datadog Agent are relevant only because they change what session/trace linking is achievable (see the "Beyond SDK init" section in the reference files).
Do NOT invent tool names. Use only CreateApiKey as described in references/common-credentials.md; every other step is done with your normal file-editing/search tools.
Do NOT write anything to memory during or after this skill. Project paths, frameworks, credentials, and ML app names are project-specific and must not be stored in persistent memory.
Ground rules while instrumenting
- Verify before you assert. If you're not sure about file content or codebase structure, read the files — do not guess.
- Check for existing instrumentation first. Before touching the backend, check whether
ddtrace/dd-traceis already initialized with LLM Observability enabled. If it is, do not add a second, competinginit/enablecall. This only means skip re-init — it does not mean skip the work. SDK presence is not the same as a well-formed trace or a session ID that actually flows; run the session-ID plumbing audit in Phase 1d and close any gap it finds, even when the SDK is already present. - Only use packages/features you've explicitly been told to add. Don't decide on your own to enable an additional Datadog product or SDK feature beyond what this skill specifies.
- Persist added dependencies into the deploy's manifest. Any Datadog package you add (
ddtrace,dd-trace) must be written into the dependency manifest the build/deploy installs from — the one identified in Phase 1e — not just installed into the local environment. A clean deploy install reads only the manifest and will crash with a missing-module error (e.g.ModuleNotFoundError: No module named 'ddtrace') if the package isn't declared there. - Don't make stylistic changes to code you're not otherwise touching.
- No package aliases when importing Datadog packages.
- Copy Datadog SDK package names, import paths, and init keyword arguments verbatim from the applicable reference section — do not paraphrase them from memory. The Python LLM Observability import is exactly
from ddtrace.llmobs import LLMObs— the module isddtrace.llmobs, notddtrace.llm_observability(that module does not exist and produces a deploy-timeModuleNotFoundError). The Node package isdd-trace, initialized with the form that matches the project: CommonJSrequire('dd-trace').init(...), an ESMimport-based form for"type": "module"/.mjsprojects, or Next.js'sdd-trace/initialize.mjsininstrumentation.ts— all shown inreferences/llmobs-nodejs.md; never forcerequire(...)into an ESM project. Use the exactLLMObs.enable(...)/.init({...})keyword arguments shown; do not add, drop, or rename kwargs based on general Datadog knowledge. - Use a real edit tool for existing files (
Edit/Write, notsed -i,awk, or scripted find/replace via Bash). If an edit-by-text-match fails, re-read the file first rather than retrying the identical edit. - Checklist discipline. Before starting the instrumentation steps, post a short checklist of the steps you're about to take. Check items off as you go, and review the checklist at the end of the run.
- After any code change, check
package.jsonfor alint:fix,fix, orformatscript (oreslint --fix/prettier --writeconfig) and run it automatically — no need to ask permission. - Verify the app still builds/runs before declaring success (see
references/common-verify-report.md). - LLM Observability session IDs propagate from the root span only. Never pass a session ID to a child span/decorator — see the "Adding spans" section in
references/llmobs-python.md/references/llmobs-nodejs.md. - Don't promise what the environment doesn't support. If no RUM SDK is detected, don't claim RUM session linking; if no dd-agent is detected, don't claim navigable APM trace linking. State the gap and the upgrade path instead (see the "Beyond SDK init" section in the reference files).
Phase 1: Analysis
Inspect the relevant application directory before asking questions or editing files.
1a. Detecting the backend LLM/agent runtime
- Python signals:
requirements.txt,pyproject.toml,Pipfile, or*.pyfiles → runtimepython - Node.js signals: a backend entry point (
server.js,index.js, an Express/Next.js/Fastify app) → runtimenodejs - If both exist (e.g. a Next.js app with a Python worker), instrument each backend runtime independently.
- If neither is present, there is nothing to instrument — stop and tell the user.
1b. Detecting the LLM framework/SDK (optional, informational)
Check dependency files for signals of: openai, @anthropic-ai/sdk / anthropic, langchain, langgraph, ai (Vercel AI SDK), boto3 + Bedrock usage, google-generativeai / google-genai, crewai, litellm, pydantic-ai, an MCP SDK, google-adk. This doesn't change the init code (ddtrace/dd-trace auto-instruments these SDKs once the tracer is initialized) — it's only used for confirming the setup with the user and for special-cased frameworks noted in the reference files (Next.js, Vercel AI SDK).
1c. Detecting the backend application framework
- Python:
fastapi,flask, ordjangodependency - Node.js:
expressdependency, ornext(Next.js API routes / server actions)
1d. Detecting existing instrumentation, and auditing session-ID plumbing
- Grep for
ddtraceinit (LLMObs.enable(,ddtrace-run) in Python, ordd-traceinit (require('dd-trace').init(,dd-trace/initialize) in Node.js. - Also grep for a frontend RUM SDK (
datadogRum.init(,@datadog/browser-rum) — not to set it up, but because its presence determines whether the opt-in RUM↔LLMObs pivot is available (that pivot reuses the RUM session ID as the LLMObs session ID). - If LLMObs init is already present, do not add a second
enable/initcall — but do not stop there. SDK presence only means "don't re-init"; it says nothing about whether a session ID actually flows. Run this audit whenever its prerequisite surfaces are present:- Session-ID plumbing — applicable whenever Phase 1a found an LLM/agent backend. Check whether a stable, conversation/operation-scoped session ID flows from the appropriate source: for web apps, the frontend sends a per-conversation ID and the backend request model/route reads it; for CLI/background jobs, the backend reuses an existing job/request/task ID or mints one UUID per invocation. A root
agent/workflowspan must then set it assession_id/sessionId(see "Beyond SDK init" and "Session ID intake by environment" inreferences/llmobs-python.md/references/llmobs-nodejs.md).LLMObs.enable()/dd-trace().init()alone does not establish this. If a RUM SDK is also present, the RUM↔LLMObs pivot is available as an opt-in — reusing the RUM session ID as thesession_id— but that is a deliberate trade-off (the LLMObs session then spans the whole browser session), not the default; don't flag its absence as a gap. - If the check finds a gap, tell the user what's missing and fix it (same reference files, same ground rules) even though the SDK itself doesn't need re-initializing. If it passes, say so explicitly. Note the RUM pivot as not applicable rather than a gap when there's no RUM SDK.
- Session-ID plumbing — applicable whenever Phase 1a found an LLM/agent backend. Check whether a stable, conversation/operation-scoped session ID flows from the appropriate source: for web apps, the frontend sends a per-conversation ID and the backend request model/route reads it; for CLI/background jobs, the backend reuses an existing job/request/task ID or mints one UUID per invocation. A root
1e. Detecting the dependency manifest to persist into
For each backend runtime found in 1a, identify the dependency manifest the build/deploy actually installs from. This is where any Datadog package you add (ddtrace/dd-trace) must be persisted so a clean deploy install includes it. It is not necessarily "whichever manifest file happens to exist": a repo can have several (e.g. an empty requirements.txt alongside a pyproject.toml, or multiple package.json files where only one is the deployed workspace), and editing a non-authoritative one is a silent no-op at deploy time.
Resolve it in this order:
- Deploy/build install command first (authoritative). Read the project's build/deploy configuration and use the install source it names:
render.yaml(buildCommand),Procfile,Dockerfile(RUN … install …),Makefile, CI workflows,package.jsonscripts. Examples:pip install -r <file>→ that requirements file;poetry install/uv sync/pdm install→pyproject.toml(+ its lockfile);pipenv install→Pipfile;npm ci/yarn install/pnpm i→package.json. - Else, the highest-priority manifest present. Python: a lockfile-backed manager (
uv.lockorpoetry.lock→pyproject.toml) >pyproject.toml[project.dependencies]/[tool.poetry.dependencies]>requirements*.txt>Pipfile>setup.py/setup.cfg. Node.js:package.json(always).
Record the manifest path and the manager that owns it. Persisting commands (poetry add, uv add, pdm add, pipenv install, npm/yarn/pnpm/bun add) write the manifest. Non-persisting commands (pip install, uv pip install, conda install) touch only the current environment — with those you must also hand-edit the manifest. If requirements.txt is generated from a requirements.in (pip-tools), edit the .in and recompile. After hand-editing a manifest that has a lockfile, regenerate the lock so a frozen deploy install picks up the new package.
Phase 2: Instrumentation — routing table
Provision DD_API_KEY via references/common-credentials.md, then follow the reference for each backend runtime found in Phase 1a:
| Need | Read |
|---|---|
| LLM Observability — Python (SDK init, spans, session ID, RUM/APM linking) | references/llmobs-python.md |
| LLM Observability — Node.js / Next.js (SDK init, spans, session ID, RUM/APM linking) | references/llmobs-nodejs.md |
If Phase 1d found existing LLMObs instrumentation on a runtime, skip only that runtime's init/enable call and credential provisioning — still act on any gap the Phase 1d session-ID plumbing audit found, using the same reference files.
Phase 3: Verification and Reporting
See references/common-verify-report.md for the verify step and the JSON report shape.
Installation
npx skills add datadog-labs/agent-skills --skill "dd-instrument-llmo" --full-depthRun this in your project — your agent picks the skill up automatically.
BEFORE IT WILL WORK
1 FOR YOU- 01A Datadog API key
A Datadog API key. If the skill can't create one for you, it asks you to make one on Datadog's API keys page and paste it in; it saves the key to the project's .env file.
License
Licensed under MIT— you can use, modify, and redistribute it under that license's terms.
View the full license file on GitHub →