langchain-ai skills-benchmarks
COMMUNITYLABSCO SUMMARY
Of the 21 skills our catalog counts, 16 are real LangChain material: three primers that pick a framework before any code is written (framework-selection, ecosystem-primer, langchain-oss-primer), core guides to LangChain, LangGraph, and Deep Agents (langchain-fundamentals, langchain-middleware, langchain-rag, langgraph-fundamentals, langgraph-human-in-the-loop, langgraph-persistence, deep-agents-core, deep-agents-memory, deep-agents-orchestration, langchain-dependencies), and three LangSmith skills for tracing, datasets, and evaluation. The other five — api-docs, database-migrations, docker-patterns, react-components, and testing-patterns — have nothing to do with LangChain and read like generic engineering references; the repository's own structure names a skills/noise/ directory of "distractor skills for interference tests," which is exactly what these five resemble.
Actually running the benchmark needs Docker, the Claude Code CLI, and three separate API keys (OpenAI, LangSmith, Anthropic) to execute its pytest or vitest task suite against a LangSmith-tracked experiment. The README gives no lighter path and no instructions for using a single skill on its own — someone who just wants the LangChain skill set would have to copy the files out of skills/main/ themselves rather than follow anything the repo documents.
READ THE FULL ANALYSIS
What actually needs a credential. Three of the sixteen LangChain skills (langchain-oss-primer, ecosystem-primer, langchain-dependencies) reference an ANTHROPIC_API_KEY, and the three LangSmith skills each need a LANGSMITH_API_KEY. One skill, react-components, is flagged as needing a local tool; the other fourteen run with nothing configured.
We can't confirm the main/benchmark split from our own data. The README describes three source folders — main (production), benchmarks (variations used only inside treatments), and noise (distractors) — but our catalog does not record which folder each of the 21 comes from beyond what the name itself suggests; if a skill from skills/benchmarks/ made it into this list, it would appear here indistinguishable from a real one.
ALSO IN THIS PACKAGE
langsmith-tracing
Traces Claude Code conversations to LangSmith, including subagent and tool executions
WHAT'S INSIDE
21 showing · 21 totalNothing else to set up — install it and go.
api-docs
The conventions to follow when building a web service other programs call — how to name things, which response codes to send back, and how to publish a machine-readable description of it.
database-migrations
Changing the structure of a live database without breaking it: every change is a numbered file that can also be undone, and a change already applied is never edited, only followed by another.
deep-agents-core
Building an assistant on Deep Agents, LangChain's kit where the planning, the file handling, the helper agents and the approval checks are already switched on and your job is to configure them.
deep-agents-memory
An agent's files can vanish with the conversation, follow the user into the next one, or sit on your real hard drive — this is how you choose, and how to mix them folder by folder.
deep-agents-orchestration
Splitting a big job up: the agent writes itself a to-do list, sends parts of the work to purpose-built helper agents, and waits for a person to say yes before anything it should not do alone.
docker-patterns
Packaging an app into a Docker container properly: build it in two stages so the finished image stays small, do not run it as the all-powerful root user, and add a check that reports whether it is still alive.
ecosystem-primer
Before any agent code is written, this settles which of LangChain's three layers the project belongs on, gets the environment ready, and names the next thing to read.
framework-selection
LangChain, LangGraph and Deep Agents sit on top of one another rather than compete; four questions, answered in order, decide which of the three a project should start from.
langchain-dependencies
Getting the shopping list right before anything is installed: the LangChain, LangGraph and Deep Agents packages a project actually needs, the versions that work together, and the old ones to stop importing.
langchain-fundamentals
A first LangChain agent end to end: it calls functions you wrote yourself, keeps track of what was said earlier in the conversation, and can hand back an answer in a fixed shape instead of loose prose.
langchain-middleware
Nothing risky happens without a person seeing it first — the agent stops with its proposed action on show, waits for a yes, a correction or a no, and then carries on from there.
langchain-oss-primer
The first stop on a LangChain project, settling three things before a line of code is written: which of the three frameworks to build on, which well-worn shape of agent to copy, and what to install.
langchain-rag
An assistant that answers out of your own files rather than out of what the model happens to know — the handful of relevant paragraphs is found and passed to it with every question.
langgraph-fundamentals
For when you want to fix the order of the steps yourself instead of letting the model improvise: the work is drawn out as a map of steps, the paths between them, the branches and the parts that run side by side.
langgraph-human-in-the-loop
A half-finished job can sit and wait for a person: it shows what it is about to do, holds there as long as it takes, and resumes with their answer — and there is a plan for when a step fails instead.
langgraph-persistence
Each conversation keeps its own saved history, so a run can be continued tomorrow or rewound to an earlier step, while a separate shared store holds the things every conversation should know.
langsmith-dataset
A test set for an AI agent is a list of inputs paired with the answers you expect back. This builds and looks after those lists in LangSmith, whether the examples are written by hand or lifted from runs that already happened.
langsmith-evaluator
Deciding whether an AI agent's answer was any good is the hard half of testing one. This is the three pieces that make that call: the grader — another model, or code you write yourself — the wiring that captures what the agent actually did, and the run itself, over a saved test set or over live traffic.
langsmith-trace
Getting an app to write down what it and its models did is one half of the job; digging back through those records afterwards is the other. Both are here, including the setup for the many frameworks that are not LangChain, where the wiring is rarely the same twice.
react-components
The current way to write a piece of a React interface: a plain function using hooks, any logic you repeat lifted out into a hook of your own, and nested parts handed in as children instead of threaded down through a chain of props.
testing-patterns
Every test written in the same three beats — set up the situation, do the one thing, check what came back — with anything outside your own code faked, and the setup that repeats kept in one place.
HOW TO GET IT
npx skills add langchain-ai/skills-benchmarksnpx skills add langchain-ai/skills-benchmarks --skill <name> --full-depthPick the skill name from the Skills tab — each entry there installs independently.