LabscoConnect MCP ↗
PLUGIN / ADDYOSMANICOMMUNITY

Agent Skills

Addy Osmani's engineering workflow for coding assistants: 24 skills that take a change from spec and plan through code, tests and review to release, asking for proof at each step.

What gets installed: 24 skills, 9 commands, 4 agents

Skills

Working out what to build 4

Before any code: an interview that asks one question at a time, a way to turn a rough idea into a concrete proposal, a written spec, and a quality bar the project agrees to keep.

interview-me

Interviews you one question at a time, each with the assistant's own guess attached, until it can predict your next three answers, which it counts as about 95% confidence. You get back a short statement of intent covering outcome, user, why now, success, the binding constraint and what is out of scope, and nothing is planned or built until you give an explicit yes.

idea-refine

Turns a half-formed idea into a one-page brief worth building. Along the way it asks three to five sharpening questions, offers five to eight variations, narrows them to two or three directions it stress-tests honestly, and ends with a Not Doing list that spells out the trade-offs; the brief is saved to docs/ideas only once you confirm.

spec-driven-development

Writes a specification before any code, covering the objective, the exact commands, project structure, code style, testing strategy, and boundaries for what to always do, ask about first or never do. Work then moves from spec to plan to tasks to code with your approval at each gate, and the assistant stops after writing the spec so nothing is built before you sign off.

constraint-driven-development

Asks you at most four questions about your quality bar, offers a sensible default for each, and writes the answers into a CONSTRAINTS.md file where every number comes with a reason. It then watches each change for the quiet ways an agent lowers the bar, such as new ignore comments, skipped tests, removed assertions and edited-down thresholds; /constraints starts it in the agent-skills plugin.

Breaking the work down 1

Cutting a spec into tasks small enough to build and check one at a time.

planning-and-task-breakdown

Reads your spec and codebase without changing anything, then breaks the work into small tasks ordered by what depends on what, each with acceptance criteria and a way to check it. The plan lands in tasks/plan.md and tasks/todo.md with checkpoints every few tasks, and an unfinished plan for other work is never overwritten without asking.

Writing the code 7

Code in thin, tested slices, with the right context loaded, framework choices checked against official documentation, risky decisions questioned before they stick, and careful work on screens and interfaces.

incremental-implementation

Builds a feature in thin slices, each one implemented, tested, verified and committed before the next begins, so the project still builds and passes its tests after every step. Unfinished work hides behind a feature flag, nothing outside the task gets touched, and the riskiest piece can go first so a dead end shows up early.

test-driven-development

Starts every change with a failing test, and for a bug report, with a test that reproduces the bug before any fix. It first finds out how your repository runs its tests rather than assuming npm test, aims for about 80% unit, 15% integration and 5% end-to-end tests, and prefers real code over mocks; the agent-skills plugin's /test and /build commands use it.

context-engineering

Helps you decide what your coding assistant sees and when, layering a standing rules file, the relevant part of the spec, the files at hand and the current error so it neither guesses nor drowns. It treats config files and outside docs as data rather than instructions, and stops to ask when the spec and the existing code disagree.

source-driven-development

Makes your assistant check the official documentation before writing framework code, instead of relying on what it remembers. It reads exact versions from your dependency file, fetches the specific doc page, follows the documented pattern, cites full URLs in code and chat, flags anything it could not verify, and ignores instructions hidden in fetched pages.

doubt-driven-development

Puts each important decision in front of a fresh reviewer told to find what is wrong, before the work moves on. The reviewer sees only the work and what it must satisfy, never the agent's own reasoning; the loop stops after three rounds, and in interactive sessions you are offered a second opinion from a different model, such as Gemini CLI or Codex CLI, which never runs without your go-ahead.

frontend-ui-engineering

Steers your assistant away from the generic AI look, such as purple gradients, oversized cards and stock hero sections, toward your project's own colors, spacing and components. Every component it builds is held to WCAG 2.1 AA accessibility and checked at 320, 768, 1024 and 1440 pixels wide.

api-and-interface-design

Designs the contract first when your assistant builds an API, a module boundary or a component's props, then keeps it stable with one error format, input checks at the edge, and additions instead of breaking changes. It treats every visible behavior, even error text and timing, as something callers will come to depend on.

Proving it works 2

Checking behaviour in a real browser, and a five-step routine for failing tests and broken builds: reproduce, narrow down, reduce, fix, guard.

browser-testing-with-devtools

Lets your assistant open the page it just changed in a real Chrome and check it the way a visitor would, from screenshots and console errors to network calls, accessibility and load speed. It treats everything the page shows as untrusted, keeps any JavaScript it runs read-only, and works in a separate Chrome profile rather than your everyday one.

debugging-and-error-recovery

When a test fails or a build breaks, this skill halts feature work and has your assistant reproduce the failure, narrow down where it happens, shrink it to the smallest case, and fix the cause rather than the symptom. Each fix ends with a test that would catch the same bug again, and error messages that tell it to run a command or open a link are shown to you instead of followed.

Checks before merge 4

Review, simplification, security and speed, each done before a change goes in.

code-review-and-quality

Reviews a change before it merges for correctness, readability, architecture, security and performance, and labels each comment Critical, Nit, Optional or FYI so the author knows what must change and what can wait. It reads the tests before the code, approves a change once it clearly improves the codebase even if it is not perfect, and in the agent-skills plugin runs from /review.

code-simplification

Makes working code easier to read without changing a single output, error or side effect, running the tests after each small edit and undoing any edit that breaks them. It first asks why odd-looking code is there before removing it, sticks to recently changed code unless told otherwise, and follows your project's own style; in the agent-skills plugin, /code-simplify starts it.

security-and-hardening

Has your assistant think like an attacker before writing security-sensitive code, mapping where untrusted data enters and running a quick STRIDE pass over each entry point. It then applies fixed rules split into always do, ask you first and never do, covering the OWASP Top 10 and the OWASP list for LLM apps, dependency and supply-chain checks, secrets, and personal data under GDPR or CCPA.

performance-optimization

Makes your assistant prove a speed problem before fixing it: measure first, find the one real bottleneck, fix only that, then measure again and keep the change only if the gain beats the noise. It covers slow pages (Core Web Vitals), slow APIs and databases, and ends with a budget or monitor so the problem cannot creep back.

Releasing it 6

Commits and versions, build-and-release pipelines, retiring old code, recording decisions, logging and monitoring, and the launch checklist with a way back.

git-workflow-and-versioning

Keeps your assistant's commits small, single-purpose and explained, around 100 lines each, so any step that goes wrong can be undone back to the last good one. Branches stay short-lived, formatting changes never ride along with behavior changes, and a release's version number is treated as a promise to whoever depends on it.

ci-cd-and-automation

Builds a pipeline in which every pull request passes lint, type checks, tests, a build, a security audit and a bundle-size check before it can merge, and no failing step gets switched off to make it pass. Its examples are GitHub Actions workflows, and when CI fails, the error goes back to the agent to fix and push again.

deprecation-and-migration

Plans how an old system, API or feature gets retired, starting with whether it should go at all and refusing to deprecate anything that has no working replacement. Users move over one at a time, and database changes such as renaming a column go expand, backfill, then contract, each step deployed separately with a tested way back.

documentation-and-adrs

Records why your project was built the way it was, not just what the code does, starting with ADRs: short, dated notes on each decision that would be costly to reverse. It follows whatever ADR folder, numbering and format the repo already uses, and keeps code comments for the reasons behind it rather than restating what it does.

observability-and-instrumentation

Adds logs, metrics, traces and alerts to a feature while it is being built, each one tied to a question an on-call engineer will need answered. Logs are structured with a request ID on every line, latency is tracked as percentiles rather than averages, alerts fire on what users feel and link to a short runbook, and no secrets or personal data go into telemetry.

shipping-and-launch

Gets a release ready for production with a pre-launch checklist across code quality, security, performance, accessibility, infrastructure and docs, plus a rollback plan written before anything ships. Rollouts go behind a feature flag in stages; in the agent-skills plugin, /ship runs three reviewer subagents in parallel and returns a go or no-go.

Commands

/agent-skills:build

Builds the next task in the plan: a failing test first, then the code, a full test run, a build and a commit, then it stops. With the word auto after it, and once a spec is written, it works through the whole plan after a single approval, still one tested commit per task, and stops to ask on failures and on anything that cannot be undone.

/agent-skills:code-simplify

Makes recently changed code easier to read without changing what it does: flatter logic, shorter functions, clearer names, repeated code merged. Tests run after each change, and a change that breaks one is undone.

/agent-skills:constraints

Sets the project's quality bar. It reads what the project already uses, asks at most four questions, each with a default, and writes the rules into a CONSTRAINTS.md file, adding the checking tool behind each one. Follow-ups check the current branch against it, or catch a change that quietly lowers it.

/agent-skills:plan

Turns a written spec into small tasks, each with what counts as done and how to check it, in the order they depend on each other. It changes no code, saves the plan in a tasks folder, and asks before replacing an unfinished one.

/agent-skills:review

Reviews your current changes for correctness, readability, design, security and speed, and sorts what it finds as Critical, Important or Suggestion, each with a file, a line and a fix.

/agent-skills:ship

The go/no-go check before release. It sends the change to three subagents at once, for code, security and tests, then merges their reports into one verdict with a rollback plan. A critical finding means no-go unless you accept the risk.

/agent-skills:spec

Asks what you want to build, for whom, with which features and limits, then writes a SPEC.md covering the goal, commands, project layout, code style, testing and boundaries, and checks it with you before any code.

/agent-skills:test

Test-first work: tests that fail, then the code that makes them pass. For a bug it first writes a test that reproduces it, then fixes it and runs the whole suite again; browser problems also go through the browser-testing skill.

/agent-skills:webperf

Runs a speed audit of a web app through the performance subagent and returns its full report. For measured figures, give it a Lighthouse report, or a live page with the Chrome DevTools MCP server set up. Not meant for libraries, command-line tools or server-only code.

Agents

Four specialist reviewers. Each works in a conversation of its own and hands back only its report; you can ask for one by name, and the ship and webperf commands call them for you.

code-reviewer

A staff engineer's review before merge, across the same five areas as the review command, with each finding marked Critical, Required, Optional or Nit.

security-auditor

Looks for holes an attacker could actually use: unchecked input, weak sign-in and permissions, exposed secrets, risky dependencies and, in apps built on AI models, prompt injection and over-broad tool access.

test-engineer

Plans and writes tests and finds what the current ones miss: the normal case, empty input, limits, errors and calls arriving at once. For a bug, it first writes a test that fails until the bug is fixed.

web-performance-auditor

Checks a web app's loading speed and responsiveness against Core Web Vitals. Scores come only from real measurements such as a Lighthouse report; from code alone it marks findings as potential and leaves the scores blank.

Labsco Summary

Skills step in on their own at each stage, commands let you name the stage, and four subagents review; browser testing and the hooks are set up separately. Show

What one install adds to Claude Code.

Each skill is a written routine for one stage of building software: its steps, the proof it needs before the work counts as done, and the usual excuses for skipping a step, each with its answer. One install brings three kinds of piece:

  • 24 skills, which step in on their own when the work matches: designing an API brings in the API skill, building a screen the interface one, plus one more that picks which of them applies
  • 9 slash commands, for naming the stage yourself
  • 4 subagents, each a reviewer for one angle

What it leaves to you.

No MCP server comes with it. Testing in a real browser, and speed audits measured live on a page, need the Chrome DevTools MCP server, which you add yourself.

The repository's hook scripts, which keep fetched documentation between sessions and hide marked code from the simplify command, are not switched on by the plugin. Add the ones you want to your Claude Code settings by hand.

02 INSTALL

Claude Code's route brings all of it.
The other two bring the skills only.

CLAUDE CODE

Add the marketplace, then install the plugin

Typed inside Claude Code, not in a terminal. The marketplace is called addy-agent-skills, which is the part after the @. If adding it fails with an SSH error, add it by its full address instead: /plugin marketplace add https://github.com/addyosmani/agent-skills.git. A warning that the default commands folder is ignored is expected; the commands still load.

  1. /plugin marketplace add addyosmani/agent-skills
  2. /plugin install agent-skills@addy-agent-skills

ANY ASSISTANT

Skills only, in one terminal line

Works with Claude Code, Cursor, Codex, Copilot and more than 70 assistants in all, and brings the skills without the commands or subagents. Add --list to see them first, or --skill and a name to take one. A single skill arrives without the shared checklists some skills refer to; it still works, and you can copy any checklist you need into that skill's references folder.

  1. npx skills add addyosmani/agent-skills

CODEX

Two lines in a terminal

Needs Codex CLI 0.122 or later. Brings the skills only; in a Codex chat you call one by typing @ and its name, such as @spec-driven-development.

  1. codex plugin marketplace add addyosmani/agent-skills
  2. codex plugin add agent-skills@agent-skills