Labsco
microsoft logo

creating-debug-tests-and-iterating

✓ Official★ 4,933

by microsoft · part of microsoft/fluidframework

When a bug will not reproduce, this drops the existing tests entirely and goes after the real running program instead, adding logging on every attempt until the cause is visible.

WHEN YOUR AGENT SHOULD USE IT

A QUICK BOUNDARY

USE FOR

  • Reproduce a stubborn bug with a script that hits the real CLI, API, or UI from outside.
  • Debug a web app end to end with Playwright instead of relying on internal mocks.
  • Add logging on every iteration until you can see exactly where behavior diverges.
  • Clean up debug logs and background jobs once the root cause is found and fixed.

DO NOT USE FOR

  • Not for verifying with mocks or test harnesses; it drives the real application end to end.

This is the playbook your agent receives when the skill activates — you don't need to read it to use the skill, but it's here to audit before installing.

*CRITICAL* Add the following steps to your Todo list using TodoWrite:

From this point on, ignore any existing tests until you have a working example validated through a new test file.

  1. Write a script that interacts with the application from the outside. The script should not call any internals. It should only interact with the external interfaces.
  2. Check to see if you require authentication. If you do, ask me for credentials. Do NOT use mock mode or test harnesses. You should be testing the real thing.
  3. Follow these steps in a loop until the bug is fixed:
  • Add many logs to the application. You MUST do this on every loop.
  • Run the debug script.
  • Analyze the output: read logs, identify errors, do whatever you need to.
  • Update the debug script. If you get stuck: did you add logs?
  1. Identify and fix the issue at hand.
  2. Clean up all background jobs, and remove extraneous logs.
  3. Make sure other tests pass.

Debug Testing

To test different kinds of applications, write scripts that test the application interfaces. Your testing should be as close to 'real' as possible.

Example

Identify the application boundary to be tested and the tools you need to test it.

CLI Tool:

./path/to/cli.sh arg1 arg2
subprocess.run(["./path/to/cli.sh", "arg1", "arg2"])
exec('./path/to/cli.sh', (error, stdout, stderr) => {
  if (error) {
    console.error(`exec error: ${error}`);
    return;
  }
  console.log(`stdout: ${stdout}`);
  console.error(`stderr: ${stderr}`);
});

API:

Start the server:

cd backend && python server.py&
cd frontend && npm run dev&

Call to the server using scripting language of choice.

Do NOT get in a loop where you just keep running other tests. In this mode, you should ignore other tests entirely until it works.

Emulators for testing

Web servers or web apps: use playwright (read the .claude/skills/webapp-testing/SKILL.md) TUI tools: use tmux with screen capture CLI tools: use bash