A skill can pass every check and still fail in use: the AI never loads it, loads it for the wrong request, or follows it and gets the job wrong. The Tests section lets you write down what “working” means for a skill or agent, have your AI tool try it, and see the results next to the skill. This guide covers how it works, how to write good tests, how to run them and how to read the results.
How it works
The Tests section has three parts, and only the first happens in Skill Builder:
- You write the tests in the skill’s or agent’s details. They are saved with it.
- Your AI tool runs them. With Connect AI on, it reads the tests, tries each one, and sends the results back. Skill Builder never runs a model itself.
- The results show in the Tests section, as counts, with a pass or fail mark on each test.
| What | Where it lives | Saved in versions | Exported |
|---|---|---|---|
| Tests | The skill or agent | Yes | Skills, as evals/ |
| Results of a test run | The project (the last 5 runs of each skill or agent) | No | No |
Tests are optional. Nothing needs them, and a skill without tests exports exactly as before.
Open the Tests section
Open a skill or agent and look under its details, below the fields. Select Tests to expand it. The heading shows how many tests there are and, once you have run them, the newest result.
Should it be used? (trigger tests)
AI tools decide whether to load a skill from its description alone. Trigger tests check that decision.
- Should use it (add with + Should use it): a request that this skill is for, written the way people really ask. “Fill in this tax form for me.”
- Should not (add with + Near-miss): a request that shares words with the skill but needs something else. “Convert this Word file to PDF.”
Change the choice next to a request at any time.
Write 8 to 10 of each. Good trigger tests:
- Vary the wording. Formal and casual, with and without the file type or tool name, short and long.
- Use near-misses, not random requests. “Write a poem” is so far from a PDF skill that it tells you nothing. The useful negatives are the ones a keyword match would get wrong.
- Make them substantial. AI tools often skip skills for simple one-step requests they can do alone (“read this PDF”), even when the description matches, so those make poor tests.
- Include competing cases. If the project has a similar skill, add requests where this one should win, and ones where it should not.
Does it do the job? (task tests)
Task tests check the instructions in the body. Add one with + Task. Each has three parts:
- The request: a realistic task, with any detail it needs. “Fill form.pdf with these details: …”
- What a good result looks like: one line, in words. “A filled PDF with every field completed.”
- Checks: statements anyone could verify from the output, one per line. “Every form field has a value.” “The date is in YYYY-MM-DD.” Add them under Checks with Add item.
Good checks are objective. “The output is high quality” can not be judged the same way twice; “lists every field from the form” can. For things only a person can judge, such as writing style, say so in the expected result and judge the output yourself.
Three to five tasks that cover the main use and an edge case or two are usually enough to start.
Run the tests
You need Connect AI turned on and your AI tool paired with Skill Builder, with the project open.
In the AI tool, run the evaluate-skill slash command with the skill’s name:
- Claude Code:
/mcp__skill-builder__evaluate-skill pdf-forms - VS Code Copilot Chat:
/mcp.skill-builder.evaluate-skill, then give the name
The command gives the AI tool a routine to follow:
- Read the tests from Skill Builder. If there are none, it writes some, shows them to you, and saves them only once you agree.
- Trigger tests. Where it can, it sends each request to a fresh subagent that has the skill installed and notes whether the skill was loaded (method: each test in a fresh context). Where it can not, it gives a fresh subagent only the names and descriptions of the project’s skills and asks which one it would use (method: simulated selection).
- Task tests. It runs each task in a fresh subagent with the skill, and again without it, then judges every check strictly from the output. A task passes when all its checks hold.
- Records the results in Skill Builder: pass or fail for each test, a short note with the evidence, the model it ran on, and the method.
- Tells you the counts, which tests failed and why, and what in the description or instructions would most likely fix them. It asks before changing the skill.
If your AI tool does not offer MCP slash commands, ask in the chat instead, for example “Run the tests of the pdf-forms skill with Skill Builder and record the results.” The AI tool can still read the tests (get_document), get the testing guidance (list_templates with guidance evaluation) and record results (update_evaluation), but it works out the steps itself rather than following the routine above.
Read the results
The newest run appears at the top of the Tests section:
- The summary, such as “8 of 10 trigger tests, 2 of 3 task tests passed”, with the AI tool, the model, the method and when it ran.
- A mark on each test: a tick for passed, a cross for failed. Tasks show the AI tool’s note beside them.
- Results out of date (highlighted) when the skill has changed since the run. Editing only the tests does not make results out of date; editing the skill does. Run the tests again before you trust the numbers.
What failures usually mean:
| What failed | Likely cause | What to change |
|---|---|---|
| A “should use it” request | The description is too narrow, or uses different words | Add the situations and words people use; put the main use first |
| A near-miss | The description is too broad | Say what the skill is not for, or narrow the wording |
| A task, also failing without it | The task is hard, or the checks are too strict | Check the task and its checks before changing the skill |
| A task that passes without it | The skill adds little for that task | Focus the skill on what the AI gets wrong on its own |
| A task that fails only with it | The instructions mislead | Read the AI tool’s note and fix the step it points to |
Results are counts, not a reliability score: 9 of 10 means nine tests passed once, on one model. A small change can move a result either way, so run the tests again after each change and compare.
Limits worth knowing
- Results come from your AI tool. Skill Builder records what the tool reports and labels it with the tool’s name, but can not check it. Read results as evidence, not proof.
- Simulated selection is close, not the same. Asking a model which skill it would use is not exactly how a tool decides to load one. Prefer a tool that can start fresh sessions with the skill installed, such as Claude Code.
- Results depend on the model. A skill that works on a large model may need more detail for a small one. The model is recorded with each run so you can compare.
- Size. Up to 50 tests of each kind per skill or agent, and 20 checks per task. The last 5 runs of each are kept.
- Agents can have tests, but they are kept in the project only: an agent is a single file, so there is nowhere to export them.
Tests in the exported skill
A skill’s tests are exported with it, next to SKILL.md:
evals/evals.json: the task tests, in the format of Anthropic’s skill-creator (prompt,expected_output,expectations).evals/triggers.json: the trigger tests, as[{ "query": "…", "should_trigger": true }].
SKILL.md never links to them, so AI tools do not load them when they use the skill. When you import a skill that has these files, from GitHub, an upload or a zip, they come back as its tests. That means you can keep tests in your repository, run them with skill-creator, and edit them in Skill Builder.
A routine for improving a skill
- Write the tests before you change anything, and run them once for a baseline.
- Change one thing: usually the description for trigger failures, the steps for task failures.
- Run the tests again and compare with the baseline.
- Keep the change if it helped, undo it if it did not, and save a version when you are happy.