Module 2 · Tools, Knowledge, Memory & Research Design · scripted
Exercise: Build a Skill
The goal: onboard one brilliant employee
You have just seen what a skill is: the folder handed to a brilliant employee on their first day. Now build a real one — for a recurring task in your own research world. Not a toy: pick something you actually do more than once and wish you never had to re-explain. Coding a transcript against your scheme. Writing the recruitment email for a study. Reviewing a manuscript with your field's checklist. Turning raw instrument output into the table your papers use. If you have explained it twice to a student, it's a skill.
What you are building
The test of success: a fresh conversation that has NEVER met you completes the task with nothing but this folder.
The exercise at a glance
- 1Step 1: write the catalog line — the name and when-to-use description — and stress-test it.
- 2Step 2: build the package through the interview — references, a worked example, a template.
- 3Step 3: the acid test — a fresh conversation gets only the folder, and you fix stumbles in the FOLDER, not the chat.
The exercise playground
For this exercise we have built an exercise playground — the Skill Builder — to help you do it: an agent interviews you about the task and writes the package with you. It drafts, you correct, the files appear in the tree as you talk, and you can download the finished skill as a zip. (The Agent tab shows exactly how that interviewer works, if you want to peek.)
You can also do the exercise in ChatGPT, Claude, or Gemini instead — if you're going to use one of those tools, click here for the instructions →.
How the playground works
- 1Open it: /exercises/skill-builder — or scan the QR code on the next card.
- 2Answer the interviewer's questions — concretely.Every vague answer becomes a stumble on someone's first day.
- 3Watch the files appear in the tree as you talk; correct anything it drafts wrong.
- 4Download the finished skill as a zip when you're done.
Scan to open the Skill Builder

The interview, flipped
You are not writing documentation. You are being INTERVIEWED by the future employee's advocate:
Answer concretely. Every vague answer becomes a stumble on someone's first day.
Step 1 · The catalog line
Before any content: the name and the description — the one line the agent reads while deciding whether to open your folder at all. Write it trigger-style ("Use this skill when…"), and add a counter-example only if you can name the actual neighbor it disambiguates from. Then stress-test it: give the line to the Skill Builder agent (or your LLM) with three task descriptions — one clearly inside, one clearly outside, one borderline — and ask which ones it would open the skill for.
Capture — end of Step 1
Copy into your course document and save:
- 1The catalog line, verbatim: name + when-to-use description.
- 2The borderline task, and whether the line routed it correctly.
Step 2 · The package
Build the body through the interview. Keep SKILL.md a genuine table of contents — pointers, not payload — and push the weight down into the folders: at least one reference (the knowledge you'd hand a newcomer), one worked example (a real input with its real finished output — examples are the facet your field most neglects), and one template with placeholders (the starting point that carries your structure). Add a script only if the task genuinely warrants one — remember the substrate: your employee has eyes and judgment, not just a shell.
Capture — end of Step 2
Copy into your course document and save:
- 1Your SKILL.md, verbatim.
- 2The one decision the interview forced that you would never have written down on your own — the "what does clean MEAN here?" moment.
Step 3 · The acid test
Now the first day actually happens. Open a fresh conversation — one that has never met you — hand it only the skill and a task, and watch. In the Skill Builder, start a new chat and point it at your saved skill; with your own LLM, upload the zip to a clean session. Do not coach. Where the employee stumbles, resist the urge to answer in chat — that fix evaporates when the conversation ends. Fix the folder instead, and run the test again.
Capture — end of Step 3
Copy into your course document and save:
- 1Where the fresh conversation stumbled on its first attempt.
- 2Which facet was missing — knowledge, example, template, steps, or the catalog line itself.
- 3The file change that fixed it, and what happened on attempt two.
Before we regroup
Copy your three captures into your course document and save — and keep the skill itself (the zip, or the folder). At the end we read the catalog lines aloud as a catalog, revisit the interview moments that surfaced knowledge people didn't know they had, and tally the acid-test stumbles by facet — that tally is a map of what our fields leave undocumented.
Copy into your course document and save:
- 1All three captures above, in your course document.
- 2The skill itself: the downloaded zip, or the folder.
Be ready to discuss:
- Whose acid test survived attempt one? What did their folder have that yours didn't?
- What did you fix in the FOLDER that you were tempted to fix in the CHAT — and why does that difference matter for the next hundred first days?
- Look at the catalog as a whole: which two skills would an agent confuse? Whose counter-example earns its place?