Engineering

Claude Code Skills: Turn Repeated Work into Reusable Playbooks

On this page
  1. What a skill is
  2. Progressive disclosure
  3. When a skill is worth writing
  4. The description does the triggering
  5. Write the body for a smart reader
  6. Scripts do the deterministic parts
  7. Four skills for one site
  8. Lessons from testing them
  9. A script that changes production needs a way back
  10. Shell scripts fail in quiet ways
  11. Let a fresh agent try it
  12. State what the agent cannot know
  13. A starting checklist
  14. Frequently asked questions

Every new Claude Code session starts from zero. It is good at working out how a project is deployed or how a post gets published, but it works it out again each time, with a different set of small mistakes. For this site, the deploy flow, the publishing steps and the server checks were re-derived in almost every session. Skills fix that: a skill is a folder of instructions and optional scripts that Claude Code loads only when a task matches. This post explains how skills work, how to write one worth the effort, and what I learned building four for this site.

What a skill is#

A skill is a directory with a SKILL.md file. The file starts with a short YAML header (name and description) and continues with Markdown instructions. Next to it you can keep scripts, reference documents and templates.

text
.claude/skills/
  blog-post/
    SKILL.md              # when to use it, and the steps
    scripts/
      check.mjs           # dry run with SEO checks
      publish-prod.sh     # publish on the server
    references/
      seo-checklist.md    # the long lists, read only when needed
    templates/
      post.example.mjs

Skills live in two places. Project skills sit in .claude/skills/ inside the repository, are committed with the code, and apply when Claude works in that project. Personal skills sit in ~/.claude/skills/ and follow you to every project. A skill runs when Claude decides it fits the task, or when you type its name as a slash command, such as /blog-post; anything you type after the command is handed to the skill as its arguments.

Progressive disclosure#

The part that makes skills cheap is how little of them is loaded. At the start of a session Claude sees only each skill's name and description, about a hundred words per skill. The body of SKILL.md is read when the skill is used. Scripts run without being read at all, and reference files are opened only when the instructions point at them. So a project can carry a lot of knowledge in skills without paying for it on every request.

SKILL.md
---
name: blog-post
description: Write, publish, update or remove blog posts on this site through the app's
  content service, with the SEO fields done right. Use it whenever the owner asks for a
  post, an article or a guide, wants a post changed or removed, or asks about a post's SEO.
---

# Blog posts

Posts live in PostgreSQL as Editor.js JSON. The app renders them, keeps revisions and
builds the sitemap, but only when a post is saved through the content service, so
never write to the posts table directly. The scripts here call that service.

## 1. Write the draft
...

When a skill is worth writing#

Three conditions, and all three should hold:

  • The task repeats. A deploy, a release, a post, a server check. A task you do once does not need a skill; a good commit message or a doc is enough.
  • It has several steps, and the order matters. Skills shine where a step is easy to forget: the test file that pins the list of migrations, the cache that only a restart clears.
  • It mixes judgment with mechanics. Deciding what belongs in a release is judgment. Scanning the diff for secrets is mechanics. A skill holds the first in prose and the second in a script.

Before adding a skill, check whether an existing one can be extended. Two skills that overlap compete for the same prompts.

The description does the triggering#

The description is the only part Claude sees before deciding to use a skill, so it carries the whole job of matching. Say what the skill does and, more important, list the situations and the words a person would use. Claude tends to under-use skills, so be more explicit than feels natural.

yaml
# Too quiet: it will rarely trigger
description: Operate the production server.

# Works: what it does, and the phrases that should bring it in
description: Operate the production server behind the site - check health, memory and
  disk, read app, database or Nginx logs, restart, take or restore backups, run SQL,
  change a setting in .env, or triage an incident. Use it whenever the owner asks how the
  server or site is doing, says something is down or slow, or mentions logs, backups,
  the database, the admin login, email settings or the certificate.

Keep the "when to use" in the description and nothing else there. The body is for the how.

Write the body for a smart reader#

Claude follows instructions better when it understands why they exist. "Restart the app after publishing" gets skipped the moment it looks unnecessary. "Restart the app after publishing, because the running process caches the blog index in memory and a publish from another process can't clear it" gets done. Explain the reason, name the failure the rule prevents, and avoid shouting in capitals. A wall of MUST and NEVER reads as noise.

Some shapes that have worked for me:

  • Numbered steps for the main flow, in the order that keeps things green.
  • A table of commands with one line each on what they do.
  • A "when something fails" section with the exact error text and what it means.
  • A list of which commands are read-only and which change something, so "don't touch anything" is unambiguous.
  • Facts that save reading: which test pins what, where a new column shows up by itself, which log files are the real ones.

Keep SKILL.md under a few hundred lines. Long lists, schemas and checklists go into references/, with a sentence in the body saying when to open them.

Scripts do the deterministic parts#

If every use of a skill would make Claude write the same helper, write it once and ship it with the skill. For this site that meant a preflight that runs lint, formatting, the relevant tests and a secret scan; a dry-run that publishes a draft into a throwaway database schema and checks the rendered page; a wrapper that runs a command on the server; a migration generator that also updates the test which lists every migration.

A few rules for those scripts:

  • Print what they did and exit non-zero on failure. Claude reads the output and acts on it.
  • Work from any directory. Resolve the repository from the script's own location, not from the current directory.
  • Refuse to do dangerous things by default. The SQL wrapper here is read-only unless you pass --write; the settings script refuses to touch the keys the containers depend on.
  • Never embed a secret. Skills are committed to the repository. Scripts read secrets from the server, the environment or a file that is git-ignored.

Four skills for one site#

This site runs on a small VM with Docker and Nginx, and I kept finding the same four kinds of work coming back:

SkillTriggers onScripts
shipdeploy, release, roll back, "is the site updated"preflight checks around the deploy script
blog-posta new post, a change to a post, a post's SEOdraft helpers, dry run, publish, live verification
prod-opshow is the server, logs, backups, settings, incidentsSSH wrapper, backup download, read-only SQL, .env changes
extend-cmsa new field, column, block type, admin screenmigration and block-type generators, checklists

Each took about an hour to write and as long to test. The testing is where the value was.

Lessons from testing them#

A script that changes production needs a way back#

The first test of the settings script wrote an empty value into the production .env. The app refused to start, and the site answered 502 for about two minutes until I restored the file by hand. The cause was a one-line mistake: the script fed the value through standard input, and something else was already using standard input. The lesson is not "be more careful". It is that any script which recreates a running service must wait for the health check and put the previous state back when the check fails. The script now does that, and it keeps the failing container's log for the post-mortem.

Shell scripts fail in quiet ways#

The live-page verifier reported that every post was missing its structured data. The data was there. The script piped a 60 KB page into grep -q, which exits as soon as it finds a match; under set -o pipefail the writer's broken pipe turned a success into a failure. Bash substring tests fixed it. Treat skill scripts like code: run them against real input and read every line of what they print.

Let a fresh agent try it#

You cannot judge your own instructions, because you already know what they mean. I gave each skill to a fresh Claude Code agent with a realistic task and asked it to report every point of friction. One read-only server check came back with fourteen: a log command that needs a terminal and would hang an agent, log files that were not named, times in UTC while the owner thinks in local time, no read-only way to list backups, no statement of which commands change the server. None of that was visible to me as the author. Every point went back into the skill.

State what the agent cannot know#

Several fixes were plain facts: a container's logs disappear when it is recreated; a test pins the exact list of migrations; the view receives the whole database row, so a new column needs no plumbing. An agent can find these by reading fifteen files, and will, every time. Write them down once.

A starting checklist#

  • Name the task in the description the way people actually ask for it, with the words they use.
  • Put the steps in order, and explain the reason behind each rule that could look optional.
  • Move anything deterministic into a script that prints its result and exits non-zero on failure.
  • Make scripts refuse dangerous actions by default and keep secrets out of the folder.
  • Run every script against real input, in a safe mode, before the first real use.
  • Give the skill to a fresh agent with a realistic task and fix what it trips over.
  • Commit project skills with the code, so they are versioned and travel with it.

Anthropic's engineering post on Agent Skills explains the design behind the format, and the anthropics/skills repository has examples to read before writing your own.

Frequently asked questions#

What is the difference between a skill and CLAUDE.md?

CLAUDE.md is loaded into every session and should stay short: conventions, commands, where things are. A skill is loaded only when its task comes up, so it can hold a long procedure and scripts without costing anything the rest of the time. Rules that always apply go in CLAUDE.md; playbooks for specific tasks go in skills.

Where should a skill live, the project or my home directory?

In the project (.claude/skills/) when it depends on that codebase or server, so it is versioned and everyone who clones the repository gets it. In ~/.claude/skills/ when it is about how you work across projects, such as a commit-message style or a review checklist.

Do skills slow Claude down or use many tokens?

Not much. Only each skill's name and description are loaded at the start; the body is read when the skill is used, and scripts run without being read. The cost appears only in the sessions that need the skill, and it is usually smaller than the cost of Claude working the procedure out from scratch.

How do I test a skill?

Run its scripts directly against real input first, in a mode that cannot cause harm. Then give the skill to a fresh Claude Code session or a subagent with a realistic task and ask for a report of everything that was unclear or wrong. Fix the skill, not the agent, and repeat once.

Comments

No comments yet. Be the first.

Leave a comment

Only used to tell you about a reply. Never shown.

Plain text; line breaks are kept.

Let's build something

Hiring for a backend role, or have a project in mind? Send a message and I will reply by email.

At most 2 links.