All posts

Agent Skills

/AI/4 min read

Loading fifty procedures eagerly costs 210,000 tokens and does not fit. Loading their descriptions costs 4,750, which makes the whole thing work — and turns the real problem into a fifty-way choice made from twenty words.

An agent skill is a folder containing instructions — and optionally scripts and reference files — that an agent loads by itself when it recognises the situation.

It exists because the alternative does not fit.

Why not just put it in the prompt

A team accumulates procedures: how a release note is formatted, which checks run before a deploy, what a data export has to include, the conventions a migration follows. Written down, each is a page or two — call it 4,200 tokens.

Fifty of those, loaded on every request:

tokens
50 procedures, fully loaded210,000
a 200,000-token context window200,000

It is over budget before the agent has read a single file of yours. And even if it fitted, 49 of the 50 are irrelevant to whatever was asked — diluting attention and paying for context that will not be used.

Loading in stages

The answer is to load almost nothing and expand on demand.

Level one is a name and a one-line description, about 95 tokens per skill. Always present.

Level two is the full instructions, loaded only for the skill that matched — 4,200 tokens.

Level three is whatever the instructions themselves point at: a longer reference, a schema, an example file. Read only if the work needs it.

tokens
50 descriptions4,750
one skill expanded4,200
total8,950

23× less than loading everything, and the descriptions cost 2.4% of the window — an amount you would not notice.

LOAD THE INDEX, NOT THE LIBRARY the window — 200,000 tokens everything loaded 210,000 — does not fit descriptions, then one skill 8,950 — 23× less the 44 skills that did not fire cost 4,180 tokens — 2.1% of the window so unused skills are not the problem. being picked wrongly is. selection is a 50-way choice made from twenty words each
The saving is large and easy. What it buys is a harder problem: choosing correctly from descriptions alone.

The description is the whole interface

Here is the consequence that matters, and it follows directly from the staging.

At level one the agent sees nothing but the descriptions. Whether a skill is ever used is decided entirely by whether its one line matches the situation — so the description is not a label, it is the trigger.

Which means the work of writing a good skill is mostly the work of writing that line. Two failures:

Too vague. "Helps with reports" matches nothing in particular and will be skipped in favour of doing the work from scratch. The instructions inside may be excellent and will never be read.

Too broad. "For any data-related task" matches constantly, including when something else was the right choice, and now a specific procedure is being applied to a situation it was not written for.

The useful shape is to say when, not what: "Use when producing the weekly operations summary — covers the required sections, the metric definitions, and the sign-off order." Someone reading that knows whether it applies. So does a model.

What unused skills actually cost

Worth doing this arithmetic because the intuition is wrong.

Of 50 skills, suppose 6 are relevant in a typical session. The other 44 cost 44 × 95 = 4,180 tokens — 2.1% of the window. Negligible.

So the reason to prune a skill collection is not token cost. It is that every added skill makes the selection problem harder: fifty descriptions that overlap are harder to choose between than twenty that do not. Keep them because they are distinguishable, delete them because they are confusable, and do not think about the tokens at all.

Carrying code

A skill can include scripts, and this is more than a convenience.

Asking the model to write the code costs around 700 tokens and produces something slightly different every time. Invoking a script costs about 40 tokens and produces exactly the same behaviour on every run — 17.5× cheaper and, more importantly, deterministic.

Which suggests the division: put the judgement in the instructions and the mechanics in a script. Anything with a single correct implementation should not be regenerated on each use.

Where it sits next to a connectivity protocol

Two different questions, often confused because both extend an agent.

A connectivity protocol answers what can the agent reach — it exposes tools and data sources that were previously unreachable.

A skill answers how should this be done here — it supplies procedure for something already reachable.

An agent with tools and no procedure will do the work in a way that is reasonable and not yours. An agent with procedure and no tools cannot do the work at all. They are additive.

What to take away

Progressive disclosure is the mechanism, and it is not complicated: keep an index in context, expand one entry when it matches.

The interesting part is what that shifts. Once loading is cheap, the constraint stops being context and becomes selection — and a skill's value is decided by one line of description that has to be specific enough to be recognised and narrow enough not to be recognised wrongly.