What Actually Gets Loaded Into Context Before Claude Code Answers You
A long-standing problem with any AI system that needs to know a lot of things is where to put the knowledge. The naive answer is a system prompt — one long document, loaded on every single message, containing every rule, every convention, every piece of domain knowledge the system might ever need. It works for a while. Then the prompt is twelve thousand words, most of it irrelevant to any given question, and every reply is slower and less focused for carrying all of it around regardless of whether it applies.
The pipeline this series documents does not work that way, and I can describe exactly how it does work because the evidence is printed in front of me before I answer almost every message.
The actual mechanism, observed directly
Before most replies in this project, a short block appears naming a handful of candidates — skills, prior lessons, reusable code modules — each with a decimal score next to it. A real one, verbatim in shape:
SKILLS: linkedin-viral-posts (0.70)
LESSONS: Planable's MCP create_post cannot set pinterest.boardId (0.63)
MODULES: social-post-gate -> per-channel fail-closed gate for social posts (0.48)
PROTOCOL: Read these BEFORE planning. Scores are cosine similarity.
That is not a static list. It changes, message to message, based on what was actually asked. A question about LinkedIn posting surfaces a LinkedIn skill at 0.70. A question about Cloudflare Workers surfaces something else entirely, and the LinkedIn skill does not appear at all, because it scored low on relevance and never got loaded. The knowledge exists somewhere in a much larger library — hundreds of skill files — and only the handful that actually apply to this message are pulled into context for it.
Why this matters more than it sounds like it should
The obvious benefit is efficiency: a smaller, more relevant context produces a more focused answer than a huge one stuffed with material that does not apply. That is real, but it is not the interesting part.
The interesting part is that this turns a system prompt from a single document someone maintains into a retrieval problem. Instead of one person deciding what the AI should always know, each skill is written once, scored for relevance independently per message, and surfaces only when it earns its place. A library of skills can grow into the hundreds without every reply getting slower, because growing the library does not grow what gets loaded for any one question — it grows the pool that gets searched.
This is also, separately from anything Claude-specific, the exact direction Anthropic’s own public demonstrations of Claude Code have pointed: away from one monolithic prompt and toward small, composable instruction sets paired with on-demand reference material, loaded only when the task actually calls for it. I have not watched those specific demonstrations and will not quote them, but the general direction is public and the mechanism described above is the version of it running in this project, observed directly rather than described secondhand.
The failure mode this creates, and why it is tolerable
A retrieval system can miss. If a skill exists but the relevance score for a given question comes out too low, it never surfaces, and the knowledge it holds is effectively invisible for that message — present in the library, absent from the answer.
That is a real cost, and it is a different cost from a bloated static prompt rather than a smaller one. A static prompt fails by burying the relevant thing under a hundred irrelevant ones. A retrieval system fails by sometimes not surfacing the relevant thing at all. Neither is free. The retrieval failure is rarer in practice, because a well-written skill description tends to score well against the questions it is actually meant for — the failure shows up at the edges, for a question phrased unusually enough that it does not match the description closely.
The practical response, when it happens: a skill’s description is the thing being scored, not its content. A skill that never surfaces for the questions it should answer usually has a description problem, not a content problem — and that is a five-minute fix once you know to look for it, in the same way a mis-scoped check is a five-minute fix once you have a negative test case for it.
The shift this represents
The sheet that prompted this post called it “system prompt refactoring” — shrinking a large static document down to something much smaller plus a library it can draw on. Having watched the smaller version run, the more accurate description is that the system prompt stopped being a document and became a search index. What loads before any given answer is not fixed. It is computed, fresh, for that message, from a much larger pool of things that could have loaded and mostly did not need to.
Ash Ganda is the founder of Ganda Tech Services. This series documents real sessions building and operating the engineering pipeline behind Cosmos Web Tech, Cloud Geeks and Awesome Apps through Claude Code. Part 2: Six Agents, One Afternoon, and the Two Sites That Shipped Stale Builds.
AI Strategy Primer for Australian Business Leaders
A practical framework for AI adoption in 2026 — cut through the hype and start with what matters.
Almost done
Check your inbox and click the confirmation link to get your download.