Blog

Your Coding Agent Knows the File. Does It Know the System?

Updated 2026-08-30 · 26 min read

How Damper gives coding agents the project context, workflow, and proof they need to make safe changes.

Add one optional field to an invoice API response.

The code change takes about five minutes. Your agent will do it correctly. It will update the DTO, adjust the serializer, write a test, and hand you a nine-line diff that looks obviously safe.

Here is what it did not check:

  • The generated Java client deserializes this response as a closed type. A new field throws.
  • The iOS app decodes the same payload with a non-failable initializer that silently drops unknown keys — so the field arrives, and disappears.
  • Three API examples in the public docs show the old response body.
  • The field is optional in the schema and nullable in the database, which are different things, and the OpenAPI spec now says one while the migration says the other.
  • Nobody ran the contract tests, because the contract tests live in a package the diff never touched.

None of these are coding errors. The agent wrote good code. It just could not see the system the code lives in.

This is the failure mode that matters. Syntax errors get caught in seconds. Context errors get caught in production, or in a customer's integration, three weeks later.

The instruction-file trap

The obvious fix is to tell the agent more. So you write it down in AGENTS.md.

That works, briefly. Then the file reaches four hundred lines and starts failing in a specific way: every task gets every rule. The agent changing a button's hover state reads your tax-rounding invariants. The agent changing tax rounding reads them too — buried on line 280, between the CSS conventions and the commit-message format.

You have not built context. You have built a prompt that grows monotonically and degrades continuously. Its failure modes are predictable:

  • Relevant rules compete with irrelevant ones for attention.
  • Nobody owns the file, so it goes stale in place.
  • Monorepo-specific behavior gets flattened into advice general enough to be useless.
  • Updating one rule means editing it in three files that have drifted apart.

The problem is not that the agent lacks information. It is that the information is not addressed to the task.

What Damper is

Damper is a project intelligence and workflow layer for coding agents. It connects to your editor through MCP and, instead of handing every agent the same large prompt, builds a work packet for the specific task in front of it.

You give it the intent and the paths you expect to touch. It returns the sections of project context that actually govern those paths, the surfaces the change reaches, the risk domains it activates, the proof required before the work is done, and the current task and approval state.

The contrast with AGENTS.md is the whole idea. An instruction file is pushed: written once, delivered always, identical every time. A work packet is retrieved: assembled per task, scoped to real paths, and refreshed when the diff grows.

Damper does not replace your repository, your tests, or your reviewers. It connects them, and it makes the connection legible to an agent.

The same task, with a context layer

Back to the invoice field. Before editing anything, the agent declares what it is about to do:

{
  "tool": "prepare_work",
  "intent": "Add optional paymentReference to the invoice API response",
  "likelyPaths": [
    "apps/api/src/invoicing/invoice.dto.ts",
    "apps/api/src/invoicing/invoice.service.ts"
  ],
  "changeKinds": ["contract_changes"],
  "contractStatus": "unsettled",
  "mode": "implementation",
  "bootstrapVersion": "2026-08-14"
}

Two paths and a sentence of intent. That is enough, because the paths are the query. apps/api/src/invoicing/*.dto.ts matches a route that already knows what a public API contract change means in this repository, and the packet comes back with the things the agent could not have guessed:

Governing context — api/openapi, sdk-generation, quality-gates/api. Not the whole manual. The three sections that own these paths.

Affected surfaces — api, docs, js-sdk, mobile, web. This is the answer to "what else does this touch," and it arrives before the edit rather than during review.

Risk domains — public_api. Which activates the compatibility checks.

Scoped checklist — contract tests, SDK regeneration, example validation, and an explicit disposition for any consumer intentionally left alone.

Fixture coverage — and this one is worth dwelling on, because it is where the system is honest about itself. If the routes covering these paths have no fixtures, or have fixtures that were never executed, the packet says so:

Note: Route fixtures are represented but not validated; run validate_context_routing before relying on them.

Representation is not validation. A route that claims to cover a path is a claim, not a test. Damper distinguishes the two rather than letting a confident-looking packet paper over an unverified one.

If the diff later expands into apps/mobile/, the agent calls prepare_work again. The surfaces change, the checklist grows, and the proof requirements grow with them. Scope expansion is a normal event with a defined response, not a thing you hope someone notices in review.

This is useful even if you never let an agent merge anything. The first and largest win is planning. The mobile-decoding question surfaces before implementation instead of after deployment. Everything downstream — the proof, the review evidence, the guarded merge — is built on that, but the planning improvement arrives on day one.

The architecture

Four layers, each with a distinct job:

Request
  ↓
Task + intent + likely paths
  ↓
Damper work packet
  ├─ governing context
  ├─ affected surfaces and dependencies
  ├─ change kinds and risk domains
  ├─ scoped completion checklist
  └─ delegation and verification guidance
  ↓
Implementation + focused proof
  ↓
Exact-candidate review and guarded merge
  1. Repository bootstrap. A short AGENTS.md that tells any agent where truth lives and what to do first. Discovery and failure behavior only — not the truth itself.
  2. Project context. Architecture, domain behavior, quality gates, and workflow rules, in structured, addressable sections.
  3. Code-derived routing and proof. Repository paths resolve to governing sections, risks, checklist scope, and verification profiles. Fixtures test those expectations.
  4. Repository enforcement. Local scripts validate an exact pull-request candidate, record independent review, and guard the merge. Damper supplies the scope and the evidence requirements; the repository enforces them.

An optional fifth layer closes the product loop — authenticated feedback, linked tasks, public roadmap and changelog. Keep it separate from the agent-control plane; they are different systems that happen to share a home.

The principle underneath all of it: give the agent the smallest complete set of verified project truth for the task at hand. Smallest, because attention is finite. Complete, because a gap is worse than a long file. Verified, because the alternative is encoding your assumptions and calling them rules.

Is this worth it for your repository?

Be honest about the answer.

Damper earns its setup cost when your repository holds knowledge that cannot be inferred from one folder: shared APIs with external consumers, several frontends, generated clients, compliance logic, background jobs, or a delivery process with real approval boundaries. If you have one service, one frontend, and no external consumers, a good AGENTS.md may genuinely be enough. Use it.

Where it does fit, the value splits by role. Developers stop rediscovering ownership and downstream effects on every task. Engineering leads turn review expectations into a workflow rather than a habit. Product and operations teams get tasks, decisions, and approval state without encoding process rules in source comments.

One limitation to state plainly: the initial setup needs someone who understands the codebase. Damper helps retrieve and maintain context, and it can draft sections from code you point it at. It should not invent the architecture it is meant to document. If you let an agent write your context layer unsupervised, you get a confident, well-structured record of what the agent assumed — which is the original problem with extra steps.


The rest of this article is the build. It is written to be followed in order, but the first three parts are where most of the value is.

Part 1: Bootstrap the repository

Use AGENTS.md as the universal entry point. If a tool expects its own file — CLAUDE.md, say — either point one at the other or give them small, distinct jobs. Do not maintain two copies of the same rules.

# AGENTS.md

This repository uses Damper MCP as the canonical source of truth for
agent-facing project context and operating rules.

Before working here, read `CLAUDE.md` and follow its complete Damper
bootstrap flow. Resolve governing context and scoped completion
requirements before editing. Refresh resolution when scope expands.

Do not duplicate durable project guidance in local Markdown. Keep
architecture, domain behavior, quality gates, and workflow rules in Damper.

The tool-specific file carries the actual sequence:

# CLAUDE.md

This repository uses Damper MCP for project context, task workflow,
quality gates, and current operating rules.

Damper instruction revision: `<bootstrapVersion from get_agent_instructions>`.

Before coding, reviewing, documenting, or investigating:

1. Call `get_project_context`.
2. For feature-sized work, call `prepare_work` with the intent, likely
   paths, change kinds, and bootstrap version.
3. For small or exploratory work, call `resolve_context` with the paths.
4. Load the returned governing sections, critical rules, and scoped
   completion checklist.
5. Check for an existing task before creating one.
6. Refresh resolution whenever touched paths, contracts, or risks expand.

If `prepare_work` is missing because the MCP client has a stale tool
registry, reconnect it. Until then, fall back explicitly to
`get_project_context`, `resolve_context`, the governing sections, and the
scoped completion checklist. If Damper is unavailable entirely, stop
context-dependent work rather than silently relying on memory.

Do not commit, merge, deploy, or complete a task without the authority
defined by the project workflow.

That last fallback paragraph matters more than it looks. Agents degrade gracefully by default — they will happily proceed on memory and produce something plausible. Naming the failure explicitly, and naming stopping as the correct response, is the difference between a missing tool being an incident and a missing tool being a silent regression.

Keep MCP configuration free of secrets: commit the server declaration or use your editor's managed connection, keep credentials in environment or secret management, and gitignore local caches such as .damper/.

Part 2: Read the code before you write the context

This is the step people skip, and skipping it poisons everything downstream. Context must describe verified code, not the architecture you remember agreeing to.

Inventory the topology

Map applications, services, workers, and mobile clients; API entry points, routes, DTOs, and middleware; tables, migrations, and seeds; jobs, queues, webhooks, and integrations; shared libraries and generated clients; frontend routes, forms, state, and permissions; tests and fixtures; public documentation and pricing claims; deployment manifests, environment variables, and rollback paths.

A reasonable first pass:

rg --files -g '!node_modules/**' -g '!.git/**'
find . -maxdepth 2 -type d | sort
rg -n 'router|controller|service|schema|dto|migration|webhook|queue|cron'
rg -n 'openapi|generated|sdk|client|contract|compatibility'
rg -n 'permission|role|account|tenant|billing|tax|money|fiscal'
rg -n 'process\.env|import\.meta\.env|Secret|ConfigMap|Deployment'

Adapt the terms to your stack. You are building a map of ownership and dependencies, not a keyword report.

Trace features by identifier, not by folder

Folder names describe how the code was filed. Identifiers describe how it actually connects. For each important feature, pick several stable ones — an endpoint path, a DTO type, a table or column, an event name, a frontend route, a feature flag, a generated SDK type, a documentation heading — and search each across the whole repository:

rg -n 'CreateInvoice|create_invoice|POST /invoices'
rg -n 'invoice\.created|INVOICE_CREATED'
rg -n 'billing_enabled|BillingSettings'

Then follow each flow in both directions. For an API feature:

route -> request schema -> service -> database/integration
      -> event/job -> response schema -> generated clients
      -> web/mobile consumers -> docs/examples -> tests

For a frontend feature:

route -> page -> form/state -> API query or mutation
      -> shared UI -> localization -> analytics -> permissions
      -> browser tests and public claims

For infrastructure:

environment declaration -> application reader -> deployment manifest
                        -> secret/config owner -> health/alerting
                        -> rollout and rollback proof

Record the source of truth, the invariants, the downstream consumers, and the verification commands for each. If you cannot prove a relationship from code, label it uncertain rather than encoding it as a rule. An uncertain note invites checking. A confident wrong rule gets obeyed.

A small matrix keeps you from missing secondary surfaces:

Feature Canonical logic Inputs/contracts Persistence Consumers Jobs/integrations Tests Docs/public claims
Invoicing invoicing/service route + DTO invoices, invoice_items web + JS SDK + iOS invoice.created webhook unit + contract API guide + pricing page

Fill one row per important feature. Those columns become your sections, routes, bundles, and checklist scope.

Part 3: Structure the context

Two failure modes bracket this step: one context page per file, and one giant project manual. Aim between them.

Sections

Three kinds, roughly.

Baselines describe repository shape, ownership, and broad invariants — overview, specification, api/architecture, web/architecture, mobile, packages.

Domain sections describe behavior that differs materially from the baseline — api/invoicing, api/subscription-billing, web/document-workflows, integrations/<provider>, mobile/authentication. Create one when behavior, risk, ownership, or verification genuinely differs. Do not create a country-, provider-, or customer-specific section when the underlying rule is generic; you will maintain three copies of one rule and they will disagree within a quarter.

Quality gates state required outcomes and proof layers — quality-gates/api, quality-gates/web, quality-gates/document-rendering. State outcomes, not command inventories; commands belong in repository scripts, which can change without a context update.

Workflow sections carry cross-cutting agent rules — workflow/context-gathering, workflow/planning-execution, workflow/agent-communication. These belong in the mandatory bootstrap packet, not in a catch-all source-path route. A **/* route looks like an elegant way to make rules universal, and it quietly suppresses the fallback architecture context that paths would otherwise have resolved to.

A template that holds up:

# <Section title>

## Scope and ownership
What this governs — and what it deliberately does not.

## Source of truth
Canonical modules, schemas, tables, configuration, public contracts.

## End-to-end flow
How data and control actually move.

## Invariants
What must stay true, including compatibility and permission behavior.

## Consumers and related surfaces
SDKs, clients, jobs, integrations, docs, analytics, support.

## Verification
Nearest meaningful tests, plus broader checks for distinct risks.

## Known pitfalls
Confirmed failure modes, migration constraints, intentional exceptions.

Attach appliesTo (affected surfaces), tags (retrieval terms), and criticalRules (short must-follow statements surfaced automatically to agents).

Never put secrets, customer data, production logs, or large source copies in context. Context is durable knowledge, not evidence storage.

Routes

Routes turn paths into governing context. This is the mechanism that makes retrieval task-specific, and it is a real upsert_context_route payload. One ordering note before you write any: the surfaces, change kinds, and risk domains you use here are validated against your completion-checklist taxonomy, so define that first — see Part 4.

{
  "name": "public-api-contract-schemas",
  "patterns": ["apps/api/src/**/*.dto.ts", "apps/api/src/**/*.schema.ts"],
  "sections": ["api/openapi", "sdk-generation", "quality-gates/api"],
  "appliesTo": ["api", "docs", "js-sdk", "mobile", "web"],
  "changeKinds": ["contract_changes"],
  "riskDomains": ["public_api"],
  "priority": 40
}

Seven fields, and the whole invoice example above follows from them.

Let the audit pick your first routes. You do not have to guess which paths need routing. Once the bootstrap is in place and agents have done a little work, run audit_context_health and read topFallbackPaths — the paths that repeatedly resolved to nothing. That is an evidence-based, ranked list of exactly what to route next, and it beats intuition. Guess only where you have no telemetry yet.

Overlapping routes compose; they do not compete. Every matching route contributes, and their sections are unioned. priority orders the result — it does not filter, and a higher-priority route does not suppress a lower one. This matters because the warning above about catch-all routes reads, on its own, as "avoid overlap," and a reader who takes that too far will under-route. The strongest pattern is deliberate layering: one broad baseline route for a subtree, plus narrow routes stacked on top for the paths inside it that carry extra contracts or risk.

Prefer narrow verified routes over broad guesses. Keep patterns repository-relative and testable. Let routes overlap when two contexts genuinely both apply, and write down why in description. Treat structured routes as authoritative and graph-derived relationships as investigation leads until you verify and promote them. And do not add a universal route for workflow rules — see above.

Bundles

Routes answer what governs this path. Bundles answer what belongs together for this capability — API contract plus SDK generation; billing plus account scope plus permissions; document lifecycle plus rendering plus frontend workflow. Define entry patterns, representative paths, included sections, and verification profiles.

Keep bundles cohesive. A bundle represents one capability or contract. A bundle that represents "most of the backend" has become a second global manual, which is where you started.

Part 4: Make it executable

Everything so far is a set of claims about your repository. This part turns claims into tests.

Fixtures

Routes without fixtures drift silently — the code moves, the patterns stop matching, and retrieval quietly degrades into fallback while every packet still looks well-formed. A fixture asserts what should resolve for an important path: sections that must be included, sections that must not be, expected surfaces, change kinds, risk domains, routing source, and verification profiles.

Negative assertions do most of the work. A supplier-document path should pull in expense rules and should explicitly not pull in outgoing-invoice behavior. Positive assertions tell you the route fires; negative ones tell you it is not over-firing.

Two rules, both learned the hard way:

Never copy the resolver's current output into a fixture. That tests the implementation against itself and will pass forever, including through the regression you wrote it to catch. Derive assertions from verified project truth.

Route representation is not scenario coverage. One representative fixture under a broad glob proves the route exists. It does not prove that the materially different risks beneath that glob resolve correctly. Write a fixture per distinct risk, not per route.

Run validate_context_routing after every route or fixture change. It executes your fixtures against the production resolver and reports pass or fail with a count.

One honest tension: the health check counts fixture coverage per route, so its uncovered-context-routes finding goes quiet as soon as every route has one fixture. Taken literally that rewards exactly the representational fixtures this section warns against. Treat that finding as a floor, not a target — clearing it means every route is represented, not that every distinct risk is covered.

Scope the completion requirements

Define this taxonomy before you write routes. The two are coupled, and the failure is silent. Every route's appliesTo, changeKinds, and riskDomains is validated against the scopes the completion checklist declares. Use a value the checklist has never heard of and nothing breaks — the route saves, resolution still runs, and your risk model is decoration until someone reads the health findings and sees the taxonomy warnings piled up. If you have already written routes, this is the first thing to reconcile.

Maintain it with patch_completion_checklist, which merges the groups, items, and scope axes you name and leaves everything else alone. Reach for a whole-document replace only when you genuinely mean to replace the whole document — resending a full checklist to change one group is how groups quietly disappear.

One naming trap is worth knowing about in advance. The taxonomy lives in a section called completion-checklist, which is managed through the checklist APIs rather than the ordinary section APIs — it is deliberately hidden from the normal section listing, and writing to it directly is rejected. It is easy to create a plausible-looking human-facing section called something like checklist, document your scopes there, and conclude you are done. Nothing fails. Nothing is wired up either.

The goal is proof proportional to risk — not running everything, and not running only what is convenient.

Scope along three axes: surfaces (API, web, mobile, SDK, docs, packages), change kinds (behavior, contract, rendering, public output, migration), and risk domains (permissions, account scope, billing, money, tax, localization, public API, integrations, rendering).

Group the resulting checklist by outcome rather than by command — core proof, cross-surface synchronization, user journeys and rendered output, financial and compliance safety, schema and migration safety, public API integrity, context maintenance. Each item names acceptable proof types: unit, integration, contract, build, smoke, rendered output, example validation. It does not name a fixed command, because commands change and outcomes do not.

When you are unsure how to scope, scope broader. The cost of an unnecessary check is minutes. The cost of a missed one is the invoice field that vanished on iOS. Intentional skips are fine — they just have to be explicit and justified rather than implied by silence.

Part 5: Run the workflow

Session start

Read the bootstrap. Call get_project_context. Call prepare_work for feature-sized work or resolve_context for small exploration. Load the governing sections, critical rules, bundles, and scoped checklist. Search for an existing task before creating one — duplicate tasks are the most common and most annoying failure of agent-driven workflows. Then lock or start the exact task according to project policy.

Before editing

List the paths you expect to touch. Classify surfaces, change kinds, and risks. For API work, declare the impact explicitly: internal, behavioral, additive, breaking, or documentation-only. Settle shared contracts before anything parallel starts. Pick the nearest proof layer and write the expected behavior first. Decide what can be delegated without splitting ownership of a contract decision.

Then check that the task is still true. A ticket is a snapshot of what somebody believed on the day they wrote it, and the gap between that day and yours is where the capability may already have shipped, the endpoint may already exist, or the bug may already be fixed. Confirm the premise against the current surface — the live tool list, the deployed schema, the actual code — before you implement it.

This is the same discipline as tracing code before writing context, applied one step earlier, and it fails the same way. An agent handed a stale premise does not stall; it builds confident, well-tested, fully documented work that solves a problem nobody has, or adds a second way to do something the system already does. The result is worse than a wasted afternoon, because now the codebase has two answers to one question and both need maintaining. Verification costs a couple of minutes and occasionally deletes the entire task.

During implementation

Keep the diff scoped. Maintain one source of truth per domain rule. Refresh context resolution when paths or risks expand. Review changes that arrive by rebase, not just conflicts — a clean rebase can still invalidate an assumption your feature was built on. Update durable context when behavior, ownership, or verification changes. Record decisions and unresolved risks on the task, where the next person will find them.

Handoff

The implementor reports: delivery state and a one-sentence outcome; the pull request and its exact head; a complete surface and diff summary; checks run; API and public-documentation impact; risks, skips, and remaining work.

The operator independently verifies: final head, base, and synthetic merge identity; task and context-scope reconciliation; exact-head independent review; required checks and any valid exceptions; merge authorization; post-merge deployment observation. The task completes only when the merge fully delivers it — not when a prerequisite lands.

Communicate for decisions, not narration

Put the communication standard in mandatory workflow context so every agent loads it before its first update. A compact routine update:

Status: <precise stage and result>
Evidence: <new proof or material observation>
Next: <next meaningful action>
Blockers: none | <stable blocker ID>
Confidence: high | medium | low — <brief reason>

Separate facts from inference from recommendation from requested action. Use a stable ack-required ID for material production, security, data, compliance, public contract, or residual-risk findings. Repeat unresolved items at every meaningful update and at final handoff.

Never infer acknowledgement from silence. An agent that treats no response as approval will eventually treat no response as approval for something that mattered.

Part 6: Enforce readiness at the exact candidate

"All tests passed" is a claim about a conversation, not about a commit. Replace it with a trusted local command that:

  1. starts from clean current main;
  2. obtains the pull request's exact head and base;
  3. builds the exact synthetic merge candidate;
  4. computes required checks from changed paths and declared risk;
  5. runs focused and broader checks in bounded parallel stages;
  6. records checksummed evidence tied to head, base, merge, and tool revision;
  7. requires an independent review receipt against the exact final head;
  8. invalidates that evidence when identity or relevant runtime changes;
  9. permits merge only through a guarded command;
  10. cleans up the worktrees, databases, caches, and processes it created.

The separation of duties is the point:

  • the implementor runs targeted checks for the surfaces they touched;
  • the operator runs canonical exact-candidate readiness once;
  • the merge command verifies the recorded evidence rather than rerunning the world;
  • high-risk or contract-changing work activates deeper proof automatically.

Keep genuinely unique integrity checks in CI. Do not re-run expensive local checks on hosted runners without a distinct reason — but do make local evidence tamper-resistant enough that the merge guard can trust it. Evidence that can be hand-edited is not evidence.

Enforcing context in CI without a key

"Keep MCP configuration free of secrets" is easy advice for an editor and harder for a build. If CI is going to enforce routing, the obvious approach is to give the runner an API key — and that is exactly what a security review will push back on.

You do not need one. Export the routes, fixtures, sections, and checklist taxonomy into a snapshot committed alongside the code, then have CI run the real resolver against that snapshot rather than calling the hosted service. The validation entry point accepts injected routes, fixtures, sections, and checklist as arguments, so the code path under test is the production one — no mock, no drift between what CI checks and what agents actually get.

The properties are worth spelling out, because they are what makes this pattern worth the setup: no credential on the runner, no network dependency in the merge gate, identical enforcement logic, and a snapshot diff that makes context changes reviewable in the pull request like any other code. The tradeoff is that the snapshot must be regenerated when context changes — which is a feature, since it turns an invisible remote edit into a reviewable commit.

Part 7: Keep it true

A context layer is a maintained system. Left alone, it becomes a confident record of how the codebase used to work — which is worse than no context at all, because agents will believe it.

Require a context disposition on every material behavior, contract, rendering, migration, or high-risk change:

  • updated — durable knowledge changed;
  • verified_current — governing context was checked and still holds;
  • not_applicable — no durable context affected;
  • deferred — tracked as explicit maintenance debt.

Set verification windows by risk: short for money, tax, permissions, public contracts, and compliance; long for stable architecture.

Verification is self-attested, so it is only worth what you put into it. Marking a section verified records the claim you made; nothing re-reads the code on your behalf. Record the paths you actually checked — the audit will otherwise flag sections that claim a repository revision without any verified paths, which is the system telling you it cannot evaluate drift. Verifying without reading the code is lying with metadata, and it is worse than leaving the section stale, because a stale section still looks stale.

Expect your health score to drop when you start doing this properly. This is the most counterintuitive thing about the audit, and it catches teams out. The score is a simple deduction from findings, so a project with no routes can score well — not because retrieval is working, but because nothing is watching it fail. Every resolution quietly falls back, and silent fallbacks generate no findings. Add routes and fixtures and the failures become visible: fallback warnings appear, uncovered routes are counted, and the number goes down.

A high score with a high fallback count does not mean you are healthy. It means the audit cannot see your gaps yet. Read the fallback and coverage numbers before you read the score, and never optimise the score by declining to add routes — which is precisely what the number rewards if you take it at face value.

audit_context_health does most of the periodic auditing for you, returning a score plus the specific problems — routes without fixtures, weak or uncovered fixtures, missing baselines, route overlap, stale knowledge, and observed fallback usage. Read the fallback list closely: a path that repeatedly falls back is a path your routes claim to cover and do not. Watch also for structured routes that omit legacy governing context, sections whose routes changed after last verification, and graph-suggested relationships still awaiting promotion.

When the diff is done, prepare_completion compares the paths you actually touched against the prepare_work snapshot and reports drift and scope expansion. Scope expansion between plan and delivery is not misconduct — it is the normal condition of software work. It just needs to be visible, because it is exactly when the original proof scope stopped being sufficient.

Track fixture coverage of enabled routes, fallback rate and repeated fallback paths, stale sections by risk, the share of completed tasks with an explicit context disposition, PR rework caused by missing context, and time from task start to a correct implementation plan.

Do not optimize route or fixture counts. They are trivially gameable and measure effort rather than outcome. Optimize retrieval precision and the number of context-related mistakes that reach review.

Optional: close the product loop

Damper can also sit inside your product. This is independent of the agent context layer, and worth keeping architecturally separate.

An authenticated feedback widget should load only for authenticated users and supported tenants, map each brand to the correct project, compute identity verification server-side with HMAC, send only approved metadata such as plan or aggregate usage, handle script deduplication and tenant switching, and degrade safely when identity metadata is unavailable. Test the authenticated, unauthenticated, multi-brand, and cleanup paths. The widget secret never reaches the browser — the server returns only the derived user hash.

If a help centre and a blog share one content store, the flag that separates them is load-bearing and it is almost always permissive by default — a new article is support documentation until someone says otherwise. That default is reasonable for the tool and wrong for you the first time you publish marketing copy into it, because the failure is invisible from the authoring side: the post looks fine, and meanwhile a customer asking your support bot a question gets answered with a launch announcement.

Set the flag explicitly at every level that carries it, even where one level already cascades. Group-level visibility that covers every post inside it is the right primitive and it is one setting away from failing — move an article out of the group, or unhide the group for an afternoon, and the protection is gone with no error anywhere. Belt and braces costs nothing here.

Public roadmap and changelog APIs can generate static documentation pages. Sanitize remote text, escape for the target markup format, filter by public state, and mark generated files clearly. Do not fetch mutable roadmap or changelog data inside deterministic PR validation; refresh those snapshots in an explicit publication workflow, or you have made your merge gate depend on someone editing a roadmap card.

A staged rollout

Phase 1 — Bootstrap. Connect Damper MCP. Add the short agent files. Inventory your surfaces. Write overview, architecture, testing, and workflow sections. Define task visibility and approval boundaries.

Phase 2 — High-value features. Trace your most important end-to-end flows. Write domain sections and quality gates. Add narrow routes for critical paths. Introduce bundles for cross-surface contracts.

Phase 3 — Executable context. Add positive and negative fixtures. Run validate_context_routing after every configuration change. Configure scoped completion groups and context maintenance. Start measuring fallbacks and staleness.

Phase 4 — Trusted delivery. Add exact-candidate local readiness. Require structured independent review. Guard merge and deployment. Record post-merge observation and task disposition.

Phase 5 — Optional product loop. Authenticated feedback capture, feedback linked to tasks, roadmap and changelog published through explicit workflows.

Most teams get the majority of the benefit from phases 1 and 2. Phases 3 and 4 are what keep it from decaying.

Common ways this goes wrong

  • Writing context before tracing code. You encode assumptions as rules, and agents obey them.
  • One giant context section. Retrieval gets noisy and updates get unsafe.
  • Routing by folder. Contracts and features cross packages and generated consumers; folders do not follow them.
  • Catch-all routes for workflow rules. They suppress the technical fallback context those paths needed.
  • Writing routes before defining the checklist taxonomy. Nothing fails. The warnings just accumulate and your risk model is decoration until someone notices.
  • Implementing a task's premise without checking it. Tickets go stale. Verify against the current surface first, or you will carefully build a second way to do something that already works.
  • Treating route representation as scenario coverage. A broad glob needs a fixture per distinct risk.
  • Copying resolver output into fixtures. Tests the implementation against itself. Always passes. Catches nothing.
  • Putting secrets or customer evidence in shared context. Context stores durable, non-sensitive knowledge only.
  • Running every test twice on every PR. Use risk-based targeted proof plus one trusted exact-candidate gate.
  • Completing a task when only a prerequisite merged. Delivery state stops meaning anything.
  • Ignoring rebased changes because nothing conflicted. No conflict is not the same as no impact.

The operating principle

The best agent workflow is not the one with the most instructions.

It is the one that retrieves the smallest complete set of verified project truth for the work at hand, converts that truth into scoped proof requirements, and refuses delivery when the evidence no longer matches the exact candidate.

Start small. Bootstrap files, a handful of source-grounded sections, narrow routes, representative fixtures. Add depth where mistakes are expensive — which you already know, because those are the areas you review most nervously. Over time it becomes a living map connecting code ownership, product behavior, risk, proof, and delivery.

The first milestone is not "map the whole repository." It is much smaller than that:

Connect one repository. Trace one important feature. Look at the work packet before an agent edits anything.

If that packet raises a question your normal process would only have found in review — or in production — the context layer is already paying for itself.

Try it on one repository

Connect one repository, trace one important feature, and inspect the work packet before an agent edits your code.