The agentic coding pattern landscape
118 named patterns, each with when not to use it. Start from the grid, or search.
The field at a glance
How many entries fall in each category and each maturity. Choose a cell to filter the list below, and choose it again to clear the filter.
| Category | foundational | established | emerging | All |
|---|---|---|---|---|
| methodology | ||||
| context-management | ||||
| verification | ||||
| control-loop | ||||
| multi-agent | ||||
| academic | ||||
| planning | ||||
| governance | ||||
| All | All categories, all:118 |
Browse the library
118 of 118
Patterns
-
- Category: verification
- Maturity: emerging
Point an agent at your checks and tell it to cheat. A second agent patches every exploit it finds, a third confirms legitimate solutions still pass, and you repeat. On one benchmark this took attack success from 62% to 0%.
-
- Category: multi-agent
- Maturity: established
A handoff transfers control. It is not a function call. A triage agent routes the conversation to a specialist, and the specialist becomes the active agent for the rest of the turn. A manager, by contrast, calls specialists as tools and keeps control.
-
- Category: context-management
- Maturity: established
Write the exact build and test commands, in order, with expected results, into the instructions file the agent reads every session. In a large real repository this single change tracked a jump in pull request success from 38.1% to 69%.
-
- Category: methodology
- Maturity: established
Agent OS keeps durable standards and feature context in a small Markdown tree, then injects only the standards that match the task. It leaves planning and execution to the host agent.
-
- Category: multi-agent
- Maturity: emerging
Peers, not workers. Each teammate is a full independent session. It claims work from a shared task list and messages other teammates directly, instead of reporting to a parent that does all the coordinating.
-
- Category: academic
- Maturity: established
An agent-computer interface is the tool surface designed for a model: bounded reads, precise edits, useful errors, and safe defaults. The interface becomes part of the agent’s capability.
-
- Category: methodology
- Maturity: emerging
Planner, Manager, Workers. State lives in files outside any context window. Workers accumulate domain expertise across assignments instead of restarting fresh. A handover command transfers a worker's knowledge to a new instance when it runs out of room.
-
- Category: academic
- Maturity: established
Agentless fixes repository issues in stages. It narrows the search space, generates several patches, filters them with tests, and selects a candidate. The simple pipeline also makes a useful baseline for agentic systems.
-
- Category: methodology
- Maturity: established
AGENTS.md gives an agent the few repository facts it cannot cheaply infer. Keep it short, testable, and specific. Treat it as context that needs evaluation, not a policy engine.
-
- Category: methodology
- Maturity: established
Aider’s architect mode separates reasoning from edit formatting. One model explains the change. Another emits the file edits. The split helps when a strong reasoner is a clumsy editor.
-
- Category: multi-agent
- Maturity: established
The architect-editor split routes a task through two roles: a reasoning-heavy planner and an edit-reliable implementer. It can improve accuracy. But the handoff loses information, so verify it.
-
- Category: academic
- Maturity: established
Best-of-N sampling generates several candidates, then picks one with tests, agreement, a reward model, or a judge. You pay for extra calls and get robustness. The selector matters as much as the sampler.
-
- Category: methodology
- Maturity: established
BMAD is a role-based, artifact-heavy workflow. Research and product framing become a PRD, architecture, epics, and story files before implementation. Its value is traceability when the work is large enough to justify the ceremony.
-
- Category: control-loop
- Maturity: established
Checkpoints make agent work reversible. Record a known-good state, try a bounded change, and verify it. If the experiment fails, restore the checkpoint. A checkpoint helps only when its scope and restoration command are explicit.
-
- Category: control-loop
- Maturity: established
A CI repair loop treats a failing check as structured feedback. Reproduce it, make one bounded change, and rerun the relevant check. Stop when the evidence says green or the attempt budget runs out.
-
- Category: control-loop
- Maturity: established
Bound the loop by something the agent does not control: iterations, turns, dollars, wall clock, or lack of progress. It is the backstop. It makes every other exit condition survivable when that condition fails.
-
- Category: methodology
- Maturity: established
Cline’s Memory Bank keeps project purpose, architecture, technology, active work, and progress in durable Markdown files. It reduces context loss across sessions. The cost is maintenance and stale notes.
-
- Category: governance
- Maturity: established
A deterministic program parses each tool call into typed effects and blocks the ones it can prove are disasters. No model is involved, so the verdict is fast and identical every run.
-
- Category: multi-agent
- Maturity: emerging
Spawn one agent per theory about a bug and tell them to disprove each other. The goal is to defeat anchoring. A single investigator finds one plausible explanation and stops looking.
-
- Category: control-loop
- Maturity: established
The loop ends when the agent emits an exact string. It is trivial to implement and trivially gameable, because the model controls the string. Where a data predicate exists, use it instead.
-
- Category: methodology
- Maturity: established
Compound Engineering turns each feature into a loop of plan, work, review, and capture. The compounding step writes reusable lessons back into the project. Later tasks spend less time rediscovering the same constraints.
-
- Category: methodology
- Maturity: emerging
Google’s Conductor extension for Gemini CLI keeps project context, track specifications, plans, implementation, review, and revert in repository files. The workflow is context-driven, with explicit human checkpoints.
-
- Category: multi-agent
- Maturity: emerging
Try the cheap path first. Escalate to the expensive one—more agents, a stronger model, a human—only when a confidence signal says the cheap answer is not good enough. Most tasks never escalate. Average cost drops, and the hard tasks do not suffer.
-
- Category: verification
- Maturity: emerging
Put non-negotiable security invariants in a versioned, machine-readable constitution that constrains generation. Do not wait to check for violations afterwards. One study reports a 73% reduction in security defects with velocity preserved.
-
- Category: context-management
- Maturity: emerging
Keep governance constraints out of the compactable part of the context. A rule that gets summarized away stops being obeyed, and the agent gives no sign that anything changed.
-
- Category: context-management
- Maturity: established
Reduce context when it helps the next decision, and keep a way to recover the evidence. Summarizing, masking old tool output, and restarting from a handoff are different interventions. None is a universal winner.
-
- Category: context-management
- Maturity: established
The window's contents are the only thing you control. Managing them deliberately is the whole job. Every context-management pattern in this library is an instance of it.
-
- Category: context-management
- Maturity: established
A context file carries durable repository facts an agent cannot cheaply infer: commands, conventions, gotchas, boundaries, and pointers to deeper material. Keep it small enough to be read and specific enough to change behavior.
-
- Category: context-management
- Maturity: established
Throw the context window away on purpose. Start a fresh agent that carries only a structured handoff. This is the primitive underneath Ralph, shift handoffs, and relay chains. Anthropic found it necessary because compaction alone was not enough.
-
- Category: control-loop
- Maturity: emerging
Convergence Loop turns a written desired state into a repeated gap check. Assess. Implement only the delta. Run checks. Stop when the gap inventory is empty or the bounded loop stalls.
-
- Category: governance
- Maturity: emerging
Agents write, test, and ship production code. No human writes or reviews the code. Oversight moves entirely to specifications, acceptance scenarios, and production telemetry. The factory inherits the quality of its test oracle, exactly.
-
- Category: governance
- Maturity: established
An anti-pattern. Subagents inherit their parent's credentials and permission mode, but not the parent's intent. Authority widens silently with every link. One approval at the top authorizes actions nobody evaluated three levels down.
-
- Category: context-management
- Maturity: emerging
Documentation-Driven Handoff makes research and plans durable interfaces between people, sessions, and agents. The file is the handoff. The chat is only where you negotiated it.
-
- Category: verification
- Maturity: established
Evaluator-Optimizer separates making a candidate from judging it. The evaluator needs a rubric and evidence. The optimizer gets bounded chances to respond.
-
- Category: academic
- Maturity: foundational
Execution Feedback Repair lets the compiler, test runner, or program report what actually happened. The agent then uses that verbatim signal to guide the next repair.
-
- Category: planning
- Maturity: established
Explore Plan Code Commit is a gate. Understand the code. Write the plan. Implement against it. Commit only after the named checks pass.
-
- Category: planning
- Maturity: emerging
Replace the spec document with a list of atomic claims. Each claim can carry a shell command that proves it. Project status stops being something you assert. It becomes something you run.
-
- Category: academic
- Maturity: established
Fault Localization First treats “where is the bug?” as a deliverable. It comes before “what patch should I write?” Reproduction, search, and coverage narrow the edit surface.
-
- Category: context-management
- Maturity: established
One session per change, like one branch per change. At the end, have the model write a summary document and the opening prompt for the next session. What survives is curated, not accumulated.
-
- Category: control-loop
- Maturity: emerging
A chain of fresh agents. Each works only while its context is still cheap. At the token budget, it writes a deliberately tiny handoff and passes the baton. The next agent reads the plan from disk and continues.
-
- Category: verification
- Maturity: established
Fresh-Context Reviewer hands a change to a reviewer who has not seen the implementation conversation. Review is separate from the author’s explanation, but the reviewer still has access to requirements, callers, and tests. Whether this separation improves defect detection needs evidence.
-
- Category: academic
- Maturity: emerging
Treat the agent's memory like a git repository: COMMIT milestones, BRANCH to explore an alternative, MERGE what worked, and retrieve history hierarchically. Reported over 80% on SWE-bench Verified.
-
- Category: control-loop
- Maturity: established
You state a success condition. A separate evaluator re-checks it after every turn. The agent keeps working until the condition resolves or a try cap fires. The actor does not decide when it is finished.
-
- Category: methodology
- Maturity: established
GSD is a compact spec-to-task workflow. It turns a vague feature into written intent, an executable plan, and a sequence of small verified changes.
-
- Category: methodology
- Maturity: established
Gstack packages a product sprint as specialist skills: clarify, plan, build, review, test in a browser, ship, and reflect. Artifacts connect the stages.
-
- Category: verification
- Maturity: established
Hooks and Guardrails move repeatable safety rules from memory into lifecycle events. A hook should make the safe path automatic and the unsafe path observable.
-
- Category: control-loop
- Maturity: established
Human-Gated Autonomy puts people at decisions whose consequences are hard to reverse, while permissions and sandboxes handle routine work inside the boundary.
-
- Category: control-loop
- Maturity: emerging
A one-time agent run that builds the scaffolding every later run depends on: a launch script, a machine-readable feature list, a progress file, and a first commit. It writes no features itself.
-
- Category: planning
- Maturity: emerging
Treat the specification as a source language, not a prompt. Write it in a structured, declarative form and compile it deterministically into code and tests. The unresolved problem is that the compiler is a language model, and a compiler has to be deterministic.
-
- Category: context-management
- Maturity: established
After each verified phase, fold the current status back into the plan file. Don't let it pile up in the conversation. The plan becomes a living document, and that makes the work resumable.
-
- Category: control-loop
- Maturity: established
Interactive Steering is the short feedback loop around an agent’s turn: interrupt a wrong direction early, queue precise context, and restart after repeated corrections.
-
- Category: planning
- Maturity: established
Before planning, have the agent interview you with structured questions until the hard parts are covered, then write a self-contained spec. Implement it in a fresh session.
-
- Category: context-management
- Maturity: emerging
Replace the markdown task list with a real issue tracker built for agents: a dependency graph in a versioned database, where ready returns work with no open blockers and finishing something unblocks whatever waited on it.
-
- Category: governance
- Maturity: emerging
A hash-chained ledger. For every agent action it records what the agent was supposed to do, what it did, and how far apart those were. It is deterministic, not model-judged, and tamper-evident by construction.
-
- Category: methodology
- Maturity: established
Kiro Spec Workflow turns a prompt into requirements, design, and dependency-ordered tasks, with review gates and synchronization checks connecting the documents to code.
-
- Category: verification
- Maturity: established
A model grades an artifact against a rubric and returns a score. No revision loop is attached. It is useful where nothing executable exists. It is demonstrably biased by surface properties that have nothing to do with correctness.
-
- Category: control-loop
- Maturity: established
Not a loop—a taxonomy of them. Classify a loop by what triggers it and what stops it. The four resulting shapes (turn-based, goal-based, time-based, proactive) each have a different correct use.
-
- Category: context-management
- Maturity: established
Memory Bank writes down the state a fresh agent cannot cheaply infer: current focus, decisions, constraints, and unfinished work. Keep it small enough to read at startup.
-
- Category: multi-agent
- Maturity: emerging
Parallel agents produce branches. One serializing integrator merges them in order, one at a time, and runs the checks after each merge. Generate in parallel. Integrate in series.
-
- Category: academic
- Maturity: emerging
Multi-Agent Debate makes disagreement an input to selection or localization. It earns its cost only when perspectives differ, rounds are bounded, and a check can arbitrate.
-
- Category: methodology
- Maturity: emerging
Spec Kit split across PM and developer agents. Read-only probing hooks ground each phase in repository evidence before it produces its artifact. Reported gains are small but statistically significant.
-
- Category: verification
- Maturity: emerging
A second model gets a clean view of the diff and the contract. The goal is to expose blind spots, not to collect a pile of opinions.
-
- Category: verification
- Maturity: established
Several reviewers run in parallel, each looking for a different class of issue. A verification step filters false positives, and a synthesizer deduplicates and ranks what survives. One reviewer serializes and truncates. A fan-out does not.
-
- Category: methodology
- Maturity: established
Keep the current behavior in living specs, describe one change as a delta, implement it, verify it, then archive the delta back into the source of truth.
-
- Category: multi-agent
- Maturity: established
One lead agent owns decomposition and synthesis. Workers handle narrow, isolated investigations. The lead picks the next wave from what it learns.
-
- Category: multi-agent
- Maturity: established
Give each agent its own worktree and branch. Integrate the results through ordinary Git review. Isolation keeps simultaneous edits from tangling one working tree.
-
- Category: governance
- Maturity: established
A second, smaller model judges each proposed action. The action runs, goes to a human, or is blocked. The classifier fills the autonomous region that permission tiers leave undefined. It is also an attack surface, with a published false-negative rate.
-
- Category: governance
- Maturity: established
Sort every action an agent can take into a few tiers by consequence and reversibility. Attach a fixed policy to each tier. Approval becomes a property of the action, not a judgment call made under fatigue.
-
- Category: planning
- Maturity: established
A runtime mode where the agent may read and explore but may not edit until a human approves its plan. It is Plan-and-Execute with the separation enforced by the tool, not the prompt.
-
- Category: planning
- Maturity: established
Plan-and-execute separates deciding what to do from carrying it out. The plan is a hypothesis that execution can revise when the repository disagrees.
-
- Category: governance
- Maturity: emerging
Fix the plan before the agent reads any untrusted content. Injected text can then influence values inside a predefined execution graph. It cannot redefine the task or invent new actions.
-
- Category: multi-agent
- Maturity: established
Three roles, not two. One expands intent into a specification. One implements. A third, separate from the implementer, decides whether the work is done. The separation exists because agents confidently praise their own output.
-
- Category: context-management
- Maturity: established
One append-only file on disk records what was done, what was tried and failed, and what comes next. The agent reads it first and updates it last, so a fresh context can pick up where the last one stopped.
-
- Category: context-management
- Maturity: established
Keep a one-line description of every capability in context. Load the full content only when it is used. The index is cheap. The contents are not.
-
- Category: control-loop
- Maturity: established
The prompt is a maintained source file, not a message. You watch the agent fail, add a "sign" at the point of failure, and commit. No prompt is perfect. Some have been tuned against observed behavior.
-
- Category: methodology
- Maturity: established
A PRP turns a feature request into a compact packet: requirements, repository examples, known traps, and executable checks. It pays off when it cuts discovery during implementation.
-
- Category: control-loop
- Maturity: established
A Ralph loop gives an agent the current task state, lets it act, and checks an external condition. Then it starts another bounded iteration. It stops when the condition passes or the budget ends.
-
- Category: methodology
- Maturity: emerging
Plan with BMAD Method interactively, then hand the resulting stories to a Ralph Loop that implements them autonomously, one story per iteration, test-first, committing as it goes.
-
- Category: control-loop
- Maturity: foundational
A ReAct-style agent chooses an action, runs a tool, reads the result, and chooses again. It is useful when the next step depends on what you discover. I would pair the loop with an external success check and a bounded failure path. A final answer tells you the agent stopped. It does not tell you the task succeeded.
-
- Category: academic
- Maturity: established
Reflexion stores an agent's verbal critique of a failed attempt and feeds that lesson into the next attempt. The memory helps only when it changes the next action.
-
- Category: context-management
- Maturity: established
A repository map gives the agent a compact index of files, symbols, and relationships. The agent gets it before it spends context reading implementation.
-
- Category: context-management
- Maturity: established
Research, planning, and implementation form a relay. First learn what the repository and sources say. Then write an executable plan. Then make and verify the change.
-
- Category: methodology
- Maturity: emerging
Spec-driven development pointed backwards. A multi-agent pipeline reads a legacy system and produces traceable operational specifications. It marks each claim with a confidence level and records the gaps it could not resolve.
-
- Category: verification
- Maturity: established
These defenses keep an agent from satisfying the check instead of the task. Freeze the tests. Hold some out. Cap the achievable score, so a perfect one is evidence of cheating. Monitor test edits. Enforce all of it with hooks, not instructions.
-
- Category: multi-agent
- Maturity: established
A role pipeline gives each phase a narrow job—research, plan, implement, review, verify—and hands explicit artifacts between phases.
-
- Category: methodology
- Maturity: established
Boomerang tasks let a parent agent delegate a focused child task, receive a result, and continue with the returned context. The parent stays responsible for the overall outcome.
-
- Category: methodology
- Maturity: emerging
A memory bank plus five specialized modes for the Roo Code extension, with prompts written in YAML rather than Markdown to cut token consumption.
-
- Category: verification
- Maturity: established
Put the agent behind a filesystem and network boundary that fits the task. Keep credentials and valuable host state outside it.
-
- Category: verification
- Maturity: established
For visual work, a screenshot is a test artifact. Render the page, inspect the image and console, make a bounded correction, and repeat against a concrete target.
-
- Category: academic
- Maturity: established
Self-refinement alternates generation, critique, and revision in one model context. It improves an artifact when the feedback is specific. Execution checks or an independent reviewer beat self-agreement.
-
- Category: context-management
- Maturity: established
A shift-worker handoff makes a long task resumable. It records the current state, decisions, failures, verification evidence, and next bounded action. An honest failed checkpoint is useful. A clean-looking but unverified handoff is not.
-
- Category: academic
- Maturity: established
A skill library packages discoverable instructions with optional scripts, templates, and references. Progressive loading gives the agent specialist behavior without putting every workflow in every context.
-
- Category: context-management
- Maturity: established
A slash command gives a repeatable workflow a named entry point, arguments, and an explicit permission boundary. It is a task interface whose implementation you can inspect.
-
- Category: verification
- Maturity: emerging
Every session starts the application and runs an end-to-end check before it touches new work. The check catches breakage the previous session did not know it caused.
-
- Category: planning
- Maturity: emerging
Team-scale agent development in four stages: product requirements, architecture, program design, then vertical slices. Humans keep reviewing. Models optimize for passing tests, and nothing penalizes a decaying codebase.
-
- Category: methodology
- Maturity: established
Spec Kit turns intent into linked Markdown artifacts—specification, plan, tasks, and implementation—through a portable command workflow. The documents organize the work. Tests still establish behavior.
-
- Category: methodology
- Maturity: emerging
Spec-driven development with the missing runtime attached: repository-native mission state, Kanban lanes, a worktree per work package, and an explicit next -> review -> accept -> merge loop.
-
- Category: methodology
- Maturity: emerging
Spec-driven development treats the specification as the durable source of intent. Plans, tasks, code, and tests derive from it. The useful question is whether the spec stays aligned after implementation.
-
- Category: methodology
- Maturity: emerging
A Claude Code workflow toolkit that runs spec → plan → tasks → implement → optimize → ship. It has tiered quality gates and specialized role agents. An academic assessment scored it highest of every framework it examined. It has only 85 GitHub stars.
-
- Category: planning
- Maturity: emerging
A spec is a numbered list of requirements. Each has a stable identifier. Code and tests that satisfy a requirement reference it. Coverage becomes a query, not a judgment.
-
- Category: control-loop
- Maturity: established
A Stop hook intercepts the agent's attempt to end its turn, exits non-zero, and feeds the same prompt back. The loop runs inside one session, so context accumulates instead of resetting.
-
- Category: context-management
- Maturity: established
Subagent context isolation gives a focused worker its own prompt, context, model, and sometimes worktree. It keeps research or review from bloating or biasing the main session.
-
- Category: methodology
- Maturity: established
Superpowers makes planning, test-first implementation, review, and debugging explicit through reusable skills. The gates improve repeatability, while each gate adds time and context overhead.
-
- Category: methodology
- Maturity: foundational
SWE-agent style treats issue fixing as a loop over a sandbox, a model-facing computer interface, and tests that decide whether the patch works.
-
- Category: methodology
- Maturity: established
Parse a product requirements document into a dependency-aware task graph, then serve that graph to a coding agent over MCP so next always returns something unblocked.
-
- Category: multi-agent
- Maturity: established
A task queue consumer claims one durable task, performs bounded work, persists the result, and acknowledges only after success. Leases, retries, and queue-age metrics make crashes visible.
-
- Category: methodology
- Maturity: emerging
Spec-as-source taken literally: the specification is the maintained artifact, and the code carries a header saying it was generated and must not be edited. It is the clearest statement of the position. The most interesting fact about it is that Tessl's own public product is something else.
-
- Category: verification
- Maturity: established
Quality measures move in one direction only. Tests may be added, never removed or weakened. A feature's status may flip from failing to passing, and nothing else about it may change.
-
- Category: verification
- Maturity: established
Test tampering is when an agent changes the test, fixture, or check instead of fixing the behavior. Protected tests and independent verification keep “green” meaningful.
-
- Category: verification
- Maturity: established
A test-driven agent loop writes and observes a failing behavior test before implementation, freezes the test during the green phase, and refactors only while it stays green.
-
- Category: context-management
- Maturity: emerging
Write large tool outputs to disk and keep only a reference in context. The agent reads the parts it needs with the tools it already has. It does not carry the whole output in every later request.
-
- Category: multi-agent
- Maturity: established
Assign a tracker issue to an agent the way you would assign it to a person. The agent plans, opens a pull request, writes the code, runs the tests, and asks for review. Someone other than the issue's author must approve it.
-
- Category: context-management
- Maturity: emerging
Trajectory memory distills prior attempts into reusable lessons, insights, and workflows. It should record why an approach worked or failed. It should not become an unreadable transcript archive.
-
- Category: academic
- Maturity: emerging
Tree search over actions expands several candidate next actions, scores the resulting states, backtracks from weak branches, and commits the best surviving trajectory. More search costs more compute.
-
- Category: planning
- Maturity: emerging
Planning moves off your machine. A cloud session drafts the plan while your terminal stays free. You review and comment on it section by section in a browser. Then you execute it remotely or send it back to the CLI.
-
- Category: verification
- Maturity: established
A verification loop turns "done" into layered evidence: targeted checks, broader checks, and inspection of the result. A failure feeds diagnosis and another bounded change.
-
- Category: verification
- Maturity: emerging
Spec-driven development with a conformance step. The spec produces tests and formal properties before any code exists, and generated code must satisfy them. The hard part: specifying an unsolved problem is harder than testing a solved one.
-
- Category: methodology
- Maturity: established
Vibe coding means directing an agent in conversation while accepting generated code with little inspection. It suits disposable experiments. Durable software still needs code ownership and verification.
Open a folder of notes
Load your own library to browse it the same way. Choose a folder or drop one on the page, or
choose several .md files. The notes are read in your browser by the same code that built the bundled library, and they
replace it until you return. Nothing is sent anywhere.
Drop a folder of notes here
Files are read in your browser. Nothing is uploaded.
type to be included. Sections are found by their headings, such as TL;DR and When To Use It.A comma list. A note whose frontmatter type isn't on it is excluded, and the diagnostics say so.