5K

Agent design patterns

An open catalogue of 203 patterns for building AI agents. Each pattern is a plain Markdown file: read it here, get it as .md for your agent, or propose your own.

203 patterns

  1. 402-First Machine Payments (Price-Before-Work Tool Purchases) Server returns an HTTP 402 price quote before doing work, and the agent pays the exact amount only if it fits its budget cap Orchestration & Control emerging
  2. Classify-Then-Act for Background Agents Sorts each task into has-default, needs-intent, or out-of-scope, builds only has-default work in a sandbox, and queues it for human review Orchestration & Control emerging
  3. Design-Time File Partition as Concurrency Control Each task declares every path it will write; overlapping tasks get a blocking dependency so only disjoint work runs at once on a shared checkout Orchestration & Control emerging
  4. Evidence Questions, Not Verdict Questions Never ask the model for allow/ask/block. Ask it narrow, typed, observational questions and let code reduce the answers to the verdict, because the collapsed question is the low-confidence one. Security & Safety emerging
  5. Non-Generative Judgment Routing with Typed Escalation Sends decision-only loop steps as batched typed questions to a non-generative judgment model and escalates what it cannot answer back to the LLM Orchestration & Control emerging
  6. One OS User per Agent Give each long-lived agent its own OS user account and a templated init-system unit, so identity, supervision and privilege come from the host instead of from a per-agent container. Security & Safety validated in production
  7. Out-of-Process Provider-Boundary Replay Capture an agent run at the model-provider HTTP boundary from outside the process, then serve those bytes back so the agent re-executes its own logic with no provider call. Reliability & Eval emerging
  8. Precomputed Code Graph Lookup Answer an agent's structural code questions from a precomputed, deterministic graph queried by intent-shaped verbs, instead of a retrieval-time search-and-read loop. Tool Use & Environment emerging
  9. Skill Activation as a Precision/Recall Measurement Measures skill routing as classification with a labelled prompt dataset, per-skill precision and recall, and CI thresholds that fail the build Reliability & Eval emerging
  10. Tier Auto-Apply by Mechanical Impact Decides which self-proposed changes auto-apply by matching the real diff's file paths and change kinds against a trusted tier table, not by finding text Reliability & Eval emerging
  11. Cross-Domain Agent Conflict Resolution A coordination layer that cross-references recommendations from independent domain agents, detects conflicts on shared resources, and resolves them through policy-as-code. Orchestration & Control emerging
  12. Deterministic Grader in the Loop Replace the LLM-as-judge in an agent's self-review loop with a deterministic scorer that returns the raw counts behind each sub-score, so the agent has something specific to fix. Feedback Loops emerging
  13. Dual-Rail Message Delivery Sends every inter-agent message on two independent channels through one entry point, alerts when the rails diverge, and requires ACKs to confirm completion Orchestration & Control validated in production
  14. Evidence-Layered Evaluation for Interactive Agents Scores interactive agent runs on a task outcome assertion and links each result to layered evidence: actions, screenshots, replays, network, and messages Reliability & Eval emerging
  15. Exact-Action Authorization Binding Bind a short-lived approval to the complete action presented to the reviewer, then rederive and compare that identity at the acting boundary before execution. Security & Safety proposed
  16. Filesystem-Mediated Host Delegation Lets a sandboxed agent write request files to a shared spool directory that a host daemon runs against a whitelist, with idempotent results the agent polls for Tool Use & Environment emerging
  17. Own-Check Fault Injection Plant a controlled fault inside a running agent pipeline and score whether the pipeline's own checks emit a detection act, keeping detected, reacted and recovered as separate verdicts. Reliability & Eval emerging
  18. Rendered UI Finish Gate Verify an agent-built interface against its intended design contract, required states, interaction semantics, and rendered output before merging. Feedback Loops emerging
  19. Persistent Test Memory Feedback Loop Preserve validated test lessons so an agent can reuse successful paths, recognize recurring failures, and retire stale knowledge. Learning & Adaptation emerging
  20. Consequence-Family Coverage Audit Audit an agent's risk policy by enumerating families of consequence and asking which rule covers each, because a hand-written risk list reliably encodes one family and stays silent on the rest Security & Safety emerging
  21. Reasoning-Token Firewall Assemble an agent's result only from answer-typed stream events, never by string-stripping interleaved reasoning tokens. Reliability & Eval validated in production
  22. Tracker-as-Desired-State Reconciliation Derive agent dispatch by re-reading workflow states in the team's issue tracker every tick, keeping only a short-lived claim locally, so missed events, restarts, and human edits all self-heal. Orchestration & Control emerging
  23. Authenticated Authority Channel Preserve a distinguishable channel for authenticated intent so retrieved content can inform reasoning without granting itself authority. Security & Safety emerging
  24. Agent-First Tool Discovery Build search indexes designed for agent consumers, returning structured tool metadata ranked by agent-relevant signals instead of human SEO metrics. Tool Use & Environment emerging
  25. Artifact-Driven Analysis Pipeline Orchestration Has an agent run independent analysis scripts, read their structured reports, and merge them into one final report or visualization Orchestration & Control emerging
  26. Black-Box Skill Invocation Shares skills through input and output schemas only and runs them on the provider side, so prompts and code never cross the boundary Security & Safety emerging
  27. Board-Mediated Async Inter-Agent Coordination Routes agent-to-agent messages through comment threads on board cards, with a prefix that creates an inbox card to notify the target agent Orchestration & Control validated in production
  28. Budget-Aware Model Routing with Hard Cost Caps Routes each request to the cheapest model that meets its needs under hard cost caps, and escalates only when quality gates fail Orchestration & Control established
  29. Capability-Escrow-Receipt Binds a signed capability listing, an atomic escrow-plus-hire step, and a signed work receipt into one loop for agent-to-agent payment Orchestration & Control experimental but awesome
  30. Context Budget as a Governed Resource Measures always-loaded context per source, alerts on trend growth, and puts hard spend caps and fan-out bounds on unattended agent runs Context & Memory validated in production
  31. Cross-Agent Lesson Sharing via Git Agents write solved problems as markdown lessons in a shared Git repo and search them before debugging, with GitHub Issues as the coordination layer Context & Memory validated in production
  32. Cryptographic Governance Audit Trail Middleware checks each tool call against a policy file, then signs a receipt of the action with ML-DSA and appends it to a tamper-evident chain Security & Safety emerging
  33. Dead-Man's Switch for Scheduled Agent Jobs Each scheduled job writes a success sentinel to its log, and a separate checker files a task when the sentinel is missing or stale Reliability & Eval validated in production
  34. Declarative Multi-Agent Topology Definition Define multi-agent systems declaratively in a single topology file — agents, flows, gates, hooks, group chats — then compile to platform-specific configurations for any agentic framework. Orchestration & Control emerging
  35. Deterministic Threat Rule Scanning Apply deterministic regex rules as a first-pass security layer to detect known threat patterns in AI agent tool calls and skill definitions. Security & Safety emerging
  36. Deterministic Zero-LLM Orchestration A deterministic code orchestrator splits goals, runs parallel coding agents, verifies with tests, and commits, spending no LLM tokens on coordination Orchestration & Control validated in production
  37. Egress Lockdown (No-Exfiltration Channel) Puts the agent behind a default-deny egress firewall with allowlisted destinations so stolen data has no outbound channel Tool Use & Environment established
  38. Local-First Credential Broker Keep raw secrets out of the agent process by injecting credentials at the network layer through a local broker, rather than handing the agent environment variables or config files. Security & Safety emerging
  39. Markdown Polis — Multi-Vendor Agent Coordination via Filesystem Constitution Coordinates agents from different vendors through versioned markdown files: capability cards, work contracts, and bandit routing that learns from settled work Orchestration & Control emerging
  40. MCP Pattern Injection Runs an MCP server that exposes framework best-practice patterns as tools, so the coding assistant fetches current patterns on demand Tool Use & Environment validated in production
  41. Output Verification Loop Verify LLM outputs by extracting individual claims, checking each against evidence sources, and returning per-claim trust scores before acting on the result. Reliability & Eval emerging
  42. Policy-Gated Tool Proxy Insert a transparent proxy between agents and tool servers that evaluates every tool call against a policy engine before forwarding, producing an immutable audit trail of all decisions. Security & Safety emerging
  43. Signal-Driven Agent Activation Watches external sources for structured signals and starts predefined agent workflows when declarative rules with thresholds and cooldowns match Orchestration & Control emerging
  44. Zero-Knowledge Verified Agent Egress Intercepts each outbound HTTP or MCP call, proves its claims against a private source of truth with a zero-knowledge proof, and blocks calls that fail Security & Safety emerging
  45. Zero-Trust Agent Mesh Gives each agent a cryptographic identity and verifies identity, signed delegation tokens, and chain depth on every inter-agent request Security & Safety established
  46. Orchestration Prompt-Writing Benchmark Scores an orchestrator on whether its sub-agent prompts give the right information fragments to the right roles, separate from end-task success Reliability & Eval emerging
  47. Commitment Ledger with Reality-Gated Credit Records agent promises and credits retrieved memory only after an externally verifiable outcome settles. Learning & Adaptation proposed
  48. Cross-Protocol Agent Discovery Aggregate agent metadata across fragmented registries and protocols into a unified, protocol-agnostic discovery layer. Tool Use & Environment emerging
  49. Session-Scoped Context Runtime for Agent Tools Interpose a context runtime that caches structured reads and normalizes tool output so sessions reuse compact representations instead of repeating raw tokens. Context & Memory emerging
  50. Denial Tracking & Permission Escalation Track repeated tool denials and auto-escalate to blanket permission prompts or fallback strategies, preventing wasted iterations in non-interactive contexts. Security & Safety emerging
  51. Unified Tool Gateway Route all agent tool calls through a single gateway that handles discovery, authentication, billing, and execution across many heterogeneous providers. Tool Use & Environment emerging
  52. Schema-Guided Graph Retrieval for Multi-Hop Reasoning Use one shared domain schema to align graph construction, schema evolution, query decomposition, and typed retrieval so multi-hop reasoning over private knowledge stays precise as domains change. Context & Memory emerging
  53. Agent Circuit Breaker Prevents agents from wasting tokens and time on repeatedly failing tools by tracking failure rates and temporarily disabling broken tool endpoints Reliability & Eval emerging
  54. Abstracted Code Representation for Review Shows reviewers pseudocode, intent summaries, and logical diffs of code changes, with drill-down to the real code to confirm the mapping UX & Collaboration proposed
  55. Action Caching & Replay Pattern Records each agent action with XPath and frame metadata so later runs replay it without LLM calls, with LLM fallback when replay fails Reliability & Eval emerging
  56. Action-Selector Pattern LLM maps user intent to a pre-approved action ID with schema-validated parameters, and tool outputs never return to the selector Orchestration & Control emerging
  57. Adaptive Sandbox Fan-Out Controller Starts a small batch of parallel sandboxes, then scales up, stops early, or refines the prompt based on early success, variance, and error signals Reliability & Eval emerging
  58. Agent Modes by Model Personality Offers separate working modes, each with its own prompts, tools, UI, and expectations, tuned to one model's working style Orchestration & Control emerging
  59. Agent SDK for Programmatic Control Exposes agent functions through an SDK and CLI so code can run the agent headless with set tools, permissions, and resource limits Tool Use & Environment emerging
  60. Agent-Assisted Scaffolding Agent generates initial files, boilerplate, and directory structure from a high-level description so developers start on core logic UX & Collaboration validated in production
  61. Agent-Driven Research Agent plans its own search queries, runs them across sources, reflects on gaps, and iterates until it can write a sourced report Orchestration & Control established
  62. Agent-First Tooling and Logging Designs tools and logs for agent consumption with one unified log stream, structured JSON output, and agent-aware CLI flags Tool Use & Environment established
  63. Agent-Friendly Workflow Design Gives agents clear high-level goals, room for implementation choices, structured I/O, and plan review before execution UX & Collaboration best practice
  64. Agent-Powered Codebase Q&A / Onboarding Agent indexes the codebase with embeddings and code graphs and answers natural-language questions about where code is and how it behaves Context & Memory validated in production
  65. AI Web Search Agent Loop Coordinator agent translates queries, spawns parallel search workers across domains and time ranges, refines iteratively, and answers with citations Tool Use & Environment emerging
  66. AI-Accelerated Learning and Skill Development Uses AI assistants as tutors that explain errors and concepts, offer alternatives, and fade support as the developer gains skill UX & Collaboration validated in production
  67. AI-Assisted Code Review / Verification Uses AI tools to flag issues, summarize change intent, and explain code so human reviewers focus on intent and business logic Feedback Loops emerging
  68. Asynchronous Coding Agent Pipeline Splits inference, tool execution, reward modeling, and learning into asynchronous workers linked by message queues so GPUs stay busy Reliability & Eval proposed
  69. Burn the Boats Removes working but obsolete features on a hard, announced deadline so the team and users move to the new approach Orchestration & Control emerging
  70. Canary Rollout and Automatic Rollback for Agent Policy Changes Ships agent policy changes to a small traffic slice first, monitors guardrail metrics, and rolls back to the last stable version automatically Reliability & Eval established
  71. CLI-Native Agent Orchestration Exposes agent capabilities as CLI commands with JSON output and exit codes so Makefiles, Git hooks, cron jobs, and CI can script and replay agent runs Tool Use & Environment proposed
  72. Code-Then-Execute Pattern LLM writes a sandboxed program or DSL script, a static taint checker verifies data flows, and an interpreter runs it in a locked sandbox Tool Use & Environment emerging
  73. Codebase Optimization for Agents Optimizes tooling, CLIs, tests, and docs for agents first, with one-command verify loops and machine-readable output, even if human DX regresses UX & Collaboration emerging
  74. Coding Agent CI Feedback Loop Agent pushes a branch, polls CI for partial failures, patches the failing files within a retry budget, and notifies when all tests pass Feedback Loops best practice
  75. Conditional Parallel Tool Execution Runs a batch of tool calls in parallel when all are read-only and in sequence when any modifies state, then returns results in request order Orchestration & Control validated in production
  76. Context Window Anxiety Management Enables a large context window but caps usage lower, and adds prompts that state the token budget so the model does not rush to finish Context & Memory emerging
  77. Context Window Auto-Compaction Catches context overflow errors, compacts and validates the transcript with a reserve token floor, and retries the request automatically Context & Memory validated in production
  78. Context-Minimization Pattern Removes untrusted user text and tool output from context after it becomes a safe structured artifact, so later steps see only trusted data Context & Memory emerging
  79. Continuous Autonomous Task Loop Pattern Runs a loop where subagents pick the next task from a todo file, execute it in fresh context, commit it, and back off on rate limits Orchestration & Control established
  80. CriticGPT-Style Code Review Runs a specialized critic model over generated code to find bugs, security flaws, and quality issues, and feeds critical findings back for regeneration Reliability & Eval validated in production
  81. Cross-Cycle Consensus Relay Each loop cycle reads a structured relay document, does its work, and atomically writes back decisions, open questions, and the single next action Orchestration & Control emerging
  82. Curated Code Context Window A search subagent finds the few relevant files for a task and injects only top snippets or summaries into the main coding agent's context Context & Memory validated in production
  83. Curated File Context Window Loads only the primary task files into the main context and has a search sub-agent rank and summarize secondary files before adding them Context & Memory best practice
  84. Custom Sandboxed Background Agent Builds an in-house background coding agent that runs in sandboxed dev environments, streams progress over WebSocket, and iterates on compiler and test feedback Orchestration & Control emerging
  85. Democratization of Tooling via Agents Non-programmers describe a tool in natural language and iterate with an AI agent that generates and fixes the code for dashboards, scripts, or small apps UX & Collaboration emerging
  86. Deterministic Security Scanning Build Loop Adds SAST, SCA, and secret scanners to the build target that the agent must run after each change, so scanner failures force the agent to fix the code Security & Safety proposed
  87. Dev Tooling Assumptions Reset Audits dev tools for human-effort assumptions and replaces tickets, PR ceremony, and sprints with immediate agent dispatch, variations, and automated tests UX & Collaboration emerging
  88. Disposable Scaffolding Over Durable Features Treats code around the model as temporary scaffolding: build the simplest thing that works now, mark it disposable, and remove it when models improve Orchestration & Control best practice
  89. Dual LLM Pattern Splits work between a privileged LLM that calls tools but never sees untrusted data and a quarantined LLM that reads untrusted data but has no tools Orchestration & Control emerging
  90. Dynamic Code Injection (On-Demand File Fetch) Expands @file or /load tokens in the prompt into file contents, line ranges, or summaries so the agent sees code without manual copy and paste Tool Use & Environment established
  91. Dynamic Context Injection Lets users inject files, folders, or saved prompts into the agent's context mid-session with @-mentions and custom slash commands Context & Memory established
  92. Economic Value Signaling in Multi-Agent Networks Attaches a token value to inter-agent requests so recipients prioritize by value, with a peer registry for discovery and ledger settlement Orchestration & Control experimental but awesome
  93. Episodic Memory Retrieval & Injection Writes a structured memory record after each episode and injects the top-k similar past memories as hints when a new task starts Context & Memory validated in production
  94. Explicit Posterior-Sampling Planner Embeds Posterior Sampling for RL in the LLM's reasoning: sample a task model, plan, act, observe reward, and update the posterior Orchestration & Control emerging
  95. Extended Coherence Work Sessions Combines models with long coherence windows with context compaction, prompt caching, and persisted state so agents stay on task for hours Reliability & Eval rapidly improving
  96. External Credential Sync Reads AI credentials from other tools' keychains and config files and syncs them into the agent's store, with near-expiry refresh, OAuth upgrade, and dedupe Security & Safety validated in production
  97. Factory over Assistant Spawns several autonomous agents in parallel with automated feedback loops such as tests and builds, and checks on them later instead of watching one agent Orchestration & Control validated in production
  98. Failover-Aware Model Fallback Classifies each model failure by reason and falls back along a model chain only for retryable errors, failing fast on auth, billing, and user aborts Reliability & Eval validated in production
  99. Frontier-Focused Development Targets only state-of-the-art models, picks the best model per use case with no user model selector, and expects to rework the product every few months Learning & Adaptation emerging
  100. Graph of Thoughts (GoT) Represents reasoning as a directed graph of thoughts so the model can branch, aggregate, refine, and revisit reasoning paths Feedback Loops emerging
  101. Hook-Based Safety Guard Rails for Autonomous Code Agents Runs small shell scripts on PreToolUse and PostToolUse hooks to block destructive commands, lint edits, and warn on context use outside the agent's reasoning Security & Safety validated in production
  102. Incident-to-Eval Synthesis Converts each production incident into executable eval cases with pass/fail criteria and gates future releases on them Feedback Loops emerging
  103. Inference-Healed Code Review Reward Replaces a binary tests-passed reward with a critic that scores correctness, style, performance, and security, explains low scores, and combines them Feedback Loops proposed
  104. Inference-Time Scaling Spends extra compute at inference time with best-of-N sampling, longer reasoning, self-refinement, search, and verification to improve output quality Orchestration & Control emerging
  105. Initializer-Maintainer Dual Agent Architecture Uses a one-time initializer agent to create the feature list, progress files, and bootstrap script, then a coding agent that resumes from them one feature per session Orchestration & Control validated in production
  106. Intelligent Bash Tool Execution Runs agent shell commands through a multi-mode executor that uses a PTY when needed, falls back to direct exec, and manages approvals, background jobs, and signals Tool Use & Environment validated in production
  107. Iterative Multi-Agent Brainstorming Spawns several agents in parallel on the same problem, often with different perspectives, then merges their ideas into one set of options Orchestration & Control experimental but awesome
  108. Iterative Prompt & Skill Refinement Combines a feedback channel, editable prompt documents, log-driven skill fixes, and usage dashboards to improve agent prompts and skills continuously Feedback Loops established
  109. Lane-Based Execution Queueing Routes work into named lanes, each with its own queue and concurrency limit, so sessions never interleave and background jobs never block users Orchestration & Control validated in production
  110. Language Agent Tree Search (LATS) Runs Monte Carlo Tree Search over reasoning steps, with the LLM generating candidate actions and scoring partial solutions to pick the best path Orchestration & Control emerging
  111. Layered Configuration Context Loads context files from enterprise, user, project, and local levels automatically and merges them into the agent's baseline context Context & Memory established
  112. Lethal Trifecta Threat Model Classifies each tool by private-data access, untrusted-content exposure, and external communication, and blocks any execution path that has all three Reliability & Eval best practice
  113. LLM Map-Reduce Pattern Processes each untrusted document in its own sandboxed LLM call with a constrained output, then aggregates the results with deterministic code Orchestration & Control emerging
  114. LLM-Friendly API Design Designs APIs with visible versions, self-descriptive names and schemas, simple calls, actionable errors, and few indirection levels so LLMs call them correctly Tool Use & Environment emerging
  115. Merged Code + Language Skill Model Fine-tunes separate language and code specialists from the same base model, then merges their weights into one model instead of one large joint training run Reliability & Eval emerging
  116. Milestone Escrow for Agent Resource Funding Holds agent funding in escrow and releases each payment only after independent verification of a measurable milestone UX & Collaboration emerging
  117. Multi-Model Orchestration for Complex Edits Splits complex code edits across specialized models: a retrieval model gathers context, a large model writes the changes, and smaller models apply them Orchestration & Control validated in production
  118. No-Token-Limit Magic Removes hard token limits during prototyping to learn what good behavior needs, then compresses context only after quality is stable and measured Reliability & Eval experimental but awesome
  119. Non-Custodial Spending Controls Puts a policy layer between the agent and the transaction signer that checks each intent against allowlists, budgets, and rate limits, and fails closed Security & Safety emerging
  120. Oracle and Worker Multi-Model Approach Uses a fast, low-cost worker model for most tool use and code generation, and lets it consult an expensive oracle model when it is stuck Orchestration & Control emerging
  121. Patch Steering via Prompted Tool Selection Tells the agent in the prompt which patch or refactoring tool to use, with usage examples, negative rules, and a fallback order Tool Use & Environment best practice
  122. Plan-Then-Execute Pattern Has the LLM fix the full sequence of tool calls before it reads untrusted data, then runs that sequence so tool outputs change only parameters Orchestration & Control established
  123. Planner-Worker Separation for Long-Running Agents Splits agents into planners that create tasks, workers that complete them in isolation, and a judge that decides each cycle whether to continue Orchestration & Control emerging
  124. Prompt Caching via Exact Prefix Preservation Keeps static prompt content first in a fixed order and only appends new messages, including config changes, so each request reuses the cached prefix Context & Memory emerging
  125. Recursive Best-of-N Delegation Runs several parallel candidate workers per subtask in a recursive agent tree, scores them with tests and a judge, and promotes the best result upward Orchestration & Control emerging
  126. Reliability Problem Map Checklist for RAG and Agents Runs a fixed 16-question failure checklist on each RAG or agent incident, maps the result to repair actions, and re-tests the same failing case Reliability & Eval proposed
  127. Rich Feedback Loops > Perfect Prompts Returns compiler errors, test failures, lint output, and human feedback to the agent after each tool call so it can plan fixes and self-correct Feedback Loops validated in production
  128. RLAIF (Reinforcement Learning from AI Feedback) Uses an AI model guided by written principles to critique outputs and label preferences, which train a reward model that optimizes the policy Reliability & Eval emerging
  129. Sandboxed Tool Authorization Filters an agent's tools through deny-first allow and deny patterns, profile presets, and subagent policies that inherit parent restrictions Security & Safety validated in production
  130. Schema Validation Retry with Cross-Step Learning Retries failed structured outputs with the validation errors as feedback and adds recent errors from earlier steps to later prompts so mistakes do not repeat Reliability & Eval emerging
  131. Seamless Background-to-Foreground Handoff Lets a user take over a background agent's unfinished work in the foreground, with the agent's branch, PR, and summaries carried over as context UX & Collaboration emerging
  132. Self-Critique Evaluator Loop Trains a judge model on its own synthetic comparisons of candidate outputs and uses it as a reward model or quality gate for the main agent Feedback Loops established
  133. Self-Discover: LLM Self-Composed Reasoning Structures Has the LLM select, adapt, and compose reasoning modules into a task-specific reasoning structure, then solve the task by following that structure Feedback Loops emerging
  134. Self-Identity Accumulation Injects a persistent identity document at session start and updates it with new user insights at session end through lifecycle hooks Context & Memory emerging
  135. Self-Rewriting Meta-Prompt Loop Has the agent reflect after each episode, draft edits to its own system prompt, validate them through guardrails, and save the new version Orchestration & Control emerging
  136. Semantic Context Filtering Pattern Extracts only the semantic or interactive parts of raw data, such as accessibility trees or relevant fields, before sending it to the LLM Context & Memory emerging
  137. Shell Command Contextualization Lets the user run a shell command with a prefix such as ! and injects the command and its full output into the agent's context Tool Use & Environment established
  138. Shipping as Research Releases reversible, instrumented features to learn whether they work, then doubles down or removes them based on usage data and feedback Learning & Adaptation emerging
  139. Soulbound Identity Verification Binds agent identity to a non-transferable credential with a committed state hash and logs signed state changes so verifiers can check continuity Security & Safety emerging
  140. Spec-As-Test Feedback Loop Generates executable tests from the spec on every spec or code commit and opens agent PRs that fix code or flag unclear spec parts Feedback Loops emerging
  141. Specification-Driven Agent Development Makes a version-controlled spec file the agent's main input, scaffolds code from it, and links each artifact back to a spec clause Orchestration & Control proposed
  142. Spectrum of Control / Blended Initiative Gives users several autonomy modes, from inline completion to background agents, and lets them switch modes per task UX & Collaboration validated in production
  143. Static Service Manifest for Agents Serves a static llms.txt or agent.json file at a well-known URL that lists services, auth, and limits so agents can plan before they call Tool Use & Environment emerging
  144. Stop Hook Auto-Continue Pattern Runs a stop hook after each agent turn that checks success criteria and makes the agent continue until tests or checks pass Orchestration & Control emerging
  145. Sub-Agent Spawning Lets the main agent spawn sub-agents with fresh context and scoped tools to work on subtasks in parallel, then merges their results Orchestration & Control validated in production
  146. Subagent Compilation Checker Spawns one subagent per module to build and check it, and returns only a short structured error list or artifact reference to the main agent Reliability & Eval emerging
  147. Subject Hygiene for Task Delegation Requires a specific action-plus-target subject on every subagent task so each delegated job stays traceable and easy to reference Orchestration & Control validated in production
  148. Three-Stage Perception Architecture Splits the agent into separate perception, processing, and action stages so each stage can be built, tested, and scaled on its own Orchestration & Control established
  149. Tool Capability Compartmentalization Splits tools into reader, processor, and writer classes with scoped permissions and blocks tool chains that combine private data, untrusted input, and external writes Orchestration & Control emerging
  150. Tool Selection Guide Maps each task type to a preferred tool: Glob, Grep, and Read to explore, Edit to modify, Bash to verify, and Task with a clear subject to delegate Orchestration & Control emerging
  151. Tool Use Incentivization via Reward Shaping Gives dense RL rewards for useful intermediate tool calls such as compile, lint, and test so the agent learns to use tools instead of only thinking Feedback Loops emerging
  152. Tool Use Steering via Prompting Tells the agent in the prompt which tool to use, how to learn a custom tool, and which shorthands map to tool sequences Tool Use & Environment best practice
  153. Transitive Vouch-Chain Trust Builds a graph of signed vouches between agents and derives trust in an unknown agent from the chain back to a trusted one, with decay at each hop Security & Safety emerging
  154. Variance-Based RL Sample Selection Runs the base model several times per sample and trains RL only on samples with score variance, skipping ones that are always right or always wrong Learning & Adaptation validated in production
  155. Verbose Reasoning Transparency Lets users open a verbose view on demand that shows the agent's interpretation, tool choices, intermediate steps, and raw tool outputs UX & Collaboration best practice
  156. Versioned Constitution Governance Stores the agent constitution in a signed Git repository where the agent can only propose changes and reviewers or CI gates merge them Reliability & Eval emerging
  157. Virtual Machine Operator Agent Gives the agent a dedicated virtual machine where it can run code, install packages, use the file system, and operate CLI tools Tool Use & Environment established
  158. Visual AI Multimodal Integration Adds multimodal models to the agent so it can analyze images, video, and screenshots and combine them with text to reason and act Tool Use & Environment emerging
  159. Working Memory via TodoWrite Keeps an explicit todo list with status, blockers, and next steps during the session so agent and user can track progress Context & Memory emerging
  160. Workspace-Native Multi-Agent Orchestration Runs agents inside the team's workspace platform so they share its memory, knowledge sources, event triggers, and integrations with humans Orchestration & Control emerging
  161. Tool Search Lazy Loading Dynamically load tools via search instead of preloading all available tools to reduce context usage Context & Memory emerging
  162. Hybrid LLM/Code Workflow Coordinator Lets each workflow pick an LLM or a code script as its coordinator, so teams prototype with the LLM and move to reviewed code when determinism matters Orchestration & Control proposed
  163. LLM Observability Sends agent runs to an LLM observability platform for span-level traces of each LLM call and tool use, plus aggregate cost, latency, and success metrics Reliability & Eval proposed
  164. Memory Reinforcement Learning (MemRL) Stores memories with learned utility scores, retrieves by similarity then re-ranks by utility, and updates scores from outcomes while the LLM stays frozen Learning & Adaptation proposed
  165. Multi-Platform Webhook Triggers Starts agent workflows from Slack, Notion, and Jira webhooks, emoji reactions, and schedules, with idempotency and signature checks on each event Tool Use & Environment emerging
  166. Progressive Disclosure for Large Files Puts only file metadata in the prompt and gives the agent load, peek, and extract tools to pull file content into context on demand Context & Memory emerging
  167. Skill Library Evolution Saves working agent code as reusable skills in a skills directory, documents and tests them over time, and loads skill details only on demand Learning & Adaptation established
  168. Workflow Evals with Mocked Tools Runs complete agent workflows against mocked tools in CI and checks which tools were called plus agent-as-judge quality criteria Reliability & Eval emerging
  169. Agent Reinforcement Fine-Tuning (Agent RFT) Trains model weights with reinforcement learning on real tool calls and custom graders to improve domain-specific tool use and multi-step reasoning Learning & Adaptation emerging
  170. Agentic Search Over Vector Embeddings Replaces vector indexes with agent-driven grep, find, and file traversal that searches current file state on demand and refines iteratively Tool Use & Environment best practice
  171. Anti-Reward-Hacking Grader Design Design reward functions with multi-criteria evaluation and iterative hardening to prevent models from gaming graders, ensuring training rewards align with actual task quality. Reliability & Eval emerging
  172. Autonomous Workflow Agent Architecture Runs multi-step engineering workflows in containers with tmux sessions, adaptive monitoring, checkpoints, and context-aware error recovery Orchestration & Control established
  173. Background Agent with CI Feedback Runs the agent in the background on its own branch and uses CI results as feedback to patch failures until green or blocked Feedback Loops validated in production
  174. Chain-of-Thought Monitoring & Interruption Streams agent reasoning and tool calls in real time so a human can interrupt and redirect early when the approach is wrong UX & Collaboration emerging
  175. CLI-First Skill Design Builds each skill as a standalone CLI with JSON output and exit codes so humans and agents use the same interface Tool Use & Environment emerging
  176. Code Mode MCP Tool Interface Improvement Pattern LLMs generate TypeScript code to orchestrate MCP tools in ephemeral V8 isolates, eliminating token-heavy round-trips and enabling efficient multi-step workflows with 10x+ token savings. Tool Use & Environment established
  177. Code-Over-API Pattern Agent writes code that calls tools and filters data inside a sandbox, so only summaries and samples return to the context window Tool Use & Environment established
  178. Compounding Engineering Pattern After each feature, codifies agent mistakes and learnings into CLAUDE.md, slash commands, subagents, and hooks so the next feature is easier to build Learning & Adaptation emerging
  179. Discrete Phase Separation Splits work into separate research, planning, and implementation conversations, each with fresh context, passing only distilled outputs between phases Orchestration & Control emerging
  180. Distributed Execution with Cloud Workers Runs many agent sessions in parallel on cloud workers, each in its own git worktree, with dependency-aware scheduling, merge coordination, and approval gates Orchestration & Control emerging
  181. Dogfooding with Rapid Iteration for Agent Improvement The agent team uses its own agent for daily work, collects feedback in low-friction channels, and ships features internally first to validate or discard them Feedback Loops best practice
  182. Dual-Use Tool Design Builds each tool with one interface and implementation that both humans and agents can call, with the same output, permissions, and logs Tool Use & Environment best practice
  183. Feature List as Immutable Contract Defines every feature up front in a JSON list with acceptance steps; the agent may only flip a feature to passing after it verifies it Orchestration & Control emerging
  184. Filesystem-Based Agent State Agents persist intermediate results and working state to files, creating durable checkpoints that enable workflow resumption, recovery from failures, and support for long-running tasks. Context & Memory established
  185. Human-in-the-Loop Approval Framework Systematically insert human approval gates for designated high-risk functions while maintaining agent autonomy for safe operations, with multi-channel approval interfaces and comprehensive audit trails. UX & Collaboration validated in production
  186. Inversion of Control Gives the agent tools, a high-level objective, and guardrails, then lets it choose sequencing and recovery while humans set policy and review Orchestration & Control validated in production
  187. Isolated VM per RL Rollout Spin up an isolated virtual machine for each RL rollout to prevent cross-contamination between parallel agent executions, ensuring safe training with destructive tool access. Security & Safety emerging
  188. Latent Demand Product Discovery Builds hackable, extensible products, watches how power users repurpose them, and turns the most frequent workarounds into supported features UX & Collaboration best practice
  189. Memory Synthesis from Execution Logs Has the agent write a structured diary per task, then runs synthesis agents over many diaries to turn recurring patterns into rules, commands, and tests Context & Memory emerging
  190. Multi-Platform Communication Aggregation Queries every communication platform in parallel through adapters that share one schema, then merges, deduplicates, and ranks the results Tool Use & Environment emerging
  191. Opponent Processor / Multi-Agent Debate Pattern Spawns agents with opposing roles on the same context, lets them critique each other, then synthesizes their positions to expose bias and blind spots Orchestration & Control emerging
  192. Parallel Tool Call Learning Uses agent reinforcement fine-tuning to teach the model to issue independent tool calls in parallel, which cuts sequential rounds and latency Orchestration & Control emerging
  193. PII Tokenization Replaces PII in tool results with placeholder tokens before the model sees them and restores the real values in outgoing tool calls Security & Safety established
  194. Proactive Agent State Externalization Gives agents note templates, completeness checks, and an external memory fallback so self-written notes keep objectives, decisions, and knowledge gaps Context & Memory emerging
  195. Proactive Trigger Vocabulary Gives each skill an explicit, documented list of trigger phrases and patterns so input routes to skills predictably, with optional proactive activation UX & Collaboration emerging
  196. Progressive Autonomy with Model Evolution Audits prompts and orchestration after each model upgrade and removes the scaffolding that evals show the new model no longer needs Orchestration & Control best practice
  197. Progressive Complexity Escalation Deploys agents on low-complexity, high-reliability tasks first and unlocks higher capability tiers when metrics and human review gates show proven reliability Orchestration & Control emerging
  198. Progressive Tool Discovery Organizes tools in a browsable hierarchy and lets the agent load names, descriptions, or full schemas only for the tools it needs Tool Use & Environment established
  199. Reflection Loop Scores each draft against a fixed rubric, feeds the critique into a revision, and repeats until the draft passes a threshold or the retry budget ends Feedback Loops established
  200. Structured Output Specification Constrain agent outputs using deterministic schemas that enforce structured, machine-readable results, enabling reliable validation, parsing, and integration with downstream systems. Reliability & Eval established
  201. Swarm Migration Pattern Main agent orchestrates 10+ parallel subagents working simultaneously on independent migration chunks, achieving 10x+ speedup for large-scale framework upgrades, lint rule rollouts, and API migrations. Orchestration & Control validated in production
  202. Team-Shared Agent Configuration as Code Check agent configuration into version control as code, enabling consistent behavior across teams, faster onboarding, and collaborative improvement through PRs and code review. UX & Collaboration best practice
  203. Tree-of-Thought Reasoning Expands a tree of candidate reasoning steps, scores partial states, prunes weak branches, and picks the best path instead of one linear chain Orchestration & Control established