Resources
65 agent frameworks you can copy, each with its source. Several of the best took a rule out of the prompt.
Spotify wrote its routing rules into CLAUDE.md and Claude could ignore them; a hook that blocks whole-file reads over 350 lines could not be ignored. 65 steps, templates and file formats, each with its source, sorted by what you are trying to build.

Most writing about AI agents is opinion about agents. This library holds only the other kind: a thing you can copy. Named steps in order, a template with placeholders, a file format, a parameter set. An essay does not get in, however good it is.
Sixty-five entries, gathered since August from companies and people who built the thing and wrote down how. Each entry names the artifact, the one detail worth stealing, who made it, and the source, so you can check it yourself. Every link was checked on 23 September 2026.
Agent workflow templates you can copy: what the strongest ones share
Several of the frameworks with measured results behind them moved a rule out of the prompt and into something the agent cannot talk its way past.
Spotify's first version of its model router put the routing rules in CLAUDE.md. The author's verdict: the rules "were advisory, not enforced. Claude could ignore them." The version that shipped uses a hook that blocks any whole-file read over 350 lines (the default) and sends it to a cheaper model instead, and the write-up reports mean savings of about 90% on those bulk reads. Checkly let an agent rewrite a service that handles about 92 million messages a day, and made one thing the main acceptance criterion: a black-box harness replaying production traffic. The result was about 13,000 lines of Go, zero incidents and 70% fewer running pods. Vercel's design.md works the same way: each round of reviewer feedback becomes a check that fails automatically next time, and in a six-page test, pages built with it had 57% fewer known layout failures (39 against 91), a sample Vercel itself calls too small to prove quality.
The pattern repeats down the list. Ordewell counts a task done only when it emits evidence markers, not when the agent says so. Cloudflare's security-audit skill never lets the agent that found a finding be the one that checks it. If a rule matters, give it teeth: a hook, a test, a file the harness loads, or a second agent that checks the first.
Two samples, so you can see what an entry looks like
A personal benchmark, with no code. Every's method, as Laura Entis writes it up: pick one task you hand to AI often, and turn each correction you make to an AI's draft into a yes/no check ("One idea per slide."). Grade the output yourself, then have an AI grade the same checklist blind, and rerun it across models whenever a new one ships. It answers a question no public benchmark does: which model is best at my work?
A hiring-signal prompt for go-to-market, not code. Bogdan Liutic's prompt reads a company's job postings and returns a buying-window verdict as JSON. It detects four signals (REPEAT_POSTING, FIRST_HIRE, STALE_ROLE, MIGRATION), throws out ghost jobs, maps each signal to a buyer, and must quote the posting verbatim. Below 60 confidence it answers NONE. The trust rule is the part to keep: backtest it on your last 15 closed-won accounts before you use it on a prospect.
How to use this library
Eight families, sorted by what you are trying to build. The first four: evals and verification; orchestration patterns; harnesses you can read or fork; and memory and context files. The other four: cost and model routing; software factories and review gates; permissions, sandboxes and security; and copyable prompts and skills for go-to-market, design and writing.
Each entry says what you copy, the one detail worth stealing, who made it, and the source. Licences are given where the source states one; where it does not, assume none. Figures are the maker's own unless the entry says otherwise, and where a maker says a number is self-reported, the entry says so too. Start with the family closest to the problem you have this week, and copy one thing, not eight.
Evals and verification
- Personal benchmark: a yes/no checklist built from your own corrections. Pick one recurring task and turn each correction into a yes/no check ("One idea per slide."). Grade it yourself, then have an AI grade the same checklist blind. Rerun across models whenever a new one ships. By: Laura Entis, Every. Source: every.to.
- Error discovery: three steps before any eval metric, plus two prompts. Log each user session as one trace in
traces/traces.jsonl. Annotate 10 traces yourself before the agent suggests anything, aim for 100, and stop at the first upstream error. Then cluster failure modes and count them. Plugin:npx skills add https://github.com/ai-evals-course/evals-skills. By: Hamel Husain and Shreya Shankar. Source: lennysnewsletter.com. - Evaluation design: seven steps from a decision to a hidden holdout. Begin with the decision the eval must inform, not a metric. Include rare, high-risk and adversarial cases, calibrate graders with inter-rater agreement, weight failures by severity, and keep a hidden holdout nobody optimises against. By: Surge AI. Source: surgehq.ai.
- Finish-the-sentence test: two lines per grader. "If this grader passes, I now know that ______." "If this grader fails, I now know that ______." A
string contains "Azure"grader also passes on the line// Don't use Azure here.By: Waldek Mastykarz. Source: blog.mastykarz.nl. - Black-box harness: the main acceptance test for an agent rewrite. Replay production data against the service and make "the harness passes against the new implementation" the main acceptance criterion. Result: about 13,000 lines of Go, zero incidents, 70% fewer running pods, within a $200 subscription's daily limits. By: Checkly. Source: checklyhq.com.
Orchestration patterns
- The Gauntlet Loop: seven steps and a meta-prompt. Define the bar as a concrete reference (his worked example uses actual Call of Duty screenshots), separate the builder and critic agents, and loop with no fixed round limit. Needs Claude Code, Codex or similar. It will not work in a standard chat interface. By: Matt Shumer. Source: somethingbig.ai.
- Advanced agentic harness: nine components, each with the failure it fixes. TypedTool (Pydantic validation) fixes hallucinated tool arguments. BudgetMulti (tokens, tool calls, wall time, cost) fixes uncontrolled cost. Includes the planner prompt: "Output ONLY JSON… Nodes may run in parallel if deps are empty." By: Bruno Gonçalves, Data For Science. Source: data4sci.com.
- Session lifetime: a 24-hour coordinator and 30-second specialists, prompts in full. The coordinator loads
preferences.mdeach morning, never executes directly, and saves durable learnings at midnight before it terminates. Built against compaction silently dropping standing rules in 30–59% of episodes (arXiv:2606.22528). By: Tomasz Tunguz. Source: tomtunguz.com. - Agents as new hires: one bot per job, named for the job. Hire bots for specific jobs, and let the name tell each one what its job is. Holly Helpdesk can handle refunds but cannot issue one. The human gets a button that opens the real Stripe request. Three more workflow templates are linked from the page. By: Claire Vo. Source: chatprd.ai.
- Production harness: a guardrail for each phase of the agent loop. Clean the input before the agent is invoked. Set schemas and execution limits at initialisation. Check for hallucinations and enforce tool boundaries inside the run. Validate output before it becomes real data. Then add observability, a learning loop and evals. By: monday.com. Source: engineering.monday.com.
- Paid Media Agent: a workspace-first build and six numbered lessons. The system prompt is a map to where things are, not a container for everything. Lesson 5: propose campaign changes, route them through human approval, and verify the change was applied. By: LangChain. Source: langchain.com.
- Harness state: two configurations of one model, with score and cost. On ARC-AGI-3, GPT-6 Astra's best score was 62.7% ($26,000) when it kept only notes it chose itself. Its best rose to 99.9% ($19,000) when the harness preserved opaque reasoning state plus compaction. The two bests come from different reasoning-effort settings; at the same setting (max) it was 62.7% against 98.6%. By: Greg Kamradt, ARC Prize Foundation. Source: arcprize.org.
- Agentic engineering: nine pieces of a coding-agent setup, as a checklist. Harnesses, project rules, research, context retrieval, LSP diagnostics, evals, telemetry, secret handling, multi-agent work. Each links to his own implementation in
my-pi. For each one you either have an answer or have found a gap. By: Scott Spence. Source: scottspence.com.
Harnesses you can read or fork
- hip: a harness whose core loop is about 200 lines of Python. Configuration is environment variables. Actions are shell commands. A subagent is a child process. By: Jonathan Chang. Source: jonathanc.net.
- Tau: three packages with one stated boundary.
tau_coding → tau_agent → tau_ai.AgentHarnessis the reusable brain,CodingSessionis the coding environment, and the TUI is one possible frontend. A Python port of Pi's minimalist coding agent, written to be read. By: Hugging Face. MIT. Source: github.com/huggingface. - omp²: four rules from a harness rewrite. One authority over session state. A controller / actor split, which subagents cross too. "The sandbox executes, it does not decide." Execution as a state stream, replacing the callback split. The write-up has 327 code blocks. By: Can Bölük. Source: stencil.so.
- Munder Difflin: a multi-agent harness that wraps the agent CLI you already run. Agents write to their own
outbox/and a router delivers to each recipient'sinbox/, inside a local git repo that no agent commits to. A circuit breaker steers, constrains, then stops an agent that loops or blows its budget. By: chaitanyagiri. MIT. Source: github.com/chaitanyagiri. - TeamAI: one git repo as the harness for the whole team.
npm install -g teamai-cli. An admin creates a shared-experience repo and grants write access. Each member runsteamai init. Every session then pulls the current skills, rules and other harness updates. By: Tencent. MIT. Source: github.com/Tencent. - Ouroboros: interview, frozen spec, ledger, staged eval. The agent interviews you until the request can be frozen into a seed spec the run cannot edit. Every action goes to a replayable ledger. Adapters for 14 runtimes. By: Q00. MIT. Source: github.com/Q00.
- Ordewell: one goal becomes an editable dependency graph. Each task gets its own runner and model. Tasks run in parallel where the graph allows. A task counts as done only when it emits explicit evidence markers, not when the agent says so. By: Ordewell. Apache-2.0. Source: github.com/ordewell.
- Pizza Bot: an inbox for long-running agents. Two queues. Unread holds completed work. Action holds durable approval requests that outlive the session that raised them. Runs are checkpointed. Local file access is granted per folder, read-only or writable. By: Pizza Bot, developed at Amazon. Apache-2.0. Source: github.com/pizza-bot-app.
- AI Systems Atlas: 78 named agent architectures, plus a notation for drawing your own. Five collections, from Agent Harnesses (18 plates) to Coding Agents (16). The part to copy is the colour grammar: neutral by default, blue for executing, red where trust ends, green when verified. By: Zolt Labs / Engineer's Codex. No licence stated for the diagrams. Source: aisystemsatlas.com.
- RL recipe for knowledge-work agents, in order. Post-train the small model (Qwen3.6-35B-A3B) first. Keep DPPO +
prompt_mean+ a harness nudge, without length penalty or curriculum learning. Then launch the 397B run. APEX-Agents Pass@1 went from 16.11% to 27.29%. By: Mercor and SkyRL. Source: mercor.com.
Memory and context files
BEHAVIOR.md: a file format for agent behaviour specs. A directory at.agents/behaviors/<name>/with YAML frontmatter (name≤64 chars,description≤1024) and a free-form body with five recommended dimensions: Intent, Evidence, Decision, Execution, Recovery. A worked template is included. By: Braintrust and Basis. Apache-2.0. Source: github.com/braintrustdata.- Domain-driven agents: a
.workflow.jsonmanifest plus aCONTEXT.mdglossary per context. Each edge declaresdirection,owner,pattern,shapeandnote.patterncan beunclassifiedwhen no single label is true. The glossary lists deliberately rejected synonyms, which stops the model inventing a fourth spelling. By: Cold Take. Source: coldtake.dev. - OKF Agent Memory: Markdown files with YAML frontmatter in
knowledge/. Frontmatter carriessources, a trust tier (generatedvsverified),statusandstale_after. Two rules sit on top: Progressive Disclosure throughindex.mdfiles, and Search-Before-Write. By: okf-memory, implementing Google's Open Knowledge Format v0.2. MIT. Source: github.com/okf-memory. - Memoryfields: agent memory as an open file format with an RFC-style spec. It ships a
SPEC.md, a skill and a command-line tool. The archival form is a zipfile. It works over local files, Amazon S3, GitHub or plain HTTP. By: Cal Paterson. Source: calpaterson.com. - Funes: a memory you own for coding agents, in six ordered steps. The article's section order is the method. Add memory to the agent you already use. Treat memory as a dataset, not a service. Ask first, wire later. The repo (
huggingface/funes) is Apache-2.0. By: David Corvoysier, Hugging Face. Source: huggingface.co. - Compaction in Pi: context compaction, specified well enough to rebuild. Keep recent turns unchanged. Summarise older ones in a separate request with a different system prompt ("a context summarization assistant"), in sections for goal, progress and key decisions. Token budget: 20,000 by default. By: Earendil Engineering. Source: earendil.com.
- Frame selection: a parameter set for choosing the frames an LLM watches. About 100–150 frames per video. Deduplicate on three channels at 16×16, 32×32 and 192×192, with thresholds of 25/255 and 45/255. Silences longer than 1.5 seconds become frame-only spans. Tool:
claude-real-video, MIT. By: Leo Huang. Source: leoaido.com.
Cost and model routing
- Uber's cost levers, with Uber's own numbers. Cap interactive sessions at 400K tokens. Lengthen the prompt cache from five minutes to one hour. Weekly agent requests rose 9.4× since February while total AI spend stayed roughly flat since April. By: Uber (from Uday Kiran Medisetty's engineering post), reported by Axios. Source: axios.com.
- Spotify's Portal and shunt: route by file size, enforce it with a hook. Large file reads and predictable code generation go to Gemini 2.5 Flash workers; debugging, architecture and editing stay on the frontier model. Two YAML modes (
bulk-reader,code-writer, temperature 0.2) plus a PreToolUse hook,check-file-size, that blocks anyReadover 350 lines by default and redirects it. The first version put these rules inCLAUDE.md, and Claude could ignore them. Tested on four scenarios in a Java monorepo, mean bulk-read savings were about 90%. By: Dimitri Mazmanov, Spotify. Source: engineering.atspotify.com. - Three-model routing: an expensive model plans, cheap ones execute, an expensive one reviews. A local model-gateway plugin routes GPT requests out of Claude Code. Astra plans, cheaper GPT-5.6 models do the volume work, and Opus reviews. The setup notes say where the proxy breaks. By: eigenwise.io. Source: eigenwise.io.
- OpenRouter checklist: ten production gotchas. From 18M+ messages of his own traffic (self-reported): the same model benchmarks very differently across providers, a vision model can have blind providers,
200 OKcan come back with no answer, and you should test from prod, not your laptop. By: Mo Moustafa. Source: mmoustafa.com. - Fable 5.1 prompting guide: a 14-item operating checklist. Re-test effort from scratch: the default is
high, andmediumcan roughly match the previous model at lower cost. Keep history append-only, because cache reads are $0.25 per million tokens. By: Anthropic. Source: platform.claude.com.
Software factories and review gates
- OpenAI's software factory: nine steps from outcome to incident bot. Two parts to copy. PRs are classified by risk, and parts of the codebase can opt in to an agent that auto-approves low-risk ones. Each change also gets a deploy agent that builds its own monitoring dashboard. By: OpenAI (Venkat Venkataramani and team), interviewed by Gergely Orosz. Source: newsletter.pragmaticengineer.com.
- Warp's software factory: Slack request to tested PR in five steps. Tag the bot in Slack. The agent opens a Linear issue, drafts a PR, records a computer-use QA video and submits it for review. Track human interactions per PR. Kickoff to PR averages 35 minutes, and the first human review arrives 3.5 hours later. By: Zach Lloyd (Warp), with Claire Vo. Source: chatprd.ai.
- PostHog's self-driving loops: six signal sources, one loop shape. Sources: MCP feedback, Slack reports, issue specs, anomaly alerts, session replays, runtime logs. Scouts deduplicate evidence, an agent investigates and opens a PR, a human gates the merge, and the loop returns after deployment to confirm the fix held. By: PostHog. Source: posthog.com.
- Nous's parallel refactor: a worktree per agent against a frozen baseline. 1,393 subagents cut non-test Python source by 34.4%, for about $19,300 in the main run (roughly $25k with follow-ups). Community review still caught removed public APIs and changed exception handling that the tests missed. By: Teknium, Nous Research. Source: nousresearch.com.
- Blast-radius review: a two-branch rule. Low-risk changes get AI review only. High-risk changes need a human. Uber's
uReviewfilters in four stages: bots comment, low-confidence comments are removed, the rest are merged and trimmed, and only the important ones reach a developer. By: Anthropic and OpenAI (the two-branch rule) and Uber (uReview), collected by Gergely Orosz. Source: newsletter.pragmaticengineer.com. - The Land PR loop: six steps with a video as the merge gate. Run
Devin Reviewup to twice. The agent records itself testing the feature in a browser. You watch the video and sendVideo approved. Land it.Used to merge up to 40 agent-written pull requests a day. By: Ryan Carson (Untangle). Source: chatprd.ai. - Airbnb's migration pipeline: validate each step, retry, tune on samples, finish by hand. Each file moves through per-file validation and refactor steps, with configurable retry loops and expanded prompt context. The first bulk run migrated 75% of nearly 3,500 test files in about four hours. Four days of "sample, tune, and sweep" refinement took it to 97%. Engineers finished the last 3% by hand, starting from the failed automated refactors. Total: six weeks, against a manual estimate of 1.5 years of engineering time. By: Charles Covey-Brandt, Airbnb. Source: airbnb.tech.
Permissions, sandboxes and security
- Graduated autonomy: four permission tiers that move both ways. T1 Probation to T4 Autonomous, driven by a 0–100 score. Promotion is slow and demotion is immediate. A 30% operator rejection rate caps safety at 70. A delegated action's tier is the minimum across the delegation chain. By: Dev Arora, Meera Kezhukoot and Sathish Kumar Prabakaran, AWS. Source: aws.amazon.com.
- AX: YAML manifests for sandboxed, network-fenced agent tasks. Task, Workspace and Gateway kinds under
apiVersion: ax.io/v1alpha1. Thenax apply -f task.yaml,ax watch,ax ssh,ax suspend/ax resume. Suspended agents keep their state, and resume is claimed at under a second. By: Google. Apache-2.0. Source: github.com/google. - Windows MXC: a JSON containment policy turned into a sandbox. A policy object sets network, UI, filesystem and execution limits. Then
createConfigFromPolicy(policy, isolationLevel, name)andspawnSandboxFromConfig(). Five containment levels, from Process to Full VM. Still in development. By: Microsoft, reported by Gergely Orosz and Ivan Klaric. Source: newsletter.pragmaticengineer.com. - Cursor's self-hosted workers: planning stays in the cloud, tools run on your machine. Run
agent worker start. The worker keeps a long-lived outbound HTTPS connection, so the vendor never connects into your network. A controller starts machines from a spawn script your team supplies, and idle machines hibernate. By: Cursor. Source: cursor.com. - Cloudflare's security-audit skill: six phases. Reconnaissance, coverage-led hunting, candidate validation, structured output, independent record verification by fresh agents, target-neutral reporting. Every finding gets one of three verdicts:
confirmed,needs_validationorrejected. The agent that checks a finding is never the agent that found it. By: Cloudflare. MIT. Source: github.com/cloudflare. - Diffium: a branch-and-diff loop for agents that touch a database. Give the agent a database branch, never production (the author's took 1.2 seconds). Take a baseline, let it work, read the diff. Rows are fingerprinted by primary key plus
md5(t::text). Tables over--row-limit(5000 by default) get a row count and no claims. By: Denislav Gavrilov. Tool MIT. Source: denislavgavrilov.com. - API onboarding descriptor: a file that tells an agent whether it can sign itself up. Nine fields, including
agentPolicy, credential mapping, scope model andgaps, an array of what does not work yet. Mintlify'smint signuptakes--firstName,--lastName,--companyand--email. By: Kin Lane (API Evangelist). Source: apievangelist.com.
Copyable prompts and skills for go-to-market, design and writing
- Hiring-signal prompt: job postings in, a buying-window verdict out as JSON. It detects four signals (
REPEAT_POSTING,FIRST_HIRE,STALE_ROLE,MIGRATION), disqualifies ghost jobs, maps each signal to a buyer and requires a verbatim quote. Confidence below 60 →NONE. Backtest against your last 15 closed-won accounts first. By: Bogdan Liutic (CR8). Source: growthunhinged.com. - Lead Magnet Wizard: one interview prompt. It interviews you about one customer problem, pitches three formats (a calculator, a scorecard, a quiz), then builds the one you pick in the same conversation. The prompt text is on a Notion page linked from the post. By: Ryan Carr (Moodboard). Source: moodboard.beehiiv.com.
- claude-ads: a paid-media skill across 12 ad platforms. Four design choices to copy: source-grounded audits, deterministic scoring (the same account scores the same twice), versioned JSON reports, and capability-gated account changes, where reads are open and writes are gated. By: AgriciDaniel. MIT. Source: github.com/AgriciDaniel.
- The Watchdog playbook: account health in five steps. A reusable Devin skill walks every customer account and gathers activity since the last check (Sentry errors, UX bugs, anomalies). It cuts the findings to the top three problems and marks each as shipped, in progress or open PR. By: Ryan Carson (Untangle). Source: chatprd.ai.
- Consultant playbook: 14 agent skills, each mapped to a cited framework. Task Decomposition runs on the Task-Based Framework (Autor, Levy & Murnane, QJE, 2003). Roadmap Sequencing runs on Weighted Shortest Job First (Reinertsen, 2009). The framework list is free. The build itself is paywalled. By: Guillermo Flor. Source: productmarketfit.tech.
- Claude to Codex:
design.mdas the handoff between two models. Claude turns a Figma file into adesign.mdof tokens. Codex gets "Take this design.md, plus what's in the Figma, and actually build me a technical design system." An optional Figma plugin syncs the coded tokens back. By: Claire Vo. Source: chatprd.ai. - Vercel's
design.md: a public file, a stylesheet and an eval loop. The file took more than 200 runs to build. In a six-page test, pages built withdesign.mdloaded had 57% fewer known layout failures (39 against 91); Vercel says six pages is too small a sample for claims about quality. The step to copy: turn each round of reviewer feedback into a check that fails automatically next time. By: Vercel. Source: vercel.com. - An agent-ready design system: split into files an agent can retrieve.
tokens.jsonanddesign.md, plus registries (components.md,patterns.md,templates.md) that only point at specs. Template files such aswizard.template.mdreference components instead of redefining them. By: not named in the source. Source: medium.com/design-bootcamp. - String Seed of Thought: a prompt that forces real variety into design output. The agent generates a long random alphanumeric string with a shell script and derives colour, layout and typography from it. It never reveals the string. The technique was published by Sakana AI. By: Anshu Chimala. Source: lennysnewsletter.com.
- Grok Bot design workflows: three workflows from voice memos and photos. FigmaBro takes a voice memo plus two Figma screenshots. Artboard structure, spacing and naming conventions are set ahead of time, so informal instructions produce organised files. By: John Bai and Peng Zheng, written up by Claire Vo. Source: chatprd.ai.
- Stripe's Kai skills: a name, a description and trigger phrases. Each skill is reviewed in an editor by its name, description and the phrases that trigger it. Projects add team defaults: a default model, scoped skill sets, and tool policies that require human approval. By: Sharadh Krishnamurthy (Stripe), interviewed by Claire Vo. Source: chatprd.ai.
- Compound Writing: turn your edits into a rules file. Have the model compare its draft with your edit and "Extract only reusable rules that would improve future drafts." File them under Voice, Structure and Content. Save them to one instruction file and load it every time. By: Katie Parrott, Every. Plugin MIT. Source: every.to.
- A self-improving assistant: one weekly task that learns from your edits. Five steps. Step 2 diffs the AI's first draft against the version you actually sent. Then it suggests skills for repeated work, filters techniques from hype, and writes the changes back to the core files. By: Daniel Blum (Melio), published by ChatPRD. Source: chatprd.ai.
- NoBuzz
/debuzz: a second model rewrites the first model's tone. Write the reply to a temp file, pass it to Gemini through the Antigravity CLI (agy -p), and print the output verbatim. Three modes:colleague,manager(about a third the length) anddirector(three to five sentences). By: adnanakil. MIT. Source: github.com/adnanakil. - i-have-adhd: a skill that puts the answer first. A single skill that constrains a coding agent's output: answer first, never a preamble. It is the smallest example of output shape kept as its own named, loadable file. By: ayghri. MIT. Source: github.com/ayghri.
- pstack: 47 skills and 2 subagents as a Cursor plugin. Includes
blast-radius(find what a change could break beyond the diff),automate-me(turns your working style into a personal-modeskill) andbro(restate the last message in plain language). By: Lauren Tan, Cursor. Source: cursor.com.
What is not in here
Essays about agent design, however good; benchmark announcements; product launches with no method attached; and anything whose only source is a newsletter relaying someone else's work. A few entries the collection once held were left out of this edition for that last reason, or because the link no longer let us check the artifact. The library is added to as new frameworks pass the same bar.
Keep reading
The rest of this article: 65 agent frameworks you can copy, each with its source. Several of the best took a rule out of the prompt.
The remaining sections are behind this point.
It costs an email address. The same address goes on Makersfuel, the MakersClaw briefing that runs Tuesday to Saturday, and every issue carries an unsubscribe link.
The address is used for the newsletter and nothing else. Privacy.

