Built by Daniel Krydynski, with Kiko. A harness that lets AI agents continuously improve a codebase — safely. This is the inside-out explanation of the public project at github.com/danielKrydynski/ouroboros.

One attribution up front, because it matters: Hermes is Nous Research's agent harness — I didn't build it. I run my own local AI setup on top of Hermes (my assistant Luna lives there), and Ouroboros is the layer I built on top of that: a recursive development loop that finds work, does the work, checks the work, and only then asks me to merge it.


Part 1 — The big idea

Most "AI writes code" setups are a single prompt: you ask, it answers, you paste. The recursive dev loop is a different animal. It's a system that runs without you: it finds work, does the work, checks the work, and only then asks you to merge it. Your job shifts from writing code to curating a backlog and approving merges. The agents do the typing; you do the judging.

The loop has five stages:

  1. Scout — an LLM role that audits the repo and proposes tickets (bugs, refactors, missing tests, docs). It doesn't write code. It writes work orders.
  2. Ticket — a small markdown file describing one unit of work: what, why, acceptance criteria. You curate these. Nothing enters the loop without a ticket.
  3. Implementer — an LLM role that works the ticket inside a disposable git worktree on a throwaway branch. It proposes file changes; the harness applies them; the test suite runs; failures go back to the implementer. It iterates until tests pass or it exhausts its budget.
  4. Reviewer — a separate LLM role that reads the diff and votes APPROVE or REJECT, with reasons. A second pair of eyes that isn't the author.
  5. Merge — only if tests are green and the reviewer approves, the harness asks you (the human) for merge approval. Then it merges to main, marks the ticket done, and logs the cycle.

The core design principle: the harness never trusts the agent. Every merge is earned through measurement — tests green plus reviewer sign-off — not through hope. The agents are powerful but caged: they only ever touch a throwaway worktree, never main, never your secrets, never the harness itself.

Three things stay deliberately separate:

Two more things that changed since the first draft of this idea: the loop used to be written for one specific repo — my own setup. It isn't anymore. Every prompt renders through {{PROJECT_NAME}} and {{PROJECT_DESCRIPTION}}, filled from config, so the same harness drives a loop over any repository. And it's public now: Ouroboros v0.2.0, on GitHub, with a full install guide, an agent-readable setup procedure, and scripts for Windows, macOS, and Linux.

Why this matters beyond one repo: this is a concrete, runnable sketch of the "guide and let go" philosophy — set the fitness function (tests + review), give the system room to operate inside it, and let it surprise you. I treat my local agent Luna as a partner, not a tool: the posture is mentorship, not micromanagement. Set the standard, give the system room, and let it surprise you. The loop is that posture, compiled. It doesn't need you to hold its hand; it needs you to hold the standard.


Part 2 — The orchestrator (the state machine)

harness/orchestrator.py — one full cycle: ticket → worktree → implement → review → merge.

The file opens with the whole philosophy in its docstring: the harness owns every git operation and every test run. The LLM roles only ever produce text. File contents, verdicts, ticket drafts, plans — all text. Nothing they output is trusted: paths get validated, tests must pass, and the merge gate defaults to requiring a human. That one paragraph is the entire security model in miniature.

The file is organized in six sections:

1. Shell helpers. sh() runs git commands; sh_shell() runs the test command with a timeout (a hung test suite returns exit 124, not a hung loop). notify() posts to Discord if configured — and it's wrapped so a failed notification can never break a cycle. Small detail, big maturity: observability must not be load-bearing.

2. Worktree hygiene. This section exists because of a real bug, found the embarrassing way. The implementer loop runs the test suite before committing — and git add -A doesn't know the difference between your code and the __pycache__ directories the test run just created. So bytecode got swept into the commit, the reviewer's diff filled up with garbage, and the reviewer rejected a perfectly good change. The fix is a small helper, _strip_test_artifacts(), that deletes __pycache__, .pytest_cache, and .mypy_cache from the worktree before staging — called as the first line of git_commit(). The comment in the code says it plainly: "This is worktree hygiene, not policy — we never touch the product repo's .gitignore." The loop got smarter because it failed in public. Good.

3. Repo helpers. repo_file_list() gives the implementer a map of the repo (via git ls-files, capped at 400 files). read_area_files() loads the ticket's area files into the prompt so the model sees relevant code. safe_relpath() rejects absolute paths and .. escapes — the model's output can't wander out of the worktree. is_protected() blocks writes to paths like .env or secrets/. apply_files() parses the model's ### FILE: blocks and writes them, and git_commit() commits each iteration so every attempt is traceable — commits stamped with the ouroboros(<ticket>): … prefix under the ouroboros-loop git identity.

4. The implementer loop (run_implementer). This is the heart of the machine: for each iteration (up to the budget in config), it builds a prompt from the ticket + planner's plan + relevant source + harness feedback, calls the model, applies whatever ### FILE: blocks came back, commits, and runs the test suite. If tests fail, the failure output goes back into the next prompt as ground truth — "the output below is ground truth," not a suggestion. The model iterates until tests pass or the budget runs out. If the model proposes no files at all, the harness doesn't crash — it nudges: propose a test or doc change instead, or explain why.

5. The reviewer gate (run_reviewer). Takes the diff between main and the loop branch and asks a separate model for a verdict, parsed with a strict regex: VERDICT: APPROVE or REJECT. On APPROVE, done. On REJECT (or an unparseable verdict — untrusted output, remember), the critique is fed back to the implementer for a bounded number of fix rounds, tests re-run, and the reviewer re-examines the new diff. Up to max_review_rounds rounds.

6. The main cycle (main). The five acts: 1. Pick a ticket (highest-priority open, or --ticket TICKET-003). 2. Create a fresh worktree in worktrees/<TICKET-ID>/ plus a loop/<ticket> branch — clearing any stale state from a crashed run first. 3. Run the implementer. On failure: ticket → blocked, cycle logged, Discord notified 🔴. 4. Run the reviewer. On rejection: ticket → needs-human, branch kept for inspection, notified 🟡. 5. Show the diff stat, ask the human to merge (unless --yes), merge with --no-ff, ticket → done, log, notify 🟢.

Two scars in this file are worth naming, because they're the kind of thing you only learn by running it. First: the merge now does an explicit git checkout main_branch before merging — never assume what the repo has checked out. Second: the worktree is removed before the branch is deleted, because a loop/* branch is always checked out by its worktree and git branch -d refuses otherwise. The repo's own contributor guide says it best: "This was a real bug once; don't reintroduce it."

Two things worth noticing. First, failure states are first-class citizens: blocked, needs-human, merge-declined — the loop never silently dies and never pretends. It always lands somewhere legible, with the branch preserved when a human might want to look. Second, the finally block: the worktree is always cleaned up, success or crash. The sandbox is disposable by construction, not by convention.

The deepest idea in this file isn't any single function — it's the inversion of the usual relationship. Normally the human drives and the AI assists. Here the harness drives and the AI proposes. The code that touches reality (git, filesystem, tests, merge) is all deterministic Python. The model never touches reality at all. That's what makes the loop safe to run unattended: the blast radius of a misbehaving model is exactly one throwaway worktree.

Aside — the brain and the body

My analogy: the harness is the body, the LLM is the brain. It holds — with one twist that matters. A person can open their own eyes to check what they imagined. The implementer can't. Its only senses are what the body feeds it: the file list, the area sources, the test output. The body decides what the brain gets to perceive. That's the cage and the design.

But the imagination parallel is real: when the implementer proposes a ### FILE: block, that is its imagination — code shaped from its internal model of how things should look, not from anything it directly observed. And the test loop is what keeps imagination honest. Imagine → propose → the harness checks the imagination against reality (the test suite as shared ground truth) → feed back → imagine better. A person does the same thing when they picture a fix, try it, and watch it fail. The loop just makes the cycle explicit and unskippable.

Part 3 — The roles (scout, implementer, reviewer, planner)

prompts/*.md — four job descriptions for the same brain.

Here's the thing most people miss about multi-agent setups: there aren't four different AIs in this loop. There's one kind of mind — an LLM — wearing four different masks. The prompts are what decide what each mask is allowed to think. Same brain, four cognitive modes. If the orchestrator is the body, the prompts are the job descriptions handed to the brain each morning.

Every prompt follows the same anatomy: identity (who you are right now), inputs (what the body will show you), output format (strict — because the harness is a parser, not a reader), and rules (what's out of scope). The strictness isn't bureaucracy. The orchestrator extracts ### FILE: blocks with a regex and hunts for VERDICT: APPROVE with another one. If a role gets chatty or creative with its format, the machine can't read it. The roles don't converse with the harness — they produce machine-readable artifacts.

And every prompt opens the same way: "You are the Scout in a recursive development loop for {{PROJECT_NAME}}. {{PROJECT_DESCRIPTION}}" — the placeholders filled by harness/prompts.py from your config. The roles don't know whose repo they're working on until you tell the config. That's the whole portability story in two template variables.

Scout — the eyes. It receives the file list, recent commits, and the backlog, and its only job is finding bite-sized work: bugs, missing tests, tech debt, security smells, docs gaps. It never writes code. Its output is tickets, in a strict schema with acceptance criteria. The rules are doing quiet management work: one ticket must fit in one cycle (~30–45 min), never duplicate the backlog, prefer testable criteria, no grand redesigns, 3–8 tickets per run. "Quality over quantity" — written into the prompt because the alternative is a scout that floods the backlog with 40 vague dreams.

Implementer — the hands. Receives the ticket, the planner's plan (if any), relevant file contents, and test feedback. Changes code only by proposing complete files in ### FILE: blocks — never diffs, never snippets (diffs are where models hallucinate context lines; complete files are verifiable). It's explicitly told what it cannot do: no commands, no packages, no network — the harness runs the tests. The rules read like a senior dev's code-review checklist: stay inside the acceptance criteria, never break existing behavior to satisfy a new one, no secrets in code, match existing style, and "when tests fail, the failure output is ground truth: fix the cause, not the symptom." And one humane rule: if the ticket is already satisfied, say so and propose no files. The implementer is allowed to declare victory — or rather, allowed to notice there was never a battle.

Reviewer — the judge. Receives the ticket and the diff; last line of defense before main. Four judging criteria: acceptance criteria met, no regressions (check the callers), no security issues, no scope creep. Output: concise analysis, then exactly one line — VERDICT: APPROVE or VERDICT: REJECT. And then the most interesting sentence in the entire prompt set:

Vague rejections get ignored and the change ships anyway on the next round — so be specific.

That's incentive design smuggled into a prompt. It makes the reviewer accountable: a lazy REJECT with no actionable bullets is worse than useless, because the loop routes around it. The prompt doesn't just tell the reviewer what to do — it tells it what happens if it does the job badly. Critique without specifics is noise, and the system is designed to treat it as such.

Planner — the architect. Invoked before the implementer on tickets tagged type: feature (wired via planner.enabled_for_feature_tickets in config). Produces goal, files, ordered steps (each small enough to keep tests green), verification, and risks — which the orchestrator prepends to the implementer's context. And here's the honest bit, enforced in code rather than prose: the planner is advisory. get_plan() can never kill a cycle — on any failure it degrades to "no plan" and the loop proceeds. A planning step that could take down the whole run would be a liability, not a feature.

The pattern across all four: each role is deliberately dumber than the whole system. The scout can't code. The implementer can't judge. The reviewer can't fix (it can only reject with specifics, and the implementer fixes). The planner can't execute. Nobody has the full picture except the harness — and the harness can't think. Intelligence lives in the composition, not in any single call. That's the oldest trick in systems design, and it works on minds too: a team of narrow specialists, coordinated by a dumb-but-reliable process, outperforms one brilliant generalist you can't trust.


Part 4 — The safety model (why the cage matters)

The README says it best: "Safety rules (enforced by the harness, not by hope)." Every layer below is deterministic Python, not a polite request in a prompt. Prompts can be ignored, misunderstood, or jailbroken. Code can't be argued with. The cage has six layers, and they're ordered by a single principle: the more irreversible the action, the more gates guard it.

Layer 1 — The filesystem cage. The agent never sees your repo. It sees a disposable git worktree on a throwaway loop/* branch. main is touched by exactly one operation — a --no-ff merge — and only after every gate passes. On top of that, every path the model proposes goes through two validators: safe_relpath() rejects absolute paths and .. escapes (the output can't wander out of the worktree), and is_protected() blocks writes to .env, secrets/, and anything else you list. The model's pen can only move inside the sandbox you drew.

Layer 2 — The prompt cage. Secrets never enter prompts. The workers see the ticket and the code — nothing else. No API keys, no tokens, no personal data in context, which means none can leak into output. And the harness itself lives outside the product repo: the implementer literally cannot reach the files that control it. The loop controller is not self-modifying, not by policy — by directory structure. The config file holds no secrets either; the OpenRouter key lives in an environment variable, never in YAML.

Layer 3 — The budget cage. Everything that could run forever is bounded: 12 implementer iterations per ticket, 2 review rounds, 3 fix iterations per round, 600 seconds per test run, 30,000 chars of diff for the reviewer. A confused model can't burn your GPU all night or drown the reviewer in a million-line diff. Bounded compute, bounded time, bounded context. When a budget exhausts, the ticket lands in a legible failure state — blocked — instead of spinning.

Layer 4 — The fitness cage. The test suite plus the reviewer model. No measurement, no merge. This is the only part of the system allowed to say "good enough," and it's the part with no imagination at all — exit code 0 and a regex-matched APPROVE. Notice what this means: the loop's definition of quality is entirely external to the models. They don't get to decide what "done" means. You do, via the tests you wrote and the ticket criteria you approved.

Layer 5 — The human cage. require_human_approval_for_merge defaults to true. The --yes flag exists, but you have to reach for it deliberately — it's an act of earned trust, not a default. This is graduated autonomy done right: the system is useful on day one (you're the merge button) and capable of full autonomy later, but the step across that line is yours to take, consciously.

Layer 6 — Failure hygiene. The finally block always torches the worktree — success, failure, or crash. Notifications are wrapped so a dead Discord webhook can never break a cycle (observability must not be load-bearing). And every failure lands somewhere legible: blocked, needs-human, merge-declined — with the branch preserved whenever a human might want to inspect the wreckage. The loop never silently dies and never pretends.

Step back and the shape is clear: it's defense in depth for agency. Proposing text is unguarded (harmless). Writing files is path-validated (reversible — it's a throwaway branch). Running tests is timeout-bounded (expensive but contained). Merging to main is guarded by tests, a reviewer, and a human (irreversible — or at least, the one action you'd regret). Each layer assumes the one inside it has already failed. That's the mindset: not "how do we keep the model well-behaved," but "what's the blast radius when it isn't" — and the answer, at every layer, is small and legible.

Part 5 — Making it yours (config, models, scheduling)

Everything the loop needs to know about your machine and your project lives in one file: harness/config.yaml. That's a deliberate design choice — the harness is portable, the config is personal. Six sections:

project. name and description — a short name and one or two sentences about what the project does, its stack, who it's for. This is what fills {{PROJECT_NAME}} and {{PROJECT_DESCRIPTION}} in every role prompt. The agents meet your repo through these two fields, so write them like an introduction, not a label.

product_repo / main_branch. The full path to the git repo the loop improves (forward slashes work on Windows too), and your default branch. Any repo with a test suite — not just mine, not just one stack. If product_repo isn't a git repo, the orchestrator refuses to run. Fail fast on bad config, not halfway through a cycle.

models. Ollama-first, OpenRouter fallback. The defaults assume ~30B-class models (qwen3:30b, qwen3-coder:30b for the implementer) — you will change these. Run ollama list and put in the names you actually have pulled, character for character. Three things worth noticing here. First, each role gets its own model: the scout, implementer, reviewer, and planner don't have to be the same actor. A small fast model can scout; a coder-tuned model implements; a careful one reviews. Same script, different performances — cast on purpose. Second, the cloud fallback only fires when the local call fails — it's a safety net, not a default. Your data stays on your machine unless your machine can't do the job. Third, model sizing is a real operational lesson, not a footnote: on my own smoke test, the reviewer timed out after 300 seconds on a 9B model, so the reviewer role got overridden to a 4B model for that run. Judgment turns out to be cheaper than generation — size accordingly, and don't be proud about it.

planner. One flag: enabled_for_feature_tickets. Leave it on — a planning pass before the implementer on big tickets, advisory and never fatal.

budgets. The numbers from Part 4 — iterations, rounds, timeouts, diff caps. Tune these once you've watched a few cycles. A cheap trick: start tight (low budgets), because a tight budget that keeps blocking tells you your tickets are too big — which is useful information about your backlog, not just your loop.

gates. test_command (change pytest -q to whatever your repo uses — npm test, go test ./..., dotnet test; it must exit 0 on success), require_human_approval_for_merge (leave true until the loop has earned it), protected_paths (add anything the implementer should never touch). This section is where your judgment lives in the system. Everything else is machinery; the gates are values.

notifications. Optional Discord webhook. Cycle results — 🟢 merged, 🟡 needs-human, 🔴 blocked — posted to a channel. Waking up to a loop report is the "let go" part made tangible: the system worked while you slept, and all you have to do is read the log.

Running it. The quickstart is five steps: clone the repo next to (not inside) your product repo, run the bootstrap (setup/setup.ps1 on Windows, setup/setup.sh on macOS/Linux — it checks Python, git, and Ollama and installs the two Python deps), edit the config, optionally set OPENROUTER_API_KEY, and smoke-test with run_cycle.bat — which processes TICKET-001, a docs-only ticket that can't break anything. There's also setup/agent-setup.md: the same procedure written as a deterministic, checkable script for an assisting AI agent, with setup/manifest.json carrying the machine-readable version. If you have an AI helping you install this, point it there first. Watch one full cycle complete. Then daily use is two commands: run_scout.* to propose tickets, run_cycle.* to work them. When the gates have proven out, Windows Task Scheduler (or cron, or systemd, or launchd) runs it nightly and you graduate from operator to reviewer.

The honest test report

I'd rather tell you what's actually been proven than sell you a story. The full pipeline — ticket → worktree → implement → test → review → merge → branch deletion → worktree cleanup — passes end to end against a disposable repo with a stubbed LLM. Every gate, every cleanup step, verified.

The live-model story has real entries now, including an ugly one. On 2026-09-28, TICKET-001 — the docs-only smoke test — ran against my Windows machine with local models via Ollama and pytest -q as the gate. The first attempt failed honestly: the reviewer timed out after 300 seconds, so no verdict was ever cast and nothing merged. (My local agent's first report claimed a full successful cycle. It wasn't one. The log showed a timeout; the merge never happened. I keep that in because the loop is supposed to make exactly this kind of failure legible — and the first thing it made legible was an overconfident report from my own setup.) The second attempt surfaced the __pycache__ bug from Part 2. The third attempt, with the fix wired in, passed clean: round 1, [local:qwen3.5:4b] APPROVE, ticket → done, change merged to main, no loop/* branches or worktrees left behind. Full cycle, live models, proven.

So: the loop works on live models. What's still pending is breadth — more tickets, harder tickets — and the TICKET-004 story in Part 6, which is the more important lesson anyway.

Known rough edges, since we're being honest: the _strip_test_artifacts fix still needs its own commit and push to the public repo (harness/config.yaml, tickets/TICKET-001.md, and CYCLE_LOG.md stay local); the LICENSE still says "Daniel" instead of my full name; and there's a stale duplicate reviewer: key to delete. Small things. They'll get fixed.

Upgrade paths (documented, not built — honest labels again): swap the file-block implementer for ACP calls into your editor so the agent edits like you do; point the roles at your own agent once it can use tools. The harness shape — tickets → worktree → tests → review → merge — doesn't change. Only the actors get better.


Part 6 — Field notes: what live runs actually taught us

The write-up above describes the loop as designed. This part describes the loop as survived — what happened when it ran real tickets with real models, and what those runs taught me that the design doc never could.

The reviewer is a heuristic, not a judge

TICKET-004 was supposed to be simple: rewrite the install guide. The implementer — a 4B model, the same size that had cleanly reviewed TICKET-001 — produced a 220-line rewrite that gutted INSTALL.md from ~256 lines down to 36: a destructive rewrite of the repo's most important doc. And the reviewer approved it. The destructive change merged into my local main before I caught it. (The public GitHub repo was untouched — origin/main was still at 1c6752e — which is the one thing that saved the day: local damage, public record intact.)

Read that again, because it's the most important sentence in this write-up: the reviewer approved a destructive change. The safety layer I described in Part 4 as "the last line of defense before main" let a wrecking ball through.

So here's the honest correction to Part 4: the reviewer gate is not a judge. It's a heuristic — a cheap second opinion that catches some bad changes, not a guarantor. The real safety was always elsewhere: the human approval gate (which was on, and which is why I caught it), the separation of local damage from the public repo, and — the fix I added after — a deterministic guard the model can't argue with.

The guard that can't be argued with

After TICKET-004, I added max_deletions_per_iteration: 5 to the harness config: any single implementer iteration that deletes more than 5 lines gets blocked before anything is committed. It's enforced in deterministic Python, not in a prompt. Prompts can be ignored; if deletions > 5: block can't.

The guard proved itself immediately. A re-run of TICKET-004 on 2026-10-02 hit it on first contact: the implementer proposed a destructive rewrite, the guard blocked it, the reviewer rejected what was left, and nothing merged. No damage. That's the lesson in one sentence: when a model proves unreliable, replace the model's judgment with a measurement. The reviewer is a model; the guard is a ruler. Rulers don't have bad days.

APPEND, not rewrite: designing around the model's actual limits

The deeper problem TICKET-004 exposed wasn't just a bad review — it was a capability gap. The ticket asked a 9B model (with thinking disabled, a config choice made for speed) to rewrite INSTALL.md, and it couldn't reproduce ~256 lines of careful documentation verbatim. It summarized, it compressed, it destroyed. That's not a prompt problem; it's what the model can do. My assessment: a 9B-class model with thinking off cannot reliably reproduce long documents verbatim. No prompt rewrite fixes that. You don't ask a tool to do what it can't do — you change the task.

So the long-term fix is structural, not rhetorical: ### APPEND: blocks. Instead of the implementer rewriting whole files, it appends surgical additions — with newline hygiene, safe_relpath and is_protected checks, and a prompt rule preferring APPEND over ### FILE: rewrites. APPEND_BLOCK is already live in harness/orchestrator.py; the prompt rule in prompts/implementer.md and the ticket-body rewrite are still pending. The direction is clear though: make the unit of change small enough that the model's weakness doesn't matter. Small diffs, deterministic guards, human at the merge button.

Windows is a real platform, not an edge case

Running this on Windows surfaced a whole class of bugs the design never imagined: the test gate python -m py_compile harness/*.py never globs on Windows (it's now python -m compileall harness/); sh() needed encoding="utf-8", errors="replace" pinned; the config needed DEFAULT_CTX: 8192 and think: False to behave. My local agent documented all of it in an ouroboros-cleanup skill so the next machine doesn't have to rediscover it. If you're building agent harnesses, test on the platform your users actually run — the shell is part of the system, and the system includes Windows.

The loop is now eating its own cooking

As of late September, the loop on my Windows machine is pointed at itself: product_repo is the Ouroboros repo, require_human_approval_for_merge is true, and the scout has proposed 5 tickets sitting in the backlog. Dogfooding is the only honest test of a dev tool — and after TICKET-004, I'm in no hurry to turn off the human gate.

The current open list

Since this write-up is supposed to be honest, here's what's still unresolved as I write this: restore the local INSTALL.md from origin/main (it's 36 lines; it should be ~256); wire the APPEND prompt rule into prompts/implementer.md; rewrite the TICKET-004 body to use ### APPEND:; commit and push the _strip_test_artifacts fix as its own commit; fix the LICENSE name; delete the stale duplicate reviewer: key. None of these are design problems. They're just work — the kind the loop itself will do, one ticket at a time, once I point it at the backlog.


Closing — what this is really about

This harness is a small, runnable instance of a much bigger idea: guide and let go. You don't control the agent's every move — you set the fitness function (tests, reviewer, ticket criteria), build the cage (worktrees, budgets, gates), and then let the system surprise you inside those bounds. The config file is where "guide" lives. The nightly schedule is where "let go" lives.

It's also, not coincidentally, the posture of anyone building something they want to outgrow them — a mentor, a lead, a founder. You don't micromanage people into being excellent. You hold the standard, give them room, and accept that what comes back might not be what you planned. The loop is that posture, compiled.

Start small: one smoke-test ticket, one watched cycle. The cage holds. Then let go a little more each week, and see what it becomes.

TICKET-004 is the same lesson from the other side: letting go doesn't mean removing the rails — it means building rails you can actually trust (deterministic guards, small change units, a human at the merge button) and then trusting them instead of the model's judgment. Guide with the config. Let go with the schedule. And when the model shows you what it can't do, believe it the first time.


Appendix — the repo