paragraph


You can hear it, right? Claude / Codex / whoever?

It's waiting for your attention. The red "blocked" indicator (if you're running herdr or something like that).

Or you're out in town and you feel that buzz in your pocket from the mobile app.

I've been getting those calls 24/7 since I got my agent factory up and running.

And no, you can't use it. It's my factory. I mean, you can, but I don't recommend it.

Not because of the mental health toll. That part's worth it. It's just that it's (for now) customized to my own workflows. What is it doing, you ask? I'll get more into that in the next article, along with why it's worth the toll.

I'm shipping many PRs faster (and most of them smaller, which is a good thing). Before the factory, in a good week I was shipping 8 Codagent PRs. In the last week I've completed 103.

Why more? Because a factory is a true loop:

So where is the human in the loop? I'm still trying to sort that out exactly. Reviewing every PR? Obviously not. Some days it goes like this:

❝

Claude: These 15 PRs are ready for you.

Me: Um, do they look good to you?

Claude (a little while later): My comments are addressed. PRs are ready.

Me: Merge them.

The short answer is, I steer. Always steering.

Steering toward what?

  • Perfecting my agent-runner workflow, from defining a change to implementing, verifying, and accepting it.

  • Improving instrumentation and analysis so I know what's working and what's not.

  • Running evals to see which variants yield better performance or lower cost.

  • All while trying to keep up with the latest damn model releases!

And there's always a fresh alert telling me to put my hands back on the wheel. A batch of changes is done. Something needs to be closely verified, not just stamped. The next big change needs to be spec'd out.

I don't mean just a literal notification; the alerts are in my head too. Wondering what decisions are waiting for me, when I can get back to my computer. Getting back up from bed because I thought of one more prompt I want to run overnight. Trying to squeeze every possible token out of my weekly allowance.

I don't always answer it. Because no, the calls are not toll-free.

But that’s a wrap on this piece. There are… more things to be done.

More smaller PRs and shorter newsletter articles…

Appendix: agent-runner changes since last article

0.4.0

Agent Runner 0.4.0 ships the v2 change workflows, interactive intake, exploratory acceptance, and automatic failure recovery, plus a long run of reliability fixes found in real workflow runs.

Highlights

  • v2 change workflows. Full and simple workflows for OpenSpec and spec-driven changes run proposal, specs, design, test plan, tasks, implementation, acceptance, archive, and PR finalization, with dedicated Lead, Crosscheck, Implementor, and Tester roles. Spec-driven keeps its artifacts outside the repository. (#58, #155)

  • Interactive intake. agent-runner -i opens a conversation with an agent that clarifies what you want, then routes it into the right workflow with that context carried along. (#59, #58)

  • Exploratory acceptance. Acceptance reads the change to decide where to look and probes it like a user would, instead of replaying a checklist written before the code existed. Acceptance and flow-testing steps get their own scratch folder. (#82, #183)

  • Automatic failure recovery. When a deterministic check fails, such as an archive commit rejected by a hook or a task left uncommitted, Agent Runner runs a bounded repair cycle with an agent, re-verifies, and shows the repair attempts and failure evidence in the run view. (#114)

  • Role-based setup. Setup recommends Lead, Crosscheck, Implementor, and Tester agents from your installed CLIs, favoring model-family diversity, with accept-all or per-role customization. planner and reviewer still work as deprecated aliases. (#57)

  • Redesigned run view. New navigation and step detail pane, pull-request links in the breadcrumb, and lower CPU use. (#61, #58)

Workflows

  • finalize-pr keeps one lead session across its CI fix loop, re-verifies CI after the last fix, treats an incomplete review-bot review as a warning, and takes a configurable ci_fix_cycles budget. (#113, #122, #130, #89)

  • verify-change stops before opening a PR when the validator stays red, records the acceptance-round validator result, and refreshes an existing draft PR's body while leaving hand-written bodies alone. (#156, #158, #171, #176)

  • Defects found by the simplify step are routed into fixes instead of being dropped. (#182)

  • verify-task-commit accepts a task verifiably delivered outside the repository, and each task delivery gets a unique record path. (#160, #170)

  • Built-in workflow prompts no longer mandate the TDD skill. (#83)

Writing workflows

  • Project workflows can call built-in sub-workflows with builtin: references, and counted loops accept max_param. (#89)

  • Built-in script steps bundle their sibling helper scripts, fixing failures such as an OpenSpec archive that could not find validate-change-name.sh. (#159, #164)

  • skip_if works on steps inside a group. (#102)

  • agent-runner run --session-dir <path> places a run's session directory where you choose. (#90)

  • Agents can start, poll, and cancel runner-owned agent calls asynchronously. (#114)

Runs and resume

  • Old run directories are cleaned up automatically, configurable with run_retention in settings. (#175)

  • Resume restarts a loop iteration when its loop variable changed, and resuming nested or failed workflows is more robust. (#133, #58)

  • Resumed Claude sessions receive the step prompt as a user message, headless Claude tasks stay in the foreground, and step completion from headless attempts is rejected. (#149, #154, #162)

Usage and cost tracking

  • Run metrics include Claude subagent usage, per-step attribution of cumulative Codex and Claude session usage and cost, uncached input tokens for Codex, and a reason for skipped agent steps that were never invoked. (#194, #143, #146, #169, #145)

  • Validator runs are instrumented and correlated with run metrics. (#114)

  • Per-step changed paths and git attribution are accurate: preexisting dirty or unchanged paths are excluded and committed renames are handled. (#131, #132, #134, #135)

Internal and development

0.3.0

Minor Changes

  • #51 Add the second-generation OpenSpec change workflows with explicit define, plan, implement, acceptance, and finalization phases.

  • #52 Replace the interactive PTY proxy with direct terminal handoff for more faithful native agent sessions.

  • #54 Record durable token usage, estimated cost, and timing metrics across workflow and child-agent execution.

  • #55 Add synchronous Runner-owned agent calls with named sessions, audit evidence, cancellation, and usage accounting.

  • #56 Add versioned workflow resolution and explicit reusable orchestration for reviews and targeted acceptance verification.

Keep Reading