You can hear it, right? Claude / Codex / whoever?
It's waiting for your attention. The red "blocked" indicator (if you're running herdr or something like that).
Or you're out in town and you feel that buzz in your pocket from the mobile app.
I've been getting those calls 24/7 since I got my agent factory up and running.
And no, you can't use it. It's my factory. I mean, you can, but I don't recommend it.
Not because of the mental health toll. That part's worth it. It's just that it's (for now) customized to my own workflows. What is it doing, you ask? I'll get more into that in the next article, along with why it's worth the toll.
I'm shipping many PRs faster (and most of them smaller, which is a good thing). Before the factory, in a good week I was shipping 8 Codagent PRs. In the last week I've completed 103.
Why more? Because a factory is a true loop:

So where is the human in the loop? I'm still trying to sort that out exactly. Reviewing every PR? Obviously not. Some days it goes like this:
Claude: These 15 PRs are ready for you.
Me: Um, do they look good to you?
Claude (a little while later): My comments are addressed. PRs are ready.
Me: Merge them.
The short answer is, I steer. Always steering.
Steering toward what?
Perfecting my agent-runner workflow, from defining a change to implementing, verifying, and accepting it.
Improving instrumentation and analysis so I know what's working and what's not.
Running evals to see which variants yield better performance or lower cost.
All while trying to keep up with the latest damn model releases!
And there's always a fresh alert telling me to put my hands back on the wheel. A batch of changes is done. Something needs to be closely verified, not just stamped. The next big change needs to be spec'd out.
I don't mean just a literal notification; the alerts are in my head too. Wondering what decisions are waiting for me, when I can get back to my computer. Getting back up from bed because I thought of one more prompt I want to run overnight. Trying to squeeze every possible token out of my weekly allowance.
I don't always answer it. Because no, the calls are not toll-free.
But that’s a wrap on this piece. There are… more things to be done.
More smaller PRs and shorter newsletter articles…
Appendix: agent-runner changes since last article
0.4.0
Agent Runner 0.4.0 ships the v2 change workflows, interactive intake, exploratory acceptance, and automatic failure recovery, plus a long run of reliability fixes found in real workflow runs.
Highlights
v2 change workflows. Full and simple workflows for OpenSpec and spec-driven changes run proposal, specs, design, test plan, tasks, implementation, acceptance, archive, and PR finalization, with dedicated Lead, Crosscheck, Implementor, and Tester roles. Spec-driven keeps its artifacts outside the repository. (#58, #155)
Automatic failure recovery. When a deterministic check fails, such as an archive commit rejected by a hook or a task left uncommitted, Agent Runner runs a bounded repair cycle with an agent, re-verifies, and shows the repair attempts and failure evidence in the run view. (#114)
Role-based setup. Setup recommends Lead, Crosscheck, Implementor, and Tester agents from your installed CLIs, favoring model-family diversity, with accept-all or per-role customization.
plannerandreviewerstill work as deprecated aliases. (#57)
Workflows
Writing workflows
Project workflows can call built-in sub-workflows with
builtin:references, and counted loops acceptmax_param. (#89)skip_ifworks on steps inside a group. (#102)agent-runner run --session-dir <path>places a run's session directory where you choose. (#90)Agents can start, poll, and cancel runner-owned agent calls asynchronously. (#114)
Runs and resume
Old run directories are cleaned up automatically, configurable with
run_retentionin settings. (#175)
Usage and cost tracking
Validator runs are instrumented and correlated with run metrics. (#114)
Internal and development
0.3.0
Minor Changes
#51 Add the second-generation OpenSpec change workflows with explicit define, plan, implement, acceptance, and finalization phases.
#52 Replace the interactive PTY proxy with direct terminal handoff for more faithful native agent sessions.
#54 Record durable token usage, estimated cost, and timing metrics across workflow and child-agent execution.
#55 Add synchronous Runner-owned agent calls with named sessions, audit evidence, cancellation, and usage accounting.
#56 Add versioned workflow resolution and explicit reusable orchestration for reviews and targeted acceptance verification.

