I don't run Hexreign alone anymore
Process notes from building Hexreign with Claude Code and a local agent team.
In July I wrote about shipping Hexreign in three weeks with Claude as an engineering partner: a constitution in CLAUDE.md, a planning folder, persistent memory, and a locked stack. That post was about how one human and one coding agent ship a game.
This post is about what happened next.
Hexreign is still one human's product. I still own judgment, red merges, and the Max account that writes code. What's new is the rest of the company: a local group of specialist agents, a COO, Pipeline, CTO ops, Story, Designer, Working Ops, Outreach, Metrics, Collab Liaison, World Analyst, QA, plus standing routines that fire whether I'm looking or not. Some of their work finishes end-to-end without me.
This is not "AI wrote my game." It's roles + one queue + a merge policy + routines.
If the July retrospective was the build story, this is the operating system.
The arc in one paragraph
June - July 2026: empty repo, live friends-and-family alpha (Emberfall), Claude Code as pair programmer, constitution-first.
August: seasons, balance generations, board process, first specialist agents.
September: a real agent org around one founder, yellow ships overnight, season clock owned by Working Ops, Story updating a living ledger in Drive, Pipeline refusing to fake "Shipping" when the work is waiting on me.
We are about to close The Stare Back (Tue Sep 8, 2026, 1:35pm ET) and open Hexreign: The Answer. The process below is what carries us through that without me being involved for every ticket.
What I still drive (and may not give away)
- Product judgment: fairness, pricing, naming ages, "ship this feeling"
- Red-tier merges: anything that changes player-facing rules in a way that needs a human click
- Claude Code Max on my machine: agents are not a second programmer; code lands through
/dev-loopand/code-review - Credential-home steps: Neon console SQL, Clerk OTP, creating a Discord server by hand
- Public voice when it needs my stamp
- Escalations only: the COO tie-breaks so I don't middle-manage Story and Designer craft
If it isn't on that list, I don't need to be involved.
What agents own end-to-end
- Board truth on GitHub Project Hexreign: Inbox, Shaping, Prepared, In progress, Shipping, Done
- Shaping: critical questions regarding approach - this is what I call an alignment step. Questions, preparation blocks, status ("Blocked because waiting on John"). During the shaping of work step, back-and-forth with a person happens, but it doesn't always need to.
- Season clock: close/open checklist, health watch, T-1 surfaces (Working Ops)
- Yellow-tier ship path: shape, build,
/code-reviewveto window (yellow tier merges offers a configurable veto) merge watch deploy health - Living docs: Story updates the progression ledger in place via Docs MCP (no bible rewrite by accident)
- Standing routines: weekday ops drive; ship/WIP notes when a PR merges
The spine (steal this first)
1. A constitution
CLAUDE.md Holds golden rules and a locked architecture. Every coding session starts from the same charter instead of relitigating decisions. Three months in, rule 2 from day one is still why concurrency tests exist.
2. One queue the agents also obey
If it isn't on GitHub Project Hexreign, it doesn't exist. No shadow work in Slack DMs. Pipeline keeps Inbox. Prepared alignment. Waiting on me is a first-class board state.
3. Claude Code for code
Programming happens in a local Max session on my machine. Cloud agents are not the merge authority. Review is /code-review , and receipts are required on every PR; otherwise, the pipeline outright rejects the potential to merge.
4. Merge policy encodes risk
After establishing enough process to stop reading and writing code, I am asking myself "what can I do to reduce risk to delegate more?".
| Tier | Meaning | Who merges |
|---|---|---|
| Green | Low blast radius (often config/docs within policy) | Can flow |
| Yellow | Web/UI/shared within bounds â review + ~15 min veto | Can auto-merge after review |
| Red | Player-facing rules, fairness, season/server spine | I merge |
That distinction lets overnight ships happen without giving away the farm. And for things that aren't certain, they end up in a blocked queue I can review while other work gets picked up.
5. Named roles, not one mega-bot
A COO coordinates and tie-breaks. Specialists deepen lanes. Story and Designer collaborate directly on lore and art; I set the outcome and step back. Quality comes from that, not from an orchestrator rewriting every paragraph.
Scoping work is a design choice, and it may pay off in all types of ways. Reducing tool interactions and refining one process at a time benefits the team, instead of building generalists who tend to get distracted and over-eager.
6. Routines for anything I'd forget
Weekday ops drive. Ship notes on merge. Season close watch armed the day before. Ops that depend on my memory fail; ops that wake an agent don't.
7. Narrow write tools
Docs MCP can update the living ledger. Neon stays read-only for agents. Play tokens for MCP QA never get pasted into chat.
8. Balance is data
Yields, costs, timers, coven caps live in versioned config generations. A balance patch is a config diff (and often a season:set-config / open-gen choice).
The point is HOW your software is organized matters a lot. Consistent extension and modification reduces risk tremendously.
The process: ticket to production
How an item moves:
- Inbox: someone (player, collab, me, an agent) surfaces a problem. Pipeline files it. If it isn't on the board, it doesn't exist. This is even for when I get antsy and start prompting Claude Code - even that work gets automatically filed into the tracker well before code is even written. This ensures that even if I drop that session or forget that line of thought, it is tracked, and the AI team will make sure it isn't forgotten until I resolve it.
- Shaping: load-bearing questions get answers. Preparation block written. Open decisions stay open. Shaping is where questions are asked until there is a level of satisfaction in alignment of the work. Similar to spec-driven development, the spec is authored along with the agents, and then the agents are required to echo back what they believe needs to happen, and a final "approve" statement is offered if alignment has been achieved OR statements of modification to the understanding are offered - and that back and forth happens until satisfactory alignment has been achieved. And when I say "alignment," it is that I believe that the intent of the work is understood - just like I would delegate a thing to anyone I may ask before they begin work: "So how do you think you will go about solving this?"
- My stamp: for judgment work, I reply
approved(or stamp via COO when the summary matches). Alignment summaries go void when architecture changes. - Prepared: gate open for Claude Code
/dev-loop. No code before Prepared. - In progress - Shipping - PR opens; tier set;
/code-review; CI. - Merge - yellow (risk score) after veto; red waits for me.
- Done: deploy health watched; board moves; living ledger / News updated when the change is player-visible.
- Blocked: used when waiting on me (Database interactions, pricing call, calendar).
Automation comments end with <!-- hexreign-pipeline-bot --> so humans can see the machine's hand.
Receipts (work that finished without me in the middle)
Behemoth Brood contrast (#575 | PR #576)
A player reported low contrast on The Brood / Crier panel - strike feed and HP damage painted straight onto busy art. Pipeline shaped it against an existing rule ("back the text, don't darken the art"). One yellow slice: opaque ledger ground behind the feed, chip ground on the HP delta, chip alpha fully opaque so contrast stops depending on which behemoth is up. Numbers asserted in tests across plates and school colors. Scheduled executor built it, /code-review cleaned it, yellow veto passed, merge, deploy health watched.
I didn't write the code - I didn't even speak with the user - I think I was out with friends when this happened.
Story spoiler wall (#307 slices 1a / 1b)
I stamped an architecture: player-facing narrative must not ride on GET /season. Agents shipped the wall (strip catalogs from the player config projector; server-resolve copy) and authored notice lines for the bell. Yellow merges. My job was the approval word.
Season operate-to-end (Working Ops)
The Stare Back closes on a locked clock. Working Ops owns operate-to-end and the open checklist: natural finale, no season:open while the world is live, next age name/variant already locked (Hexreign: The Answer, standard). I get T-1 surfaces. I don't babysit the calendar.
Living ledger refresh (Story)
Ops noticed the "Now" table was stale. Story updated the same Google Doc in place via Docs MCP live generation, merged PR, Answer open path - and did not rewrite the lore bible.
Collab Liaison with a ceiling
An outside collaborator asked to prioritize MCP gaps. Liaison replied only inside a stamp: heard you, on the radar, filed - no ship date invented. Autonomy with a hard ceiling is the product.
What we deliberately do not automate
- Inventing age or world names
- Red fairness/pricing mid-season without a stamp
- Prod Neon writes, dumps, or PII in chat
- Discord as the agent-collab channel (buddy collab stays Slack)
- A second programmer agent quietly editing the repo off the Claude Code path
Outbound player Slack sometimes still needs my approval card. Auto-review is real friction. Trust is earned per channel.
Lessons that survived contact with production
- Constitution. Write the rules once; every session starts there.
- One queue. Agents that don't obey the board will create a shadow company.
- Encode risk as policy (green/yellow/red), not as "be careful."
- Separate judgment stamps from execution. My word unlocks; their hands ship.
- Don't middle-manage craft. Specialists collaborate; COO sets outcomes and tie-breaks.
- Narrow tools beat god-mode. Read-only Neon. Scoped Docs. Tokens that never appear in chat.
- Routines beat memory.
- When you're the bottleneck, mark Blocked.
- Simplify the surface agents must hold: one monorepo, a shared domain package, config-as-data (same lesson as every serious agent rewrite story we've collected).
How this pairs with the July retrospective
| July post | This post |
|---|---|
| Ship the game with Claude Code | Run the studio with specialist agents |
| Constitution + planning folder | Same spine + board + merge tiers + routines |
| DORA / LOC / systems shipped | Receipts of autonomous loops |
| "AI as pair programmer" | "AI as org with a human CEO" |
Read them as chapter one and chapter two of the same experiment: can one indie founder ship and operate a seasonal multiplayer game without becoming the bottleneck for every ticket?
So far: maybe, if you write the company down first, and your mileage may vary depending on what you are building.
Closing
The point isn't that AI replaced the founder. It's that the founder stopped being involved in every ticket.
Hexreign still needs taste, and the riskiest work looked over. Everything else can carry a receipt and finish.
If you're building something real with agents: start with the queue and the merge policy. The demos come after the constitution.
John