Case study · Mike Butz
Human-directed AI development: the Unproven workflow
How I directed AI-assisted development of Unproven’s Inventory Core, from bounded requirements to independent review and human acceptance.
I directed the accepted IC01–07 Inventory Core work for Unproven. The result combines a plain C# domain model with a development process that gives coding agents specific tasks and retains human judgment over acceptance. It is an implemented subsystem in a game prototype, rather than a production game release.
The problem
An inventory operation has to keep identity, ownership, storage and equipment consistent. If a request cannot proceed, it also needs an exact, useful reason. As features accumulated and work moved between AI contexts, I needed those rules to survive the handoff.
The workflow was part of that engineering problem. A broad request to “make inventory work” was insufficient: requirements, allowed changes, validation and acceptance all needed to be explicit.
My role and the agent roles
I owned the player experience, priorities and scope, including requirements such as five armor regions and a necklace plus two ring slots. I worked with ChatGPT on architecture and requirements, then used Codex for implementation. A separate AI reviewer context compared the result with the requirements. I retained the human validation and acceptance decision.
Each task identified its baseline, active slice and stage, requirements, permitted changes, tests, stop conditions and Git permissions. One production slice stayed active at a time. High and medium review findings blocked progression until correction and fresh review.
- Define and reviewMike sets goals and scope; requirements and design are reviewed.
- Implement and validateA coding agent works within the task boundary and records focused and regression results.
- Independent reviewA separate context checks requirements, code and evidence.
- Correct and recheckA failed requirement leads to a bounded root-cause correction and fresh review.
- Human acceptanceMike validates the result before it enters the accepted baseline.
Repository instructions and human gates guide this process. They are not an independently enforced autonomous-agent platform, and I do not claim to have manually written every line of its implementation.
The implemented architecture
The inventory domain is plain C#, separate from Unity presentation and adapters. Persistent items use stable IDs; stacks use a variant ID and quantity. A validated ownership graph records where an item belongs.
Invalid requests can be rejected without publishing a partial mutation. Revision checks reject stale work, successful changes commit before event publication, and typed result categories make rejection reasons part of the contract.
The accepted slices cover identity and registry, ownership, publication and revisions, storage, resource requirements, reservations, and equipment. IC06 is reservations. Its ledger tracks exact input claims and predicted output guarantees, supports cancellation and fulfillment, and fails closed when reconstruction cannot establish valid state.
Two corrections the existing tests missed
Public result categories, 28 August 2026. An early IC07 review found that an equipment conflict and invalid armor placement could both fall back to an “unavailable” result. Existing tests checked rejection, without requiring the correct reason. The correction added precise assertions; a historical targeted run passed all three cases, including a genuine unavailable control.
Coordinator authorization, 30 August 2026. A later, separate review found raw transfers across ownership roots involving equipment that should have been rejected even with current revisions. Fourteen invalid plans were accepted and one was rejected for the wrong reason: the targeted red run passed zero of 15 cases. After correction, all 15 passed.
That second requirement also covered ordering and preservation: reject before concurrency checks, leave authority unchanged, and publish no events. It was a different defect from the earlier result-category problem.
The practical lesson was to verify the boundary and the exact public result, rather than treating compilation or a generic rejection assertion as enough.
Evidence and results
The accepted IC01–07 baseline was recorded on 31 August 2026. Its saved final independent-review runs report:
| Historical test scope | Passing cases |
|---|---|
| Focused IC07 checks | 119 / 119 |
| Inventory Core suite | 487 / 487 |
| Project EditMode suite | 671 / 671 |
These scopes are nested and must not be added together. A read-only inspection on 6 October 2026 checked source and stored evidence; it did not launch Unity or rerun the tests. The results establish their historical execution scope, not a fresh runtime result today.
The evidence records a process that rejected apparently successful work, identified a specific failure and checked a correction before acceptance. Independent review, automated validation and my acceptance remain distinct steps.
Limits and next direction
IC08 remains design and review work, without implementation authorization. The accepted core excludes UI, visible equipment and statistics, combat, potion consumption, world-loot cutover and physical save orchestration.
Broader work includes deterministic world queries with seed/settings fixed for a service lifetime; terrain ownership, rendering and collision readiness boundaries; and tree generation, chopping, felling and snapshots. These implemented subsystems do not establish complete player-action, loot, terrain takeover or Meadow integration.
The workflow also costs time: requirements, review evidence, rework and repeated validation all add overhead. I scale that rigor with risk. Persistent state and cross-system authority receive full gates; isolated gameplay and disposable visual work can use lighter processes.
There is no measured productivity improvement or commercial-release claim here. A fresh commit-linked test run, a recorded behavior walkthrough and approved public evidence excerpts would strengthen external verification. The private review packet stays separate from this page.