Skip to content

Punch-list story ​

Run YAML worklists through Kitsoki stories with model policy and independent verification.

The punch-list story consumes a punch-list/v1 manifest, enforces profile/model policy, delegates live story driving, gates results on deterministic verification, and writes a report. This demo drives the top-10 GPT-5.5 dogfood manifest through a no-LLM flow so the full loop is reproducible while the policy and evidence surfaces stay visible.

Question

Can a Kitsoki story model a realistic repo workflow?

Watch for

The handoffs between rooms, operator decisions, host calls, and replayed results.

Why it matters

The workflow is replayed from fixtures instead of improvised for the site.

Key beats

  1. Punch-list is a generic worklist runner. A YAML manifest names the stories to drive, the model policy to enforce, and the deterministic checks that decide whether each item passes.

    Start with the punch-list story — screenshot
  2. This demo opens the punch-list story itself. The fixture seeds the real top-10 GPT-5.5 dogfood manifest so the recording stays deterministic and free of live model calls.

    One story, many work items — screenshot
  3. Start a fresh session. The flow fixture supplies the manifest path and stubs the live driver, the same no-LLM posture used by story tests.

    Create the run — screenshot
  4. The first room explains the contract: load punch-list/v1 YAML, enforce model policy, drive each item, verify independently, and produce a report.

    Choose the manifest — screenshot
  5. The start action parses the manifest and rejects bad policy before any live work can start: duplicate IDs, missing story paths, Claude profiles, and LLM-spending verifiers are blocked.

    Load and lint — screenshot
Full recorded walkthrough (11 steps)
  1. The board shows processed, passed, partial, failed, skipped, and pending counts, then selects the next pending item by priority.

    Board view — screenshot
  2. Processing an item walks through policy, drive, verify, and record. The live driver handoff must include trace and model evidence before verification is allowed to mark the item complete.

  3. The same loop repeats item by item. At the midpoint, five top-10 entries have passed and five remain pending, with no live model calls in the recording.

  4. By the final queue beat, nine items are recorded as passed and the last item is selected: the story-qa workflow through dogfood-marathon.

  5. When no pending items remain, punch-list writes a markdown report under .artifacts/punch-list/ with all ten items, their status, story, trace path, and verifier summary.

  6. The shipped story is generic: swap in any punch-list/v1 manifest, run it through Studio MCP, and trust deterministic verification plus trace/model evidence rather than a maker self-report.