Fix a real bug in slidey, end to end
kitsoki's bug-fix pipeline reproduces, fixes, tests and validates a real slidey bug — a grid-cards narration-timing drift.
The bugfix pipeline run against the real slidey repo (a Node/Vue narrated-deck renderer), fixing a frame-timing drift: grid `cards` scenes with more than six items had their narration desynced because the frame estimator over-counted each card beyond the table by ten frames. Every room is a deterministic state-machine beat: reproduce the drift with a failing unit test, propose the one-line root-cause fix, implement it in an isolated worktree, run slidey's real node --test suite, self-review the diff, validate end to end, and close out — ready to open a PR. The content is replayed verbatim from a real fix run; the render is no-LLM and deterministic.
Can a Kitsoki story model a realistic repo workflow?
The handoffs between rooms, operator decisions, host calls, and replayed results.
The workflow is replayed from fixtures instead of improvised for the site.
An interactive clip from the kitsoki story deck — replayed from a real run, no LLM. Open full-screen →
Key beats
kitsoki's bug-fix pipeline, pointed at the real slidey repo — a Node/Vue narrated-deck renderer. We'll fix a grid-cards narration-timing drift end to end: reproduce → propose → implement → test → review → validate → done. Every room is a deterministic state-machine step kitsoki runs the same way every time — this whole walk is replayed from a real fix run, no LLM in the loop now.
This card is the bug-fix pipeline, bound to an external checkout — the slidey Node repo — so the agent reads and edits real renderer code in an isolated git worktree, and CI runs the project's real node --test suite. The ticket is slidey-128. Let's run it.
Click New session to create a fresh, independently-traced run of the bug-fix pipeline against slidey issue 128.
We're on the drive view: each operator accept advances the pipeline exactly one room, and the state badge tracks where we are. The ticket — slidey-128, grid `cards` scenes with more than six items desync their narration — is already loaded. Let's reproduce it.
bf.reproducing — the agent confirms the bug is real before touching code. The frame estimator and the renderer must produce identical frame totals (a documented lock-step contract). For a grid cards scene the estimator sums `cards_item_<i> ?? 30` over every card, but the table only keys items 0–5; the renderer holds `?? 20`. A failing unit test on an 8-item grid shows the 20-frame overcount: estimate 484 vs render 464. Accept to move on.
Full recorded walkthrough (13 steps)
bf.proposing — root cause confirmed: the estimator's `?? 30` fallback disagrees with the renderer's `?? 20` for cards beyond cards_item_5. The fix is a one-line change in `src/timing.js` aligning the fallback to 20, plus a self-deriving regression test. Accept the proposal.
bf.implementing — the agent edits `src/timing.js` in an isolated worktree: the grid-cards per-item fallback changes from `?? 30` to `?? 20`, with a comment documenting the lock-step contract, plus the regression test in `test/timing.test.js`. The change is committed to the fix/cards-item-timing-drift branch. Accept to run the tests.
bf.testing — CI runs the project's real suite: `node --test test/timing.test.js`. The new regression proves the 8-item grid estimate now matches the renderer-true total, and the module is green — 19 passed, 0 failed. The fix is proven, not just plausible. Accept to self-review.
bf.reviewing — the agent reviews its own diff against the worktree before anyone else sees it: scope is one fallback value, in-table grids (≤6 cards) and every other scene type are untouched, and the regression derives its expectation from the renderer's own rule. Accept to validate end to end.
bf.validating — a full check: the whole slidey suite passes (94 passed, 0 failed, 3 pre-existing skips), the targeted timing suite is 19/19, and the originally-reported symptom is gone — a >6-item grid cards scene now estimates exactly the frames the renderer emits, so narration stays in sync. Outcome: pass. Accept to close out.
bf.done — the pipeline has reproduced, proposed, implemented, tested, reviewed and validated the fix. The close-out artifact captures the lessons and the metrics. The branch fix/cards-item-timing-drift is ready to open a pull request against slidey. One more accept resolves the ticket.
The fix is committed to fix/cards-item-timing-drift, the ticket transitions to resolved, and the run reaches its PR-ready close-out. The whole arc — a real slidey bug, reproduced to PR-ready fix — ran as one deterministic, replayable graph.
You watched kitsoki's bug-fix pipeline take a real slidey bug from a symptom to a tested, validated, PR-ready fix — every room a deterministic, replayable state-machine step, the content drawn from a genuine fix run. Hit '?' anytime to replay this tour.
The high-value slidey integration: kitsoki, pointed at a real Node deck-renderer repo. The agent reproduced the bug from the estimator/renderer lock-step contract, found the divergent fallback (estimator counted 30 frames where the renderer holds 20 for cards beyond cards_item_5), wrote the one-line fix plus a self-deriving regression test, and the targeted node --test suite went green (19/19; 94/94 whole-repo). What you see here is that real run, replayed deterministically — same frames every time, no API cost.