Skip to content

Agent action transcripts ​

Every agent call's full tool-use stream — captured once, replayed byte-identically.

Every agent call — task, decide, ask, converse — runs a rich native execution stream of tool calls, reasoning, and MCP submissions. kitsoki captures that stream byte-verbatim to a per-call sidecar and renders it in the observer: typed rows with full tool I/O, a latency waterfall, running token/cost accrual, the decide guardrail arc, an honest cassette-vs-live diff, and a run-wide rollup.

Question

Can the runtime overrule a model before state advances?

Watch for

The reject, nudge, resubmit, accept arc inside one bounded agent call.

Why it matters

The model proposes; the host contract decides whether the result is valid.

Key beats

  1. Every run begins here, in the story library — each card is a deterministic story graph kitsoki runs the same way every time. We'll demonstrate agent-action transcripts on the bug-fix pipeline.

    Start at the story library — screenshot
  2. This story hands a ticket to an autofix agent that reads the repo, edits code, and runs the build, then a judge gates the result. Those agent calls run rich native execution streams — exactly what the transcript drawer captures and replays.

    The bug-fix pipeline — screenshot
  3. Click New session to create a fresh, independently-traced run of the pipeline.

    Spin up a run — screenshot
  4. The agent-action transcripts live in the read-only observer view, where every recorded call can be inspected and replayed. Switch to it to open this run.

  5. Every agent call — task, decide, ask, converse — runs a rich native execution stream: tool calls, reasoning, MCP submissions. kitsoki captures that stream byte-verbatim to a per-call sidecar, renders it here, and replays it deterministically from the cassette. We lead with the one move a coding agent can't make, then walk the full per-call stream.

Full recorded walkthrough (17 steps)
  1. Lead with the move no coding agent can make: the host overruling the model mid-call. In the decide call the model submits its verdict via mcp__validator__submit, typed as a GUARDRAIL row — accept/reject + confidence, not a generic tool call. Here the first submit is REJECTED on a schema violation: the runtime, not the model, is in charge.

  2. On rejection the host injects a coaching nudge as the next turn's stdin — invisible in a raw stream-json tee, because the input prompt is never echoed back. kitsoki writes a synthetic _kitsoki row so the full submit -> reject -> NUDGE -> re-submit -> accept round-trip, with iteration boundaries, reads as one sequence.

  3. Beyond the guardrail, every call's full stream is here too. Take the autofix task: select it in the trace and its detail pane gains an 'Agent actions (N)' affordance — N is the captured event count, an 18-step Read / Grep / thinking / Edit / Bash arc. Click it to open the drawer.

  4. The drawer renders the native stream as typed rows: assistant reasoning, and each tool call's full input and output — the file Read, the Grep pattern, the Edit diff, the Bash command and its stdout. Every row collapses on its header, so a long arc stays scannable.

  5. Each step is a typed row anchored on a header you can click to expand or collapse its detail. No 200-rune preview, no names-only rollup — the command that ran, the file that was edited, and the bytes that came back, exactly as captured.

  6. Toggle the drawer's waterfall mode to see per-step latency. The offsets come from a parallel .timings sidecar stamped at capture — never re-derived from replay timing, so the bars are byte-stable.

  7. Each bar's width is that step's duration. The slow Bash or the long model turn stands out at a glance — the same headline view a dedicated agent-observability tool gives you, over data we already hold.

  8. The header accrues input/output tokens and cost across the whole call, not just the terminal total — so 'this call retried twice before it finished' is visible in the running tally, not buried.

  9. Because the transcript replays byte-identically from the cassette, the drawer can diff a fresh live run against the recorded one and flag tool-path drift — a capability no external tool has, because none has deterministic replay. Under pure replay there is no live run to compare, and the control says so honestly.

  10. The Actions view mode rolls every transcript-bearing call across the whole run up under its turn and room — the kitsoki analog of a session replay. Click to open it.

  11. Each row is one agent call — its verb, call_id, and action count — and expands into its own full drawer. The whole run's agent activity in a single grouped view, the unit an operator actually reasons about.

  12. You've seen the complete agent-action-transcripts stack: per-call typed rows with full tool I/O, the latency waterfall, running cost accrual, the decide submit -> reject -> nudge -> accept guardrail arc, the honest cassette-vs-live diff, and the run-wide session rollup — all over data captured once and replayed deterministically. Hit '?' anytime to replay this tour.