Trace introspection, live
View modes, waterfall, decision-first detail, annotation, replay.
Every trace-introspection capability on a live run: view modes, the latency waterfall, observation category chips, decision-first detail with the confidence bar and alternatives, annotation, and deterministic replay.
Can the observer make a complex run debuggable?
View modes, event kinds, latency, world mutation, annotation, and replay.
Every important runtime action is recorded as a structured event.
Key beats
Every run begins here. A story is a deterministic graph of rooms that kitsoki runs the same way every time, recording each decision and call as it goes. We'll demonstrate trace introspection on the bug-fix pipeline.

This story triages a ticket, lets an agent patch the code, then judges the result against a confidence gate — looping until it passes. It exercises every kind of decision, agent call, and host call, which makes it the ideal run to introspect.

Click New session to create a fresh, independently-traced run of the pipeline.

The trace tools live in the read-only observer view — built for inspecting a run while it's live or after it's finished. Switch to it to introspect this run.
kitsoki records every decision, agent call, host call, and world mutation as a structured, immutable event stream. This tour walks the full introspection stack — everything you can see, annotate, and replay from a finished run.
Full recorded walkthrough (18 steps)
Tree shows the event sequence, Timeline is a latency waterfall, Graph is the state diagram — three co-equal projections of the same immutable event stream. Switch anytime.
Click the Timeline tab to see a duration-proportional waterfall. Each bar's width is the call's duration_ms within its turn window.
Bars are colored by observation kind — purple for decisions, blue for agent calls, amber for host calls. A bottleneck stands out immediately instead of requiring manual millisecond arithmetic.
Switch back to Tree view to see each event in sequence with its kind badge.
Expand a turn's start event and read its routing detail. The 'Direct' row says yes — this turn was matched to its intent deterministically and advanced without ever calling an LLM. Across a typical run roughly four turns in five resolve this way, with zero agent calls. The model is a callee at named decision points, not the dispatcher for every step — the line a structured-output wrapper can't draw, because there the model is always in the loop.
Expand a world.update row and the change is laid out key by key — added, removed, or changed, with the before and after side by side. This is the state delta the runtime actually applied, not a sentence an agent wrote about what it did: kitsoki owns the world, so every mutation is structured, diffable data you can audit — never prose you have to trust.
Every event has a semantic kind: decision, agent-call, host-call, narration, world-mutation, routing, or lifecycle. Click a chip to filter down to one category — the same taxonomy drives row colors, the waterfall, and a future graph layout.
For a gate or routing event, the pane leads with the verdict — available intents, chosen intent, confidence, reason — not the raw prompt. The prompt is still there as a collapsed evidence drawer.
The bar fills to the model's confidence and marks the configured threshold with a tick. Green means the engine auto-fired; red means it bailed to human review.
Below the winning verdict, the pane lists every runner-up intent the model considered — with its own confidence score. Seeing the margin between winner and second-place reveals whether the decision was crisp or borderline.
Score, label, or comment on any gate or agent call. Annotations are stored in a sidecar stream — the deterministic trace is unchanged, but every scored decision becomes a labeled training datapoint.
Re-run one recorded agent call against a different LLM or local model and diff the verdict — without touching the original run. This makes the pluggable-operator moat visible: same recorded input, different operator, auditable output diff.
You've seen the complete trace-introspection stack: semantic kind taxonomy, latency waterfall, decision-first detail, confidence bars, runner-up alternatives, human annotation, and single-call replay. Hit '?' anytime to replay this tour.