Skip to content

Authoring Guide ​

How to write a kitsoki application. The companion documents:

  • state-machine.md — the conceptual model.
  • kitsoki docs app-schema — the authoritative YAML reference.
  • testing.md — flow and intent fixtures.
  • hosts.md — every built-in host handler.

If you only read one source other than the schema, read that.


1. Anatomy of an app ​

testdata/apps/cloak/
├── app.yaml             single source of truth (or top of an include tree)
├── recording.yaml       (state, input) → (intent, slots) for replay/static harnesses
├── flows/
│   ├── winning.yaml     Mode 2 deterministic flow tests
│   └── losing.yaml
├── intents/
│   └── go_intents.yaml  Mode 1 intent-recognition fixtures (optional)
└── prompts/             optional — Markdown prompt templates for host.agent.ask
    └── shell_repair.md

Kitsoki reads only the YAMLs. Markdown prompts under prompts/ are referenced by host.agent.ask via relative path. Anything else is your test/documentation infrastructure.


2. Minimal runnable app ​

app:
  id: tiny
  version: 0.1.0
  title: "Tiny App"

world:
  counter: { type: int, default: 0 }

intents:
  increment:
    description: "Add one to the counter."
    examples: ["add one", "++", "bump"]
  show:
    description: "Show the counter."
    examples: ["show", "what's the count?"]

root: main

states:
  main:
    view: |
      counter = {{ world.counter }}
    on:
      increment:
        - target: main
          effects:
            - increment: { counter: 1 }
      show:
        - target: main

Save as tiny.yaml, then:

kitsoki run tiny.yaml                    # opens TUI; auto-picks harness
kitsoki test flows tiny.yaml             # no fixtures yet → exits 0
kitsoki viz tiny.yaml                    # writes tiny-viz.dot

3. Authoring loop ​

The loop most authors settle into:

  1. Sketch the graph in YAML. Start with rooms and intents. Use placeholder views.
  2. kitsoki inspect or kitsoki turn to probe. kitsoki turn is especially useful — one stateless turn, no DB, JSON output.
  3. Add a flow fixture under flows/. Use intent: blocks (no LLM needed) to lock the state logic.
  4. kitsoki test flows until green.
  5. Add intent fixtures for natural-language inputs. Run kitsoki test intents --harness static to lock pass-rates.
  6. kitsoki viz to sanity-check the graph shape.
  7. kitsoki render -o APP.md to produce review-friendly docs.
  8. kitsoki run to play it for real.

The first four steps catch the vast majority of mistakes before the LLM ever sees the app.

3.1 Avoid ceremony steps ​

A room costs the operator a turn, so it must earn it. A room exists to do exactly one of:

  • Decide — offer a choice between two or more reachable forward paths (a decision gate), or
  • Side-effect — invoke: a host or mutate world: in its on_enter:, or
  • Collect — gather slot input the next step needs.

A room that does none of these is ceremony — it makes the operator click begin/continue/ok/next to advance the one place they could go anyway. The two usual shapes, and the fix:

  • Pass-through landing room. An idle/start/welcome whose only job is a slotless, guardless intent that routes onward. → Make the first real room the root: and delete the landing room. (cherny-loop did this: root: configuring, no idle+begin.)
  • No-choice checkpoint. A room with a single forward intent and no alternative. → If its on_enter: does real work, auto-advance with emit_intent: from on_enter: (see state-machine.md → the emit_intent effect) instead of waiting for a click; if it does no work, fold it into the neighbouring room.

The engine already draws this line: a state is a decision gate only when it has ≥1 forward intent reachable only by operator input (not an emit_intent: auto-advance target) — isDecisionGate in internal/machine/machine.go. One forward path is not a decision; don't spend a turn on it.

Genuine multi-choice checkpoints are not ceremony — an accept / refine / abandon review gate, a router hub with many actions, or a deliberate review-before-a-costly-action pause all earn their turn. The standard look re-render and quit/abort escape hatches are exempt.


4. Top-level fields (cheat sheet) ​

The full reference is kitsoki docs app-schema. The high-altitude shape:

app:        { id, version, title, author, license }
world:      { <name>: VarDef }                   # typed key/value bag
intents:    { <name>: Intent }                   # global intent library
root:       <state-name> | <inline-State>        # initial state
states:     { <name>: State }                    # the graph
off_path:   { trigger, banner, return }          # optional escape hatch
hosts:      [ <host-name>, … ]                   # allow-list
proposals:  { <name>: ProposalKind }             # draft → review → execute
phase_templates: { <name>: PhaseTemplate }       # repeated-room compression
phases:     { template, graph }                  # template instantiations
checkpoint_intents: { <name>: { description } }  # injected into _awaiting_reply states
include:    [ <glob>, … ]                        # merge other YAMLs

What lives where:

  • Reusable intents → top-level intents:. State on: maps reference them by name.
  • State-specific intent overrides → states.X.intents:.
  • Anything that touches the network or filesystem → host.* invocation, declared in the top-level hosts: allow-list.
  • Anything that survives a turn → world: with a typed default.
  • Anything per-turn that the user supplied → slots: on the intent.

5. Common patterns ​

5.1 Catch-all transitions ​

Always end an intent's transition list with a default: true so the machine never has to emit GUARD_FAILED for a benign case:

on:
  go:
    - when: "slots.direction == 'south'"
      target: bar
    - when: "slots.direction == 'west'"
      target: cloakroom
    - default: true
      target: foyer
      effects:
        - say: "You can't go that way."

5.2 Wildcard intent handler ​

Inside a state where any non-listed intent should behave the same way, bind "*":

on:
  "*":
    - target: .                # stay
      effects:
        - increment: { disturbance: 1 }
        - say: "It's too dark to do that."

target: . means "stay in the same atomic state".

5.3 Calling a host ​

hosts:
  - host.run

states:
  shell_idle:
    on:
      run:
        - target: shell_done
          effects:
            - invoke: host.run
              with:
                cmd: "git status"
                cwd: "{{ world.workspace_root }}"
              bind:
                last_output: stdout
                last_code:   exit_code
              on_error: shell_error
            - say: |

            {{ world.last_output }}

5.4 Deterministic Starlark glue ​

For ad hoc procedural logic, prefer host.starlark.run before shell. It gives you a real language with a typed sidecar, a narrow capability sandbox, deterministic replay, and flow-test coverage that executes the real script.

Use it for:

  • shaping JSON/YAML into world fields;

  • branching on structured HTTP responses;

  • inspecting a small part of the working tree;

  • writing bounded review artifacts;

  • replacing dense Pongo or inline shell data munging.

    hosts: - host.starlark.run

    states: enriching: on_enter: - invoke: host.starlark.run id: derive_widget with: script: scripts/derive_widget.star capabilities: http: methods: [GET] cassette_required: true inputs: widget_id: "{{ world.widget_id }}" bind: widget_name: name on_error: enrich_failed once: true

The script must define def main(ctx): ... and return a dict. The same-path sidecar, not the script comments, declares the enforced contract:

# scripts/derive_widget.star.yaml
inputs:
  widget_id: { type: string, required: true }
outputs:
  name: { type: string }

The default Starlark surface is intentionally tiny: ctx.inputs, read-only ctx.world.get(key), and deterministic json, math, and decode-only yaml. External surfaces are opt-in with with.capabilities: grant http, fs, vcs/probe/github, or exact host.verbs only when the script needs them. Ungranting is real enforcement: an ungranted ctx.fs or ctx.http attribute is absent before any filesystem/network adapter can run.

Flow fixtures should run the real script. If the script uses HTTP, add starlark_http_cassette: so the test replays the network boundary without a socket. If it uses ctx.fs or ctx.probe, use the inspection replay support (starlark_inspect_cassette:) instead of replacing the whole handler with a canned result. A whole-handler stub proves only the room wiring, not the script logic. Add at least one negative capability check for any new capability family: the test should still fail when a harmless mock/replay adapter is present but the script tries an ungranted ctx surface.

Full reference: ../architecture/starlark.md and ../architecture/hosts.md#hoststarlarkrun.

5.5 Bounded CodeAct loops ​

Use host.agent.codeact when the step is still exploratory, but the agent should act only by emitting Starlark snippets over an explicit capability map. This is narrower than host.agent.task: no open tool loop, no sandbox: block, and no shell unless a granted Starlark host/probe surface provides a specific read-only operation.

hosts:
  - host.agent.codeact

agents:
  triager:
    system_prompt: "Inspect with scoped Starlark snippets and submit a verdict."

states:
  triaging:
    on_enter:
      - invoke: host.agent.codeact
        once: true
        with:
          agent: triager
          goal: "Triage {{ world.ticket_id }} and return a verdict with evidence."
          budget: 6
          capabilities:
            world: read
            vcs: read
          schema: schemas/triage_verdict.json
        bind:
          triage: payload
          codeact_steps: steps
        on_error: triage_failed

The runtime runs each emitted snippet through the same Starlark evaluator as host.starlark.run. If the agent tries ctx.fs without an fs grant, the step fails early and the error is fed back to the next step; it does not fall through to a mock or touch the live filesystem. That is the point of CodeAct: exploratory agent reasoning with deterministic, reviewable actions.

Flow tests usually replay or stub the whole host.agent.codeact result with a host_cassette, because the live loop would spend LLM tokens. Once the loop stabilizes into deterministic logic, promote the final snippet to host.starlark.run with a sidecar and prove the promoted flow has no host.agent.codeact dispatch.

Full reference: ../architecture/starlark.md and ../architecture/hosts.md#hostagentcodeact.

5.6 LLM-backed effect ​

host.agent.ask runs claude -p against a prompt template file with templated {{ args.X }} placeholders; bind stdout back into world. Full contract and the ask_with_mcp / talk variants in hosts.md. End-to-end worked example (shell-repair) in kitsoki docs llm-guide §11.1 LLM-backed effects.

Prompt files are first-class, extensible templates: they can {% extends %} / {% include %} other prompts, and you mark the sections a project is expected to specialize with {% block spec_&lt;name&gt; %} extension points. A project then drops an overlay that extends your base prompt and fills those blocks — specializing your story without forking it. This is the supported way to inject project-specific context (repo conventions, domain rubric, house tone); see prompts.md for the search path, the spec_ convention, --prompt-overlay, and kitsoki prompts spec.

5.7 Operator clarification gates ​

When a story needs fresh human input mid-flow, prefer the story-visible operator-aware path whenever the runtime provides one. The ask belongs in the state machine as an invoke:/transition that can be rendered, traced, stubbed, and replayed; it should not be hidden inside an agent prompt as AskUserQuestion or a live-only branch.

The standard shape is:

  • The agent or deterministic step produces the questions as data and binds them into world.
  • A story effect invokes the operator-aware ask handler (for example host.operator.ask, once available) from the same room, guarded on there being unanswered questions.
  • If a live operator surface is attached, the handler forwards through the shared OperatorPrompter seam. That means web, TUI, and MCP Studio use the same runtime path; Studio can route the prompt through MCP elicitation or its session.answer fallback.
  • If no operator is attached (flow tests, cassettes, headless replay), the handler returns a stable "not answered" result instead of blocking. The room then falls back to its ordinary typed-answer UI, skip/regenerate verbs, or a needs-human exit.
  • Flow fixtures stub the ask handler by invoke id: for both outcomes: answered-by-operator and no-operator/fallback. Cassettes record the same host call shape, not a separate prompt-only behavior.

This keeps "real" and headless behavior close enough to test. The surface may change how the operator answers, but the story graph, world binds, trace events, and fallback room stay the same.

Do not:

  • Re-enable or depend on AskUserQuestion inside dispatched agents. It is headless-unsafe and hard-denied.
  • Add an agent-only MCP instruction that changes the story outcome when the tool happens to be present. If MCP-aware asking is useful, expose it as a story host call/effect so flows and cassettes can exercise it.
  • Replace an existing free-text clarification room with a live-only modal. The modal is an acceleration path; the room remains the durable fallback and review surface.

5.8 Background job ​

hosts:
  - host.run

world:
  result:      { type: string, default: "" }
  last_job_id: { type: string, default: "" }

states:
  running:
    on_enter:
      - invoke: host.run
        with:
          cmd: "long-running-script.sh"
        background: true
        bind: { last_job_id: job_id }
        on_complete:
          - set: { result: "{{ world.last_job_result.stdout }}" }
          - say: "Job complete. Output: {{ world.result }}"

When the job finishes, the orchestrator fires the on_complete: effects in a synthetic turn and posts an inbox notification. Full lifecycle in background-jobs/.

5.9 Posting to a transport ​

hosts:
  - host.transport.post

effects:
  - invoke: host.transport.post
    with:
      transport: "{{ world.transport }}"   # "tui" / "jira" / "bitbucket"
      thread:    "{{ world.thread }}"      # PLTFRM-12345 / PR/repo/42 / session-uuid
      phase_id:  "phase_a"
      title:     "Phase A complete"
      body:      "Result: {{ world.result }}"

The transport handles markup conversion (Markdown → Jira wiki for Jira, etc.). See transports.md for the registry.

5.10 Template interpolation: how complex values render ​

{{ ... }} expressions inside YAML strings are evaluated by the expr-lang engine against world and slots. How the result is spliced back into the surrounding string depends on its type:

Value typeRendering
string, int, float, boolGo's fmt %v (the usual default).
map[string]any, []anyencoding/json.Marshal — sorted keys for maps, compact form.
nilempty string.
anything else%v (fallback).

The map/slice case matters when you pass a structured world slot into a host call argument or a prompt:

- invoke: host.run
  with:
    cmd: "consume.py"
    args: ["--input", "pr_status={{ world.pr_status }}"]

If world.pr_status is {state: "FAILED", build: "..."}, the rendered arg is pr_status={"build":"...","state":"FAILED"} — parseable JSON, sorted keys, ready for the downstream CLI. Without the JSON rule it would render as Go's map[build:... state:FAILED] repr, which no standard parser can read.

On marshal failure (cyclic graph, unsupported type) the renderer falls back to %v so a corrupt slot doesn't crash the template. Implemented in internal/expr/expr.go::anyToString.


6. Scaling up: includes, phases, proposals, imports ​

For non-trivial apps, four features compress the YAML:

  • Includes — include: ["rooms/*.yaml"] merges other YAMLs into the manifest. Duplicate state or intent names error at load. Use for same-app file splitting.
  • Imports — imports: { &lt;alias&gt;: { source: ./bugfix } } embeds another app as an aliased sub-story. Private world; explicit projections through world_in: and per-exit set:; named exits; state/intent/prompt overrides; rebindable host_interfaces:. Use for cross-repo composition and reusable mini-stories. Full reference: imports.md.
  • Phase templates — declare a reusable room shape once, instantiate it once per phase. See state-machine.md §9.
  • Proposals — declare a draft → review → execute lifecycle once, reuse it for every "user-confirms-then-runs" pattern. Schema in app-schema.md under ProposalKind.

Use them when you'd otherwise be copy-pasting the same five states. Don't use them on a five-state app.

6.1 Synonyms — bare strings, templates, and enum-value tiers ​

The full story (latency budget, calibration workflow, cache behaviour) lives in semantic-routing.md. This section is the authoring cheat-sheet.

Intents and enum slots accept a synonyms: block. Each intent synonyms: entry is either a bare phrase or a {slot_name} template:

intents:
  ford:
    title: "Ford the river"
    examples: ["ford", "ford the river"]
    synonyms:
      - wade
      - "walk it"

  propose_purchase:
    title: "Draft a purchase"
    slots:
      items:      { type: string, required: true }
      total_cost: { type: int,    required: true }
    synonyms:
      # Bag-style bare strings — match anywhere in the input.
      - "buy supplies"
      # Templates — positional, with {slot_name} captures fed to the
      # slot's typed parser. Multiple templates supply alternatives.
      - "buy {items} for {total_cost}"
      - "purchase {items}"
      - "spend {total_cost} on {items}"

  pick_profession:
    slots:
      profession:
        type: enum
        values: [banker, carpenter, farmer]
        synonyms:
          banker:    ["banker", "rich guy"]
          carpenter: ["carpenter", "builder"]
          farmer:    ["farmer", "farmhand"]

An optional top-level app.routing: block tunes the matcher:

app:
  id: my-app
  routing:
    enabled: true            # set false to skip semroute entirely
    semantic_high_bar: 0.80  # confidence floor for direct submit
    semantic_mid_bar:  0.65  # confidence floor for slot-fill (Phase 4)
    stopwords_extra: ["yall", "wagon"]

At load time internal/semroute compiles every declared synonym — plus every examples: entry, which it treats as an implicit synonym — into a per-app index. Each foreground turn runs TryDeterministic → TrySemantic → tryTurnCache → harness, so "let's wade across" matches the wade synonym above and resolves ford without an LLM call.

Confidence bands decide what the orchestrator does with a hit:

BandWhat triggered itOrchestrator action
1.00Display string or unique example (deterministic tier)SubmitDirect, no LLM
0.90Bare-string synonym matched exactly one intentSubmitDirect, no LLM
0.80Template matched and every {slot} capture parsed cleanlySubmitDirect with the parsed slots, no LLM
0.65Template matched but ≥1 capture was named-but-unparseableComputeClarification for the unparseable slot
0.50Two or more intents tiedAMBIGUOUS_INTENT disambiguation card
0Nothing matched, or every match was on a non-allowed intentFall through to the turn cache, then the LLM

A worked template trace: with the three propose_purchase templates declared above, an input of "buy 6 oxen and 200 lbs food for 240" matches "buy {items} for {total_cost}". The {items} capture (6 oxen and 200 lbs food) feeds the string slot (no parser specialisation — the raw text becomes the value). The {total_cost} capture (240) feeds the int parser. Both parse OK, so the verdict is band 0.80 and the orchestrator submits propose_purchase{items, total_cost=240} directly.

The 0.65 band is conservative on purpose. If the user typed "buy 6 oxen for fjord" the literal anchors still align (the {total_cost} capture lands on fjord), but the int parser refuses the captured tokens. The verdict downgrades to 0.65, the orchestrator runs ComputeClarification targeting total_cost, and the TUI prompts the user for the cost without throwing away the items they already typed.

Template authoring rules:

  • Every {slot_name} must reference a real slot: on the same intent. The compiler refuses an unknown name with a clear *CompileError.
  • Captures must be separated by literal tokens. "buy {items}{total_cost}" is a compile error because the matcher cannot split the run between the two captures.
  • Leading and trailing captures are fine ("{x} dollars", "spend money on {items}", plain "{x}").
  • Templates are positional — they match input in the order the author wrote them. Bare-string synonyms remain bag-style; pick the shape that fits the phrase.
  • Multiple templates per intent are encouraged. Within an intent, the matcher prefers the template that fills the most slots (most-specific-wins); declaration order breaks fill-count ties.

Enum-slot synonyms: are consumed by internal/slotparse. The slot parser runs three tiers in order — direct stem match, synonym word-bag containment, and Damerau-Levenshtein-1 fuzzy match — so the same pick_profession fixture above will route "banker", "rich guy", "money man", and even the typo "bankr" all onto the canonical banker value without an LLM call. Adding a new synonym is one line of YAML; the matcher rebuilds its per-slot index at app load.

The turn-result cache reads routing.cache_* to size and expire its rows. Calibrate the synonym library with kitsoki replay-routing &lt;app.yaml&gt; --target 0.30 and grow it with kitsoki inspect --synonym-suggestions --cache-db &lt;path&gt; — see semantic-routing.md for the full workflow.


6.2 host.agent.extract — tiered resolver for effects ​

host.agent.extract solves a different problem than transport-level routing: it resolves any free-text field inside an effect into a typed payload. Use it when a state transition needs to extract a structured value from something the user typed, from a file, or from a background tool output.

effects:
  - invoke: host.agent.extract
    with:
      input: "{{ world.user_input }}"
      schema: ./schemas/direction.json
      resolvers:
        - synonyms: ./synonyms/directions.yaml
        - slot_template: ./templates/directions.yaml
        - llm:
            prompt: ./prompts/extract_direction.md
            agent: extractor
    bind:
      submitted:   world.extracted_direction
      resolved_by: world.extract_tier
    on_error:
      - invoke: host.transport.post
        with:
          transport: tui
          body: "Could not extract a direction from your input."

Synonyms file format (synonyms/directions.yaml):

"go north,head north,north": { direction: "north" }
"go south,head south,south": { direction: "south" }
wade: { action: "wade" }

Keys are case-insensitive, comma-separated phrase lists. Values are the typed payload. Keep the file next to the app's other YAML; the path in resolvers: is relative to the file that contains the effect.

Result shape:

FieldNotes
submittedThe typed payload. null on no-match.
resolved_bysynonyms | slot_template | llm | no_match
claude_session_idClaude session ID when the LLM tier matched.

On no_match, Result.Error is set so on_error: fires. Use this to show the user a helpful fallback.

Progressive determinism — after any LLM-tier resolution, run:

kitsoki extract suggest-synonym <session-id> <call-id>

The command prints a YAML snippet with the exact phrase→payload mapping that will move the next identical input to the deterministic tier. Add it to your synonyms file to shrink the LLM dependency.


7. Authoring tooling ​

CommandWhat it does
kitsoki inspectRead-only JSON snapshot of a stored session.
kitsoki turnOne stateless turn. Great for probing (--state X --intent Y --world …).
kitsoki vizDOT or Mermaid graph of the state machine.
kitsoki render -o APP.mdMarkdown documentation derived from the YAML — overview, Mermaid diagram, transition tables.
kitsoki test flowsMode 2 deterministic tests.
kitsoki test intentsMode 1 intent pass-rate tests.
kitsoki run --warp &lt;path&gt;Boot the TUI directly into a primed mid-game state from a YAML "warp basis". See imports.md.
In-TUI /warpSlash command equivalent. /warp &lt;state&gt; world.X=Y for inline; /warp file:&lt;path&gt; to load a basis.
kitsoki docs apply-proposalLLM-facing guide for "implement this prose proposal against app.yaml".
kitsoki extract suggest-synonym &lt;session-id&gt; &lt;call-id&gt;Propose a synonym entry from a recorded LLM-tier host.agent.extract call.
In-TUI Edit modeHot-reload editing — see developer-guide.md §8.

kitsoki render is one-way: the Markdown never feeds back into the engine. Re-run after every change to keep APP.md in sync.


8. Pitfalls ​

  • Forgot to declare a host. The loader rejects invoke: host.X unless hosts: [host.X] is at the top level.
  • Default branch missing. When no when: matches, the user gets GUARD_FAILED. Always provide a default: true arm.
  • relevant_world: [foo] for an undeclared world key. The loader rejects this — foo must exist in world:.
  • State name collisions across includes. The loader merges conservatively; rename one of them.
  • Background job referencing world.last_job_result outside on_complete:. That variable is injected only into the completion turn; outside of it the value is empty.
  • Forgetting default: on the last transition for an intent inside a phase template. The expander adds one for cycle_budgets, but for non-budgeted arcs it's your responsibility.
  • Editing app.yaml while a session runs without saving. Hot reload triggers on the watched file's mtime; if your editor writes through a temp file (vim by default), enable backup-style writes or run kitsoki run --no-watch (when implemented) and reload by hand.

9. Choosing tool profiles for agents ​

When an agent declares Bash in its tools: list and is used with host.agent.ask or host.agent.decide, you must also supply a bash_profile:. Pick the profile that gives the LLM exactly the capability it needs — no more.

read-only — the LLM can only run commands on a built-in allowlist: grep, find, cat, head, tail, ls, git, jq, rg, wc, stat, awk, sed, sort, uniq, echo, and a handful of others. Use this for diagnosis / code-review agents that only need to inspect the repository. The loader enforces the allowlist; multi-command chains (;, |, &&, backticks) are always rejected regardless of profile.

agents:
  code-reviewer:
    system_prompt_path: prompts/review.md
    tools: [Read, Grep, Glob, Bash]
    bash_profile: read-only

commands: [...] — an explicit argv0 allowlist you maintain. Prefer this when the agent needs a tool not on the read-only list but you still want the guarantee "this agent cannot run rm, curl, or arbitrary binaries." Useful for CI-diagnoser patterns that need kubectl or docker inspect but nothing else.

agents:
  ci-diagnoser:
    system_prompt_path: prompts/diagnose_ci.md
    tools: [Read, Bash]
    bash_profile:
      commands: [git, jq, grep, kubectl]

sandboxed_write: &lt;dir&gt; — the LLM may write, but only under a per-call scratch directory. Network is denied via the HTTP_PROXY env var trick (best-effort: raw TCP connections are not blocked). Use this for "build the project and inspect the output" patterns where the agent needs to produce temp files without touching the working tree.

agents:
  build-inspector:
    system_prompt_path: prompts/build_inspect.md
    tools: [Read, Bash]
    bash_profile:
      sandboxed_write: ""   # empty → system TempDir; or supply a base path

For host.agent.task and host.agent.converse, bash_profile is not consulted — those verbs allow unrestricted Bash by design. Use the effect-level sandbox: block when the subprocess itself needs runtime supervision:

- invoke: host.agent.task
  with:
    agent: implementer
    sandbox:
      min_strength: supervised
      repo: read_only
      rw: [".artifacts/my-run", ".worktrees"]
      hidden: [".env", ".git/config"]
      degrade: warn
    acceptance: { schema: schemas/result.json }

The shippable supervised backend gives process-group cleanup, timeout/cancel, temporary HOME/XDG dirs, provider/Kitsoki env allowlisting, and trace-visible policy. It records filesystem/network policy as degraded when it cannot enforce it; stronger fs_confined/os_confined/vm_confined backends are future runtime implementations.


10. Where to next ​

  • The schema — kitsoki docs app-schema.
  • Worked examples — testdata/apps/cloak, testdata/apps/dev-story, testdata/apps/proposal_smoke, testdata/apps/background_jobs.
  • Embedded operator manual — kitsoki docs llm-guide.
  • The state machine in depth — state-machine.md.
  • Agent declaration reference — hosts.md §Agent declaration.