Harness Profiles
Status: operator-facing reference for the
harness_profiles:block and the live provider/model switch. A harness profile is a named, operator- selectable bundle of the agent-selection axes — which backend CLI is forked, which endpoint it talks to, and which model it defaults to — that a live session can switch between from the TUI (/provider,/model) or the web header picker. The switch takes effect on the next turn.
A kitsoki session has four orthogonal agent-selection axes, each documented on its own page:
- backend — which coding-agent CLI is forked (agent-backends.md):
claude | copilot | codex. - provider — an env retarget of the forked CLI's subprocess (agent-providers.md).
- plugin — an alternate component that answers (agent-plugin.md), e.g.
builtin.local_llm(llama.cpp). - model — the
--modelpassed to the call.
Historically each axis was frozen at startup and reachable only through flags, env, or per-story YAML — never from a live session. A harness profile collapses these axes behind one operator-facing name so an operator picks a profile (and optionally a model) instead of learning the taxonomy.
Configuration
Profiles are declared in .kitsoki.yaml (and its local override — below; the same file that carries story_dirs and the implicit-root root: block — see imports.md "The blank root that grows"), loaded on both kitsoki run (TUI) and kitsoki web.
Shared file and local override
Config loads from two layers — the same dichotomy as Claude Code's settings.json + settings.local.json:
| File | Tracked? | Holds |
|---|---|---|
.kitsoki.yaml | checked in | the team baseline — profiles that load for everyone with zero setup (ambient claude-native, no ${VAR}). |
.kitsoki.local.yaml | gitignored | your personal, secret-bearing, machine-specific overrides — synthetic/codex/local-llm profiles whose ${VAR} env would otherwise hard-fail the shared file for a teammate without the key. Copy .kitsoki.local.yaml.example. |
Load (internal/webconfig; LocalConfigPath derives the sibling path by inserting .local before the extension) deep-merges the local file on top of the shared one, local wins, before validation — so ${VAR} expansion and the backend/model/effort/default_profile checks all run once on the effective config:
story_dirs,default_profile— a non-empty local value replaces the baseline's;default_profilemay legally name a profile only the local file declares.harness_profiles— merge by name: baseline-only profiles survive, local profiles are added, and a profile declared in both is replaced whole by the local one (restate every field you want — you never restate the other profiles).agent_launch_policy— replaced whole by the local file. Keep machine paths such as protected checkout roots and allowed capsule directories in.kitsoki.local.yaml.agent_user_delegation— replaced whole by the local file. It records the local macOSrun_as_userwrapper setup, but runtime delegation is currently disabled: the block is parsed for compatibility, does not affect launch binaries or warnings, and does not select a backend/model.
A missing local file contributes nothing, so a fresh checkout runs off the shared baseline alone.
Guided first-run setup
The dev-story landing room includes a no-LLM setup path for the local override:
provider
or:
setup local harness profile
It runs profile_setup_discover.py to inspect the checked-in and local config, installed backend binaries/KITSOKI_AGENT_*_BIN overrides, and credential sources by presence only. The review room shows the effective profiles, backend status, common env-var presence, and a redacted .kitsoki.local.yaml patch. Nothing is written until the operator confirms.
Apply writes only .kitsoki.local.yaml. It can set default_profile to an existing profile, add a codex/OpenAI-compatible profile that references an env var such as ${OPENAI_API_KEY}, or add a builtin.local_llm profile. It refuses raw key material and refuses to write the local override if git tracks it.
# .kitsoki.yaml — checked in; the baseline that must load for everyone.
default_profile: claude-native # the profile new sessions start on
harness_profiles:
claude-native: # your native Anthropic Claude Code subscription
backend: claude # (ambient auth; the default — no secrets)
model: opus
models: [opus, sonnet, haiku] # claude's short model names
effort: medium
efforts: [low, medium, high, xhigh, max] # claude supports --effort
# .kitsoki.local.yaml — gitignored; deep-merged on top, local wins.
# default_profile: synthetic-claude # (optional) start new sessions here
harness_profiles:
synthetic-claude: # claude-code pointed at synthetic.new
backend: claude # base URL omits /v1 (claude appends /v1/messages)
model: hf:zai-org/GLM-5.2
models: [hf:zai-org/GLM-5.2] # static fallback
models_endpoint: https://api.synthetic.new/openai/v1/models # the full always-on list
quota: # optional local provider guardrail
window: 1m
tokens_per_window: 120000
max_concurrent: 1
reserve_tokens: 30000
state_path: .artifacts/quota/provider-state.json
lease_timeout: 45m
env:
ANTHROPIC_BASE_URL: https://api.synthetic.new/anthropic
ANTHROPIC_AUTH_TOKEN: "${SYNTHETIC_API_KEY}"
synthetic-codex: # the codex CLI pointed at synthetic.new
backend: codex # OpenAI base URL includes /v1
model: hf:zai-org/GLM-5.2
models: [hf:zai-org/GLM-5.2]
models_endpoint: https://api.synthetic.new/openai/v1/models
env:
OPENAI_BASE_URL: https://api.synthetic.new/openai/v1
OPENAI_API_KEY: "${SYNTHETIC_API_KEY}"
codex-native: # codex's own config/auth
backend: codex
model: gpt-5-codex
models: [gpt-5-codex, gpt-5]
llama-local: # a local llama.cpp model (plugin path)
plugin: builtin.local_llm
model: h200/gpt-oss-120b
| Field | Meaning |
|---|---|
backend | claude | copilot | codex; empty ⇒ claude. Ignored when plugin is set. |
model | default --model for the profile; when the profile is active it supersedes story-local agent model defaults so the selected provider receives a compatible model id. |
models | static catalog the /model command and web dropdown list. When set, the model pick must be a member (of the full catalog, incl. fetched — below). |
models_endpoint | OpenAI/Anthropic /models URL (e.g. https://api.synthetic.new/openai/v1/models); its always-on model ids are fetched and merged into the catalog at selection time (auth from this profile's env), so the full live list is offered — not a hand-maintained subset. A fetch failure falls back to models. Cached per profile. |
effort | default reasoning effort (low|medium|high|xhigh|max), applied where the backend/model supports it (claude --effort). |
efforts | catalog the /effort command + web effort dropdown list — declare only on profiles whose backend/model supports effort (codex ignores --effort today). Empty ⇒ no effort control. |
env | env overrides merged onto the forked CLI subprocess. ${VAR}-expanded at load time (an unset var is a hard error, mirroring providers:). Never recorded in traces. |
quota | optional provider guardrail for live CLI-backed calls. It throttles before the provider sees the request, persists usage assumptions, and coordinates parallel kitsoki processes through a JSON state file. See Provider quota control. |
plugin | routes through an agent plugin (e.g. builtin.local_llm) instead of forking a backend CLI. |
default_profile (top-level) | the profile new sessions start on; must name a declared profile. Omitted ⇒ the flag-derived static default (today's --agent/--model). |
Secrets never live in the file: env values use ${VAR} interpolation against the process environment. With no harness_profiles: block the static flag/env path is preserved byte-for-byte.
Agent launch policy
agent_launch_policy: is the local preflight guard for external coding-agent launches. It is separate from harness profiles: profiles select a backend/model; launch policy decides whether that backend may start in a working directory.
Example local policy:
# .kitsoki.local.yaml
agent_launch_policy:
enabled: true
require_capsule: true
protected_roots: [.]
allowed_roots:
- ./.worktrees/capsules
- /tmp/kitsoki-capsules
protected_branches: [main, master, trunk, "integration/*", "staging/*"]
When enabled, kitsoki run, kitsoki web, kitsoki mcp, and kitsoki agent launch all reject launches in protected roots/branches and can require an opened capsule. See agent-launch-policy.md for the exact semantics and trace fields.
The macOS run_as_user wrapper path is currently disabled. Launch policy alone is not a filesystem sandbox: it blocks unsafe start locations before launch, but backend processes currently run as the invoking user after launch.
Live TUI and web sessions also install an automatic fallback ladder when no explicit harness_ladder: is declared. The default priority is:
claude-native(claudebackend,opus)codex-native(codexbackend,gpt-5.5)synthetic-claude(claudebackend,hf:zai-org/GLM-5.2)synthetic-codex(codexbackend,hf:zai-org/GLM-5.2)
synthetic-codex stays in the profile catalog for manual selection and diagnostics, but it is intentionally last in automatic fallback. A 429, quota, rate-limit, timeout, 5xx, or provider transport failure is logged, recorded in the trace, and backs off that provider/harness lane. Kitsoki skips any remaining effort or model-strength rungs on that lane and tries the next configured profile instead.
Provider quota control
quota: is an optional, provider-neutral control loop for profiles that have finite or opaque limits. Kitsoki estimates a call before launch, reserves capacity in a persisted state file, starts the agent only when capacity is available, then updates its assumptions from the usage the backend reports.
It works for any CLI-backed profile (claude, codex, copilot) because it sits at the shared agent subprocess boundary. API-key profiles can leave it off or set high limits. Subscription profiles with inspectable quotas can still use it as the local coordination layer; provider-specific quota polling can update the configured numbers later without changing stories.
harness_profiles:
synthetic-claude:
backend: claude
model: hf:zai-org/GLM-5.2
quota:
window: 1m # token bucket window
tokens_per_window: 120000 # estimated/learned tokens admitted per window
max_concurrent: 1 # simultaneous in-flight calls for this profile key
reserve_tokens: 30000 # minimum reservation for small prompts
state_path: .artifacts/quota/provider-state.json
lease_timeout: 45m # stale in-flight reservation reap timeout
The limiter key is profile | backend | model | endpoint, so changing models or endpoints gets separate accounting. Profiles that intentionally share the same API key should share state_path; separate keys or machines can use a different file.
The persisted state lives at .artifacts/quota/provider-state.json by default and is guarded by an advisory .lock file. That gives two useful properties:
- Restart learning: observed usage, backoff, and the rolling window survive a
kitsoki webrestart. - Cross-process coordination: two local kitsoki processes sharing the same
state_pathsee the same in-flight reservations and token window.
Reservations have a lease. If a process crashes after reserving capacity but before recording completion, the next reservation reaps the stale entry after lease_timeout instead of blocking forever. Set it longer than your largest expected agent call.
The control loop updates itself from evidence:
usage.total_tokens, when present, is authoritative.- Otherwise it sums normalized token fields such as
input_tokens,output_tokens, cache-read/cache-creation tokens, and reasoning tokens. - Future reservations use at least the observed average if it is larger than the prompt-size estimate.
- A response or error containing
429,quota,rate limit,rate_limit, ortoo many requestsrecords a backoff through the currentwindow.
Operational signals are emitted as structured logs:
provider.quota.throttle— a call is waiting, withreasonoftokens,concurrency, orbackoff.provider.quota.usage— a call finished and updated the persisted state.
The state file is operational data, not source. Keep it under .artifacts/ (or another gitignored runtime path) unless you are intentionally capturing it as debug evidence.
Why env works for codex/openai, not just claude
A provider/profile's env is merged onto every backend CLI's subprocess environment (internal/host/agent_runner.go, envWithProvider), not only the claude one. So a claude-backed profile sets ANTHROPIC_BASE_URL/ ANTHROPIC_AUTH_TOKEN (which the claude CLI reads) and a codex-backed profile sets OPENAI_BASE_URL/OPENAI_API_KEY (which the codex CLI reads). That is what makes "synthetic.new on codex" real with no engine change — the merge was already backend-agnostic; only the variable names are backend- specific.
Selecting a profile
| Surface | How |
|---|---|
| TUI | /provider lists the profiles (active one marked) and /provider <name|n> switches; /model lists the active profile's catalog and /model <id|n> switches the model; /effort does the same for the reasoning effort (where the profile declares efforts:). See docs/tui/README.md. |
| Web | A provider dropdown, a dependent model dropdown, and (where the profile declares efforts:) an effort dropdown in both the Observe and Drive headers; switching fires the runstatus.session.set_selection RPC. See docs/web/README.md. |
Both drive the same orchestrator API — Profiles(), Selection(), SetSelection(profile, model, effort) — exposed to the web via the optional HarnessController driver interface.
Resolution & precedence
The selection is held per session behind a mutex and resolved once per dispatch. Precedence (highest first):
per-effect
with: { provider }/ named story provider › active profile model/env › agent model default › flag-derived static default
So an operator-selected profile like synthetic-claude can replace story-pinned Claude model names with the profile's compatible hf: model id. A call that explicitly names a story provider: still selects that provider instead of the session profile (applyProvider in internal/host/agents.go).
Next-turn semantics
A switch rebuilds the selection lazily: every interpretive call from the switch on uses the new selection; the one already in flight finishes on the old one. There is no mid-flight cancellation. A single snapshot per dispatch means a concurrent switch can never tear one call.
Trace
Every agent.call.start already stamps the model; a session that selected a profile also stamps profile (AgentCalledPayload.Profile) and, when set, effort, so a transcript line reads agent.decide · profile=claude-native · model=opus · effort=high. The web trace renders these as a chip on each agent row, so the trace provenance matches the picker. Only the profile name, backend, model, and effort are recorded — never the env secrets.
Real-provider runbook
The three providers the feature targets, end to end:
Put your synthetic.new key in the environment the kitsoki process inherits — it is referenced as
${SYNTHETIC_API_KEY}in.kitsoki.yaml, never inlined:export SYNTHETIC_API_KEY=sk-... # in your shell / profile / .envrc(synthetic.new's Anthropic-compatible base is
https://api.synthetic.new/anthropic— claude-code appends/v1/messages— and the OpenAI-compatible base ishttps://api.synthetic.new/openai/v1. Use explicithf:model ids (for examplehf:zai-org/GLM-5.2) to avoid rejected alias names.)Native Anthropic Claude Code —
claude-native: ambient auth, nothing to set. Verified live: a realkitsoki turnroutes free text → intent with theclaudebackend (no cassette).Auth-precedence gotchas (verified end-to-end):
claudehonorsANTHROPIC_BASE_URL+ANTHROPIC_AUTH_TOKENeven with an active subscription, sosynthetic-claudeworks with env alone.- A model will misidentify itself ("I am Claude Sonnet 4" from GLM-5.1) — that is hallucination, not a routing bug. Confirm the route by the billed model / the server's
modelfield, never by the model's self-description. codexignoresOPENAI_BASE_URL/OPENAI_API_KEYwhen logged into a ChatGPT account (it rejects the non-OpenAI model). To runsynthetic-codexyou must declare a[model_providers.synthetic]block in~/.codex/config.tomland select it;codex-native(the subscription) needs none of this.
Launch with the profiles and pick live:
kitsoki web --stories-dir stories # or: kitsoki run <story>/app.yamlIn the web header, pick
synthetic-claude(and a model from its catalog), orsynthetic-codex, then submit a turn. The trace row showsprofile=…and the chosen model's real answer. In the TUI,/provider synthetic-codexthen drive.
Automated tests never exercise these — real-LLM verification is operator-driven and gated, per CLAUDE.md.
Backward compatibility
Fully additive, default-off. Existing flags/env/providers:/agent_plugins: keep working unchanged; a profile is a named bundle of the same overrides applied through the same merge points. With no profiles declared, /provider and /model report the single flag-derived default and the web picker hides.