Skip to content

Testing ​

  • agent-evals.md — story-local benchmarks for bounded host.agent.* call sites.
  • routing-tuning.md — free-text routing fixtures, live profile tuning, report summaries, and session-mined route cases.
  • story-upgrade.md — CI-safe capsule checks and gated live-LLM smoke runs for upgraded project stories.
  • model-harness-eval-pilot.md — pilot process for comparing model/harness behavior with intent reports and agent evals.
  • open-source-repo-catalog.md — landing page for reusable OSS repo candidates, implemented external bake-off manifests, harnesses, and durable results.