Skip to content

BugSwarm + GLM-5.2 bugfix comparison report ​

Generated at: 2026-07-06T00:00:00Z.

This report is generated offline from committed evidence. It does not call BugSwarm, Docker, or any LLM. Missing cells are reported as pending.

Research Question ​

Compare the Kitsoki bugfix pipeline with raw prompts on GLM-5.2, using total token usage as the primary cost axis, and prepare the same success-rate comparison for a BugSwarm corpus alongside the existing OSS oracle corpus.

The current committed evidence does not yet contain the full GLM-5.2 matrix. This report therefore separates observed results from missing cells instead of imputing raw-prompt or BugSwarm numbers.

Method ​

The headline matrix has one row per (corpus, treatment) bucket. A cell is counted as attempted only when its quality is solved, partial, or failed; pending and blocked are excluded from the model-quality denominator. Token totals are summed only from committed cell evidence that records real usage.

Inputs:

  • Report JSON: docs/case-studies/bugswarm-glm52-bugfix-report.data.json.
  • Report Markdown: docs/case-studies/bugswarm-glm52-bugfix-report.md.
  • GLM-5.2 bakeoff cells: tools/bugfix-bakeoff/results/cells.
  • Arena supporting rollup: tools/arena/results/round-1/rollup.json.
  • OSS oracle corpus: tools/arena/corpus/cost-bench.manifest.yaml.
  • Source catalog: tools/arena/corpus/sources.yaml.
  • OSS arena GLM rollup: not supplied.
  • BugSwarm source: tools/arena/corpus/bugswarm.seed.yaml.
  • BugSwarm verification report: not supplied.
  • BugSwarm arena rollup: not supplied.

Primary metrics:

  • success rate: solved / (solved + partial + failed).
  • partial rate: reported separately because hidden oracles can be implementation-coupled.
  • total tokens: provider-neutral primary cost measure.
  • USD cost: secondary; only shown where committed cell evidence provides it.

Corpus Coverage ​

corpustasksrepositoriesverified/imported status
OSS oracle corpus2612frozen and locally validated
BugSwarm1n/aadapter-ready; converted verified tasks: 0; verification report: 0/0 (none)

The OSS oracle corpus remains the active internal benchmark source. It covers the pre-registered public OSS targets plus existing hidden-oracle bugfix fixtures. BugSwarm is represented separately in the source catalog, so its fail/pass CI artifact sampling process does not get collapsed into the OSS oracle denominator.

BugSwarm source contract:

  • import explicit exported artifact metadata with tools/arena/scripts/bugswarm_to_arena.py.
  • require image_tag, repo, failed_job_id, and passed_job_id.
  • treat the failed job as RED and the passed job as GREEN inside the artifact image.
  • keep imported tasks unattempted until Docker verification proves both sides still reproduce.
  • verify with tools/arena/scripts/bugswarm_verify_source.py; dry-run mode records the Docker commands, while --execute runs each side in separate fresh containers.

Source Mix ​

The OSS oracle source and BugSwarm are kept as separate source families so the report can show blended overall treatment totals without hiding which evidence came from deterministic GitHub-content oracles, hidden bugfix fixtures, or containerized fail/pass CI artifacts.

source componenttasksreposoracle kindssplitrepositories
pre_registered_oss_targets2010github_contentheldout:4, training:16ansible/ansible, grafana/grafana, kubernetes/kubernetes, microsoft/TypeScript, microsoft/vscode, python/cpython, pytorch/pytorch, rust-lang/rust, tensorflow/tensorflow, vercel/next.js
armed_bugfix_fixtures62external_bakeoffheldout:2, training:4kitsoki, query-string
BugSwarm containerized_fail_pass_ci_artifacts11fail/pass artifact scriptsverification-gatedsquare/okio

Blend policy:

  • Keep OSS oracle tasks and BugSwarm artifacts as separate source families in denominators.
  • Report overall GLM-5.2 treatment totals only after both Kitsoki and raw-prompt arms have attempted cells.
  • Use total tokens as the primary cross-source cost axis; USD remains secondary and evidence-dependent.
  • Do not count dry-run BugSwarm verification as RED/GREEN proof.

Reproducibility Ledger ​

Status: reproducible. Generator: tools/arena/scripts/glm52_bugswarm_report.py sha256 08d7ac0776bf1edcca5c1860061be66745ed3b483f58fc37fb004bae9d494e43.

Regenerate with:

python3 tools/arena/scripts/glm52_bugswarm_report.py --generated-at 2026-07-06T00:00:00Z --json-out docs/case-studies/bugswarm-glm52-bugfix-report.data.json --markdown-out docs/case-studies/bugswarm-glm52-bugfix-report.md
artifactkindstatusbytessha256
tools/arena/corpus/cost-bench.manifest.yamlfilepresent20,768e9ac3482ba56eae8e837b32e5f106565fbc8a03915d9c2535e563f87f69ffb49
tools/arena/corpus/sources.yamlfilepresent3,00893125ed682ff0b829b75642fcefd6a4a8f2a2993f99122629dea51b70590e566
tools/arena/corpus/bugswarm.seed.yamlfilepresent1,185f174df93ace8ad08aa7a5f90033301c09719a7b7230617b6b2c6379c13821a2e
tools/arena/results/round-1/rollup.jsonfilepresent10,97849f8d3cb25601a32a51d44a44cc940bd24f0261af887cc31fdcbc95443745724
tools/bugfix-bakeoff/results/cells/*glm-5.2*.jsondirectory-glob1 match(es)n/an/a
tools/bugfix-bakeoff/results/cells/bug9-glm-5.2-kitsoki.jsonfilepresent1,554ae6b0d8529e1b3fc183a999e93788ffa852b9881dbc8e6a3248aa1fe89839046
optional-oss-arena-rollupmissing-optionalmissingn/an/a
optional-bugswarm-verificationmissing-optionalmissingn/an/a
optional-bugswarm-arena-rollupmissing-optionalmissingn/an/a

Validation commands:

  • python3 tools/arena/scripts/glm52_report_gate.py --report-json docs/case-studies/bugswarm-glm52-bugfix-report.data.json
  • python3 tools/arena/scripts/glm52_report_gate.py --report-json docs/case-studies/bugswarm-glm52-bugfix-report.data.json --require-publishable
  • python3 tools/arena/tests/test_glm52_bugswarm_report.py
  • python3 tools/arena/tests/test_glm52_report_gate.py
  • python3 -m py_compile tools/arena/scripts/glm52_bugswarm_report.py tools/arena/scripts/glm52_report_gate.py
  • python3 tools/arena/tests/validate_corpus.py tools/arena/corpus/cost-bench.manifest.yaml
  • python3 tools/arena/tests/run_no_llm.py

GLM-5.2 Headline Matrix ​

corpustreatmentnattemptedsolvedpartialfailedpendingsuccess ratetokens
bugswarmkitsoki100001n/an/a
bugswarmraw-prompt100001n/an/a
oss-oraclekitsoki1101000.0002,890,980
oss-oracleraw-prompt100001n/an/a

Overall GLM-5.2 Treatment Rollup ​

treatmentnattemptedsolvedpartialfailedpendingsuccess ratetokens
kitsoki2101010.0002,890,980
raw-prompt200002n/an/a

Kitsoki vs Raw-Prompt Comparisons ​

scopestatusKitsoki attemptedraw attemptedsuccess deltatoken rationotes
bugswarmpending00n/an/aKitsoki GLM-5.2 arm has no attempted cells.; Raw-prompt GLM-5.2 arm has no attempted cells.
oss-oraclepending10n/an/aRaw-prompt GLM-5.2 arm has no attempted cells.
overallpending10n/an/aRaw-prompt GLM-5.2 arm has no attempted cells.

Research Claim Ledger ​

Status: partial (3 supported, 3 pending).

Publication gate:

python3 tools/arena/scripts/glm52_report_gate.py \
  --report-json docs/case-studies/bugswarm-glm52-bugfix-report.data.json \
  --require-publishable
claimstatusfindingmissing evidence / caveat
overall-token-usagependingThe claim is not yet answerable from committed evidenceRaw-prompt GLM-5.2 arm has no attempted cells.; No delta or token ratio is published while the comparison is pending
overall-success-ratependingThe claim is not yet answerable from committed evidenceRaw-prompt GLM-5.2 arm has no attempted cells.; No delta or token ratio is published while the comparison is pending
bugswarm-success-ratependingThe claim is not yet answerable from committed evidenceKitsoki GLM-5.2 arm has no attempted cells.; Raw-prompt GLM-5.2 arm has no attempted cells.; No delta or token ratio is published while the comparison is pending
bugswarm-reusable-sourcesupportedImported BugSwarm task count: 1Execute-mode RED/GREEN verification is still required before live GLM-5.2 cells
oss-source-mixsupported20 tasks over 10 public targets; 6 armed bugfix fixture tasksGLM-5.2 headline cells currently cover only the committed bugfix fixture row
observed-oss-kitsoki-glm52-cellsupported1 attempted cell(s), 2890980 total tokensThis is not a Kitsoki-vs-raw comparison until the matching raw-prompt arm is attempted

Threats To Validity ​

Status: blocked (5 active, 2 high severity).

threatcategoryseveritystatusmitigation
missing-raw-glm52-arminternalhighactiveCommit raw-prompt GLM-5.2 cells for every headline task and regenerate the report
bugswarm-unverified-artifactconstructhighactiveRun bugswarm_verify_source.py --execute, apply the verification report, and regenerate with --bugswarm-verification
single-observed-glm52-cellexternalmediumactiveSchedule the remaining GLM-5.2 cells and report denominators by source family
partial-is-not-solvedconstructmediumactiveKeep partial rate separate from success rate and adjudicate oracle-coupled failures before publication
supporting-round-not-glm52externallowactiveKeep supporting round results out of headline GLM-5.2 denominators

Completion Audit ​

Status: incomplete (4/8 requirements proven).

requirementstatusfindingnext
report-artifactprovenThe report is generated offline from committed inputsdone
oss-sourceprovenThe report references the frozen OSS oracle corpus and keeps it separate from BugSwarmdone
bugswarm-sourceprovenImported BugSwarm task count: 1done
bugswarm-execute-verificationmissingVerification mode=none; verified=0/0Run bugswarm_verify_source.py --execute and apply the verification report
oss-kitsoki-glm52proven1 attempted cell(s), 2890980 total tokensdone
oss-raw-glm52missingNo attempted cell is committed. Pending task(s): kitsoki-bug9-bugfix-test-repairRun the generated gap-plan commands, land the rollup, and regenerate this report
bugswarm-kitsoki-glm52missingNo attempted cell is committed. Pending task(s): bugswarm-square-okio-140452393Run the generated gap-plan commands, land the rollup, and regenerate this report
bugswarm-raw-glm52missingNo attempted cell is committed. Pending task(s): bugswarm-square-okio-140452393Run the generated gap-plan commands, land the rollup, and regenerate this report

Study Protocol ​

Status: pending-evidence. Candidate: glm-5.2. Primary cost metric: total_tokens.

Success metric: solved / (solved + partial + failed).

corpustasktreatmentgate
oss-oraclekitsoki-bug9-bugfix-test-repairraw-promptready-to-plan
bugswarmbugswarm-square-okio-140452393kitsokiexecute-verify-bugswarm
bugswarmbugswarm-square-okio-140452393raw-promptexecute-verify-bugswarm

Execution steps:

  • oss-raw-glm52: ready; Schedule missing OSS oracle raw-prompt GLM-5.2 cells with the frozen corpus manifest. Report regeneration argument: --oss-arena-rollup .artifacts/arena/glm52-oss/rollup.json. Commands:
    • python3 tools/arena/scripts/oss_to_arena_spec.py --report-json docs/case-studies/bugswarm-glm52-bugfix-report.data.json --corpus tools/arena/corpus/cost-bench.manifest.yaml --out .artifacts/arena/oss-glm52.yaml
    • python3 tools/arena/arena.py plan --spec .artifacts/arena/oss-glm52.yaml
    • python3 tools/arena/arena.py run --spec .artifacts/arena/oss-glm52.yaml --out .artifacts/arena/glm52-oss
    • ARENA_PAIRED_TASK_ENABLE_CODEX=1 python3 tools/arena/arena.py run --spec .artifacts/arena/oss-glm52.yaml --out .artifacts/arena/glm52-oss --live
  • bugswarm-execute-verification: required-before-live; Prove BugSwarm failed/passed scripts still reproduce in fresh containers. Report regeneration argument: --bugswarm-verification .artifacts/bugswarm/verification.json. Commands:
    • python3 tools/arena/scripts/bugswarm_verify_source.py --source tools/arena/corpus/bugswarm.seed.yaml --out .artifacts/bugswarm/verification.json --execute
    • python3 tools/arena/scripts/bugswarm_apply_verification.py --source tools/arena/corpus/bugswarm.seed.yaml --verification .artifacts/bugswarm/verification.json --out .artifacts/bugswarm/arena-source.verified.yaml
    • python3 tools/arena/scripts/glm52_gap_plan.py --report-json docs/case-studies/bugswarm-glm52-bugfix-report.data.json --json-out .artifacts/arena/glm52-gap-plan.json --markdown-out .artifacts/arena/glm52-gap-plan.md --bugswarm-source .artifacts/bugswarm/arena-source.verified.yaml

Live controls:

  • The report generator, gap planner, and tests are offline and must not run Docker or LLMs.
  • The operator must run no-LLM arena.py plan and non-live arena.py run before any --live command.
  • Live commands must be explicit and include ARENA_PAIRED_TASK_ENABLE_CODEX=1.
  • GLM-5.2 raw-prompt variants must use backend=claude so paired_task_runner dispatches through the synthetic-claude profile.
  • BugSwarm live cells require execute-mode RED/GREEN verification before model scheduling.

Committed GLM-5.2 Cells ​

tasktreatmentqualitytokenscostevidence
bug9kitsokipartial2,890,980$15.575450tools/bugfix-bakeoff/results/cells/bug9-glm-5.2-kitsoki.json

Committed OSS GLM-5.2 Arena Cells ​

tasktreatmentqualitytokenscostevidence
nonenonependingn/an/an/a

Committed BugSwarm GLM-5.2 Arena Cells ​

tasktreatmentqualitytokenscostevidence
nonenonependingn/an/an/a

Evidence Gaps ​

  • No committed raw-prompt GLM-5.2 result exists for the OSS oracle corpus.
  • BugSwarm artifacts have been imported but none are verified RED/GREEN yet.
  • No BugSwarm verification report is attached to this generated report.
  • Some imported BugSwarm tasks are missing committed GLM-5.2 Kitsoki or raw-prompt result cells.

Evidence Closure Packet ​

Generate the offline execution packet for the pending headline cells with:

python3 tools/arena/scripts/glm52_gap_plan.py \
  --report-json docs/case-studies/bugswarm-glm52-bugfix-report.data.json \
  --json-out .artifacts/arena/glm52-gap-plan.json \
  --markdown-out .artifacts/arena/glm52-gap-plan.md \
  --bugswarm-source tools/arena/corpus/bugswarm.seed.yaml

The packet emits no-spend arena.py plan / arming commands and, only after a spec passes audit, explicit ARENA_PAIRED_TASK_ENABLE_CODEX=1 ... --live commands for operator execution.

corpusstatuspendingnext
oss-oracleready-to-plan1Run glm52_gap_plan.py; it can generate an OSS paired-task spec from the frozen corpus manifest.
bugswarmneeds-execute-verification2Run bugswarm_verify_source.py --execute and apply the verification report before scheduling live GLM-5.2 cells.

Interpretation ​

  • Committed GLM-5.2 Kitsoki evidence contains 1 attempted OSS oracle cell(s), 2890980 total tokens, and no solved cell yet.
  • The GLM-5.2 raw-prompt arm remains pending; the report must not compute a token ratio from missing data.
  • BugSwarm is reusable as an imported source with 1 task(s) in the supplied source file.

Provenance and References ​

Local evidence:

  • tools/bugfix-bakeoff/results/cells — committed Kitsoki/raw-prompt bugfix cells and usage evidence.
  • tools/arena/corpus/cost-bench.manifest.yaml — frozen reusable OSS task source and deterministic oracle metadata.
  • tools/arena/corpus/sources.yaml — adapter contract, required metadata fields, and verification contract.

Upstream references:

BugSwarm seed provenance:

Bottom line: the committed GLM-5.2 evidence is not yet sufficient to claim Kitsoki beats or loses to raw prompts. The report is useful now as a reproducible evidence ledger and corpus scaffold; the headline comparison still requires raw-prompt GLM-5.2 cells and verified BugSwarm cells.

Supporting Codex-Native OSS Round ​

The existing arena round-1 results are supporting evidence for the Kitsoki-vs-raw-prompt harness and token accounting, but they are not GLM-5.2 cells. They should not be used to answer the GLM headline.

treatmentnattemptedsolvedfailedsuccess ratetokens
kitsoki44220.50021,459,517
raw-prompt44220.500537,743