Agentic-Coding Evaluation Lab Methodology Report
This page is an automatically rendered view of a single validated, public-safe evidence file — a demonstration of the reporting layer, not a hand-written document. Before anything renders, the evidence is independently re-validated and re-scanned, and rendering refuses if any required honesty label has drifted.
What this demonstrates: deterministic report rendering, label checking, and refusal behavior on one validated evidence file.
What this does not demonstrate: any claim beyond that one evidence file. The campaign header below records the file’s provenance; every example report published on the public site is synthetic and demonstrates the reporting layer, not model performance.
Campaign Header
campaign_id: synthetic-local-promotion-001
created_at: 2026-07-09T17:34:57+00:00
evidence: evidence/campaigns/synthetic-local-promotion-001/evidence.json
Provenance: synthetic
Sanitization status: public-safe
Sanitization checked: true
Redactions required: false
Capability Table
| Role | Successes / trials | Rate | Wilson CI |
|---|---|---|---|
| incumbent | 5/8 | 62.5% | 95.0% CI [0.305738, 0.863158] |
| candidate | 6/8 | 75.0% | 95.0% CI [0.40927, 0.928522] |
Paired Delta
| Point estimate | Bootstrap CI | Seed | Iterations | Alpha |
|---|---|---|---|---|
| 0.125 | 95.0% CI [-0.25, 0.5] | 12345 | 2000 | 0.05 |
Decision
Decision: false
Label: enforcing:superiority-by-margin
Rule: rule=bootstrap_ci_low_gt_margin
Margin: 0.1
not promotable — failing: beats_incumbent_beyond_band, no_regression_must_not_break
Enhanced Estimators
Enhanced estimators are additive diagnostics for new analyses. They are explicitly non-gating here and do not strengthen or relabel the primary result.
| Estimator | Label | Status | Rendered values | Gate role |
|---|---|---|---|---|
| Two-stage bootstrap | enhanced:two-stage-bootstrap | reported | estimate 0.125; 95.0% CI [-0.375, 0.625]; seed 12345; iterations 2000; alpha 0.05 | non-gating additive estimator |
| Wilcoxon signed-rank | enhanced:wilcoxon-signed-rank | reported | statistic 2; p_value 1; n 3 | non-gating additive estimator |
| GLMM logistic | enhanced:glmm-logistic | optional-dependency-missing | formula n/a; backend n/a | non-gating additive estimator |
Sign Test
The sign test is two-sided, reported-only, and does not gate.
| Wins | Losses | Ties | p_value | alternative | reported_only | Gate role |
|---|---|---|---|---|---|---|
| 2 | 1 | 1 | 1 | alternative:“two-sided” | reported_only:true | does not gate |
Raw Outcomes Summary
Tasks: 4
Replicates: 8
Task classes: fix_failing_test=1, multifile_refactor=1, repo_reasoning=1, small_edit=1
| Task class | Tasks | Replicates | Incumbent successes | Candidate successes |
|---|---|---|---|---|
| fix_failing_test | 1 | 2 | 2 | 1 |
| multifile_refactor | 1 | 2 | 1 | 1 |
| repo_reasoning | 1 | 2 | 1 | 2 |
| small_edit | 1 | 2 | 1 | 2 |
Reproducibility Manifest
Seeds
| Name | Seed |
|---|---|
| bootstrap | 12345 |
| power_simulation | 12345 |
Task Set
| task_set_hash | task_count | replicates_per_task | fixture_policy | task_classes |
|---|---|---|---|---|
| sha256:6c8e84bfd10233d52ea472d69b230a49b7051a89aa1989e2e8334a3056c5a180 | 4 | 2 | synthetic | fix_failing_test, multifile_refactor, repo_reasoning, small_edit |
Models
| Role | Label | Version | Quantization |
|---|---|---|---|
| incumbent | openai/qwen3.5-9b | qwen3.5:9b | n/a |
| candidate | openai/qwen2.5-coder-7b | qwen2.5-coder:7b | n/a |
Cost
| Currency | Total | Basis |
|---|---|---|
| USD | 0 | synthetic local example; no provider-metered cost |
Preflight Results
Status: passed
Checked at: 2026-07-09T00:00:00+00:00
| Model role | Resolved | Model version | Notes |
|---|---|---|---|
| incumbent | true | qwen3.5:9b | |
| candidate | true | qwen2.5-coder:7b |
Core
| Core version | Core content hash |
|---|---|
| 0.2.0 | sha256:7ec59090e3196b225ca2d68bc76cb6703b867f259f87e9d9e2bcafce896bf6ca |