OH OpenCode Harness

DeepSeek-only diagnostic benchmark · June baseline

Turning failed evals into an agent reliability map.

OpenCode Harness ran DeepSeek through smoke, long-context, and repair suites. The results are intentionally published with failures because the point is diagnosis, not a polished leaderboard.

Results

Benchmark snapshot

Three suites, one provider, concrete failure modes.

These are the published pre-feedback-loop results. The site keeps the baseline visible until an authenticated post-fix run is available.

Smoke 1/4 25.0% pass rate
Long Context 1/4 25.0% pass rate
Repair 0/2 0.0% pass rate

What the harness exposed

The score is less important than the failure signature.

Marker-following drift

Cases missed required completion markers such as ZH_MODEL_SUMMARY or reached the step limit before closing.

Tool-loop overrun

Several summaries ended with intent to inspect more files instead of finishing the requested answer.

Long-context synthesis gap

The focused safety explanation passed, while broader repository mapping and eval-flow synthesis were less stable.

Repair finalization gap

Repair cases did not reliably close with the expected markers after editing copied fixture workspaces.

Public evidence

Reports are committed. Raw traces stay local.

Suite Case Outcome Failure Steps
Smoke chinese-coding-task Fail expectation_mismatch 5
Smoke patch-proposal-no-write Pass — 6
Smoke repo-map-orientation Fail expectation_mismatch 7
Smoke tool-calling-stability Fail expectation_mismatch 4
Long context chinese-long-context-summary Fail max_steps 10
Long context cross-file-eval-flow Fail expectation_mismatch 7
Long context repo-wide-module-map Fail expectation_mismatch 5
Long context security-and-permissions-context Pass — 6
Repair repair-calculator Fail expectation_mismatch 5
Repair repair-text-utils Fail expectation_mismatch 8
.\scripts\run-deepseek-benchmark.ps1 -SuiteSet all
python -m opencode_harness diagnose eval-runs/path-to-run/report.json --output eval-runs/deepseek-diagnosis.md
python -m opencode_harness diagnose-compare --before eval-runs/before/report.json --after eval-runs/after/report.json --output eval-runs/deepseek-before-after.md

Reliability iteration

The harness now measures recovery, efficiency, and cost.

Verifier feedback

A failed test result can return to the same agent session as a structured observation before final classification.

Quality and efficiency

Reports now include input/output Tokens, model request time, total case time, and optional price-based cost estimates.

Broader task evidence

Coding Agent Core v0.2 adds 15 verified tasks across repair, testing, refactoring, dependency migration, and security.

Read Reliability Case

Next reliability work

Use the diagnosis to improve the agent loop.

The harness now includes finish-marker reminders, a final-step guard, trace-aware failure-mode diagnosis, before/after diagnosis comparison, and repair verifier feedback. The next iteration should rerun DeepSeek and publish the reliability delta.

Open Full Diagnosis