Benchmark snapshot
Three suites, one provider, concrete failure modes.
These are the published pre-feedback-loop results. The site keeps the baseline visible until an authenticated post-fix run is available.
DeepSeek-only diagnostic benchmark · June baseline
OpenCode Harness ran DeepSeek through smoke, long-context, and repair suites. The results are intentionally published with failures because the point is diagnosis, not a polished leaderboard.
Benchmark snapshot
These are the published pre-feedback-loop results. The site keeps the baseline visible until an authenticated post-fix run is available.
What the harness exposed
Cases missed required completion markers such as ZH_MODEL_SUMMARY or reached the step limit before closing.
Several summaries ended with intent to inspect more files instead of finishing the requested answer.
The focused safety explanation passed, while broader repository mapping and eval-flow synthesis were less stable.
Repair cases did not reliably close with the expected markers after editing copied fixture workspaces.
Public evidence
| Suite | Case | Outcome | Failure | Steps |
|---|---|---|---|---|
| Smoke | chinese-coding-task |
Fail | expectation_mismatch |
5 |
| Smoke | patch-proposal-no-write |
Pass | — | 6 |
| Smoke | repo-map-orientation |
Fail | expectation_mismatch |
7 |
| Smoke | tool-calling-stability |
Fail | expectation_mismatch |
4 |
| Long context | chinese-long-context-summary |
Fail | max_steps |
10 |
| Long context | cross-file-eval-flow |
Fail | expectation_mismatch |
7 |
| Long context | repo-wide-module-map |
Fail | expectation_mismatch |
5 |
| Long context | security-and-permissions-context |
Pass | — | 6 |
| Repair | repair-calculator |
Fail | expectation_mismatch |
5 |
| Repair | repair-text-utils |
Fail | expectation_mismatch |
8 |
.\scripts\run-deepseek-benchmark.ps1 -SuiteSet all
python -m opencode_harness diagnose eval-runs/path-to-run/report.json --output eval-runs/deepseek-diagnosis.md
python -m opencode_harness diagnose-compare --before eval-runs/before/report.json --after eval-runs/after/report.json --output eval-runs/deepseek-before-after.md
Reliability iteration
A failed test result can return to the same agent session as a structured observation before final classification.
Reports now include input/output Tokens, model request time, total case time, and optional price-based cost estimates.
Coding Agent Core v0.2 adds 15 verified tasks across repair, testing, refactoring, dependency migration, and security.
Next reliability work
The harness now includes finish-marker reminders, a final-step guard, trace-aware failure-mode diagnosis, before/after diagnosis comparison, and repair verifier feedback. The next iteration should rerun DeepSeek and publish the reliability delta.
Open Full Diagnosis