Skip to content

Draft: production coding runs and benchmark-driven hardening - #39

Merged
dollspace-gay merged 271 commits into
mainfrom
feature/g4-conversational-runs
Sep 2, 2026
Merged

Draft: production coding runs and benchmark-driven hardening#39
dollspace-gay merged 271 commits into
mainfrom
feature/g4-conversational-runs

Conversation

@dollspace-gay

@dollspace-gay dollspace-gay commented Aug 29, 2026

Copy link
Copy Markdown
Member

Summary

  • Complete the G4 conversational coding product path with durable objectives, repository grounding, detailed designs, exact-target gates, independent review, user correction, restart recovery, bounded provider recovery, deterministic context compaction, prompt-cache accounting, effect receipts, live progress, and explicit default-off provider failover.
  • Add real Rust-owned HarnessBench and Terminal-Bench 2.0 adapters with thin external boundary glue; retain benchmark workspaces, traces, usage, provider identity, exact upstream revisions, and an evidence-backed failure journal outside Git.
  • Complete the unchanged 106-task HarnessBench diagnostic campaign and apply only general, regression-tested harness corrections. Hidden, contradictory, or defective oracle expectations retain their honest lower scores instead of receiving task-name or hidden-vocabulary hacks.
  • Add plain-English documentation validation, public POSIX and PowerShell bootstrap installers, transactional self-update and rollback, three-platform native packaging, H0-H4 qualification machinery, and candidate-bound SBOM/provenance/signature evidence.
  • Keep generated benchmark reports, sandboxes, credentials, release binaries, caches, and local provider/Crosslink bootstrap metadata outside Git.

Current validation

  • HarnessBench completed all 106 tasks at pinned upstream revision 1025086a446653702b80cfb48babbeec35db6b2c: outcome 0.8969, process 0.9286, security 1.0000, combined 0.8331. Forty tasks were perfect and 64 scored at least 0.9.
  • The retained diagnostic HarnessBench report has SHA-256 26a981ef968443504a0d47d420e34783fa99c2b9d8b1d876661415175a8dbd3e. The Rust publisher verifies exact pinned-catalog coverage, selected report paths and hashes, aggregate scores, elapsed time, and usage.
  • The frozen Terminal-Bench 2.0 k=5 baseline is running serially at concurrency 1. Its current immutable prefix contains 232/445 completed trials: 128 reward 1, 87 reward 0, and 17 completed without a score. Accuracy is 59.53% across scored trials and 55.17% across all completed trials. This is provisional until the full baseline closes.
  • Terminal-Bench failures have produced task-neutral corrections for provider-turn recovery, exact workspace mapping, nested-repository grounding, prerequisite installation, external-effect verification, bounded command output, cancellation/reaping, independent-input validation, and closed mutation contracts. The running baseline is not changed or restarted as those fixes land.
  • H0 has real candidate-bound Linux, macOS, and Windows controllers, exact shard interchange, independent-review admission, and deterministic three-host aggregation.
  • H1 now has 25 genuine native routes: both sides of journal, blob, retained Git snapshot, lease, patch, D1 gate, and F0 promotion commits; active-projection, committed-journal, referenced-blob, and retained-snapshot corruption containment; and provider, tool, and worker death plus retry exhaustion. The six new dependency routes exercise the real executable-backed provider transport, ordinary grounded receipt-backed product tool, daemon worker supervisor, durable scheduler, and fresh journal replay. All six diagnostics passed with 36 retained evidence files under /home/doll/.local/state/peritus/qualification/h1/dependency-routes.5nwGZV. The artifact-quota route additionally forces a real durable-catalog quota race after object publication, proves rollback removes the losing object and metadata, and retains a passing six-class report at /home/doll/.local/state/peritus/qualification/h1/blob-finalize-disk.fifreA/report.json (SHA-256 0f114d8b503be7259cae2ff3a8666dec1094d8ad257867a3ebee748b57be1f19).
  • H2 has real packaged-host controllers on all three platforms, including installed TUI PTY/ConPTY lifecycle, upgrade, rollback, uninstall, checksum rejection, state preservation, and sandbox activation.
  • Exact checkpoint b0e65ee1 is fully green across Gate A, Foundation, native H0 security, and native H2 package workflows on Linux, macOS, Windows, and Verus. Current signed head 30b15608 adds the six dependency routes and real artifact-quota rollback route, then restores full-workspace Verus verification by classifying the qualification-only dependency admin boundary consistently with its product-runner effect dependency. A fresh hosted runner wave is in progress.
  • Local validation for 30b15608 passes exact all-feature daemon Verus verification, strict all-target/all-feature Clippy, 54 daemon unit tests plus 24 integration tests, all public daemon subprocess conformance, ordinary-API policy, documentation validation, architecture validation, formatting, diff hygiene, the 500-line source ceiling, all six dependency operator runs, and the real artifact-quota rollback diagnostic.

Benchmark integrity rule

The final report will contain a benchmark-gotchas table. For every underspecified or hidden expectation it will record the published contract, the hidden expectation discovered only after scoring, the honest retained result, and the shortcut refused. Installing ordinary prerequisites such as Python, Make, R, compilers, or normal packages inside an authorized disposable task environment and keeping them discoverable on PATH is legitimate agent execution. Peritus will not read verifier or reference-solution internals to solve a task, key behavior to task names or private benchmark vocabulary, alter fixtures, resources, deadlines, or scoring, or add score-only behavior.

Why this is still a draft

  • Complete all 445 frozen Terminal-Bench trials, classify every nonpass and painful pass, and rerun unchanged focused comparisons against the corrected candidate.
  • Run the pinned final complete HarnessBench and Terminal-Bench campaigns sequentially with one frozen revision-bound binary.
  • Publish honest before/after aggregates that separate Peritus defects, legitimate model/capability limits, provider or infrastructure failures, and benchmark underspecification.
  • Finish the remaining genuine H0-H4 evidence, including the reference-machine eight-hour H3 soak or an honest external-host limitation, independent final audit, and exact-candidate release bundle.
  • Keep the final exact-head hosted runners green and link every production-readiness claim to reproducible commands and retained evidence.

Issue: #31

Run review through a read-only D0 tool loop, teach the Claude account router its inert host-tool protocol, and retain relocation-safe external benchmark evidence. Task 049 now passes unchanged with native success and outcome/process/security/combined 1.0/0.9867/1.0/0.9867.
Require the next missing repository-grounding operation at the provider boundary so repeated premature terminals cannot consume the review/fix window. Preserve explicit output alternatives as at-least-one groups instead of flattening them into cumulative requirements.
Partition Rust and Verus work by architecture layer, split H2 scenarios and release construction, and execute H0 probes in bounded worker partitions. Enforce a ten-minute ceiling for every hosted workflow job while preserving required Gate A status names and complete qualification coverage.
Run Rust shard controllers outside the artifact tree they validate so Windows never replaces a live xtask.exe. Split H2 into one scenario per native job and keep the Linux-only H0 operator test off macOS.
@dollspace-gay
dollspace-gay marked this pull request as ready for review September 2, 2026 23:34
@dollspace-gay
dollspace-gay merged commit 300d6e0 into main Sep 2, 2026
311 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant