Draft: production coding runs and benchmark-driven hardening - #39
Merged
Conversation
Run review through a read-only D0 tool loop, teach the Claude account router its inert host-tool protocol, and retain relocation-safe external benchmark evidence. Task 049 now passes unchanged with native success and outcome/process/security/combined 1.0/0.9867/1.0/0.9867.
Require the next missing repository-grounding operation at the provider boundary so repeated premature terminals cannot consume the review/fix window. Preserve explicit output alternatives as at-least-one groups instead of flattening them into cumulative requirements.
Partition Rust and Verus work by architecture layer, split H2 scenarios and release construction, and execute H0 probes in bounded worker partitions. Enforce a ten-minute ceiling for every hosted workflow job while preserving required Gate A status names and complete qualification coverage.
Run Rust shard controllers outside the artifact tree they validate so Windows never replaces a live xtask.exe. Split H2 into one scenario per native job and keep the Linux-only H0 operator test off macOS.
dollspace-gay
marked this pull request as ready for review
September 2, 2026 23:34
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Current validation
1025086a446653702b80cfb48babbeec35db6b2c: outcome 0.8969, process 0.9286, security 1.0000, combined 0.8331. Forty tasks were perfect and 64 scored at least 0.9.26a981ef968443504a0d47d420e34783fa99c2b9d8b1d876661415175a8dbd3e. The Rust publisher verifies exact pinned-catalog coverage, selected report paths and hashes, aggregate scores, elapsed time, and usage./home/doll/.local/state/peritus/qualification/h1/dependency-routes.5nwGZV. The artifact-quota route additionally forces a real durable-catalog quota race after object publication, proves rollback removes the losing object and metadata, and retains a passing six-class report at/home/doll/.local/state/peritus/qualification/h1/blob-finalize-disk.fifreA/report.json(SHA-2560f114d8b503be7259cae2ff3a8666dec1094d8ad257867a3ebee748b57be1f19).b0e65ee1is fully green across Gate A, Foundation, native H0 security, and native H2 package workflows on Linux, macOS, Windows, and Verus. Current signed head30b15608adds the six dependency routes and real artifact-quota rollback route, then restores full-workspace Verus verification by classifying the qualification-only dependency admin boundary consistently with its product-runner effect dependency. A fresh hosted runner wave is in progress.30b15608passes exact all-feature daemon Verus verification, strict all-target/all-feature Clippy, 54 daemon unit tests plus 24 integration tests, all public daemon subprocess conformance, ordinary-API policy, documentation validation, architecture validation, formatting, diff hygiene, the 500-line source ceiling, all six dependency operator runs, and the real artifact-quota rollback diagnostic.Benchmark integrity rule
The final report will contain a benchmark-gotchas table. For every underspecified or hidden expectation it will record the published contract, the hidden expectation discovered only after scoring, the honest retained result, and the shortcut refused. Installing ordinary prerequisites such as Python, Make, R, compilers, or normal packages inside an authorized disposable task environment and keeping them discoverable on
PATHis legitimate agent execution. Peritus will not read verifier or reference-solution internals to solve a task, key behavior to task names or private benchmark vocabulary, alter fixtures, resources, deadlines, or scoring, or add score-only behavior.Why this is still a draft
Issue: #31