Repository navigation
Disambiguate EIP Complexity Template language for better use with LLMs - #147
Open
danceratopz wants to merge 10 commits into
Open
danceratopz wants to merge 10 commits into
danceratopz wants to merge 10 commits into
Conversation
Use the same Cryptography label in the criterion and checklist row.
Define additional execution-layer testing work as the assessment scope, including execution-client support for cross-layer changes. Assess dependent EIPs relative to their required prerequisites, including same-fork proposals. Attribute inherited work to its defining EIP while retaining target-required changes and regression or interaction tests. Explain that this baseline sets neither activation order nor bundle cost. Require supplied evidence, supporting sources and material gaps to be explicit. A dependency reference does not establish its semantics, and missing evidence does not establish either low or high complexity. Define test families and baseline tests without assuming particular test files exist. This separates behavioral breadth from fixture counts. Place provenance requirements in Notes. Record source, rubric, baseline and assumptions, plus model and instruction versions for automated use. Define the conditional baseline as the stated base fork plus required mechanisms, including dependencies absent from the requires header; record both without assuming missing semantics.
Apply the tested revision 3 definitions and score-selection guidance, with a smaller set of clarifications following the local studies: - PAT excludes constructing prerequisite suites, not target-required updates. - INV can apply alongside changed expectations and requires a distinct output. - SGAS excludes state-access and BAL-size charges alone. - Modified opcode semantics explicitly include emitted logs. Retain the original v3 PAT scoring levels and PERF wording. Preserve score domains, tier thresholds, the cross-EIP bonus and table formatting. The replacement v3a run evaluated the substantive candidate before final editorial cleanup. Remove study-only wording and the code-only bonus instruction; retain UNSP in the total. Wording changes alone are not evidence of improved assessment quality. Resolve the EDGE exception contradiction, include added CRYP mechanisms and protocol account/state encodings, restore the SAO matrix note, and count target-specific XEIP combinations with unchanged rules. Preserve numeric levels and the bonus formula.
Update the revision number after the assessment guidance and criteria. Explain that revision 3 disambiguates the text to improve AI assessments while remaining usable by humans. Preserve earlier revision history and note that scores are not directly interchangeable across revisions. Disclose substantive interpretation choices for encoding, cross-EIP combinations and normative omissions; unchanged numeric options do not imply interchangeable scores.
Recognize renewed benchmarking of existing workloads after changes to gas costs or resource limits, even when the underlying mechanisms are unchanged. Attribute only the validation required by the changed capacity or bound. Keep the PERF definition and scoring levels unchanged. Evaluate this note as v3b, with the current committed template as the study snapshot.
Clarify that deploying a newly introduced system contract does not modify an existing system contract, even when its address already has an empty account. Assess deployment under +SC and activation work under FORK; changes to contracts in the prerequisite baseline still count under ~SC.
danceratopz
marked this pull request as draft
October 7, 2026 10:48
Ask for coordinated cases the target adds, and exclude citations and tests of the other EIP's own behavior, instead of excluding "unchanged inheritance". The previous wording led automated helpers to reject almost every interaction, including recorded system-contract and access-list behavior for EIP-7928, so no assessment received a bonus.
The template did not say whether the bonus requires a base score of 3. Make explicit that it depends only on the number of qualifying EIPs, which matches how the bonus has been computed. Tested wording moves base cross-EIP scores by 0.07 on average across 48 EIPs.
Define a new invariant as an output that did not exist before the target, with examples, and assign changed costs, limits or expected values to Patterns affecting pre-existing tests. Attribute prerequisite assertions to the prerequisite. Drop the "score both only when both are required" sentence, which could be read as a gate on either criterion.
danceratopz
force-pushed
the
eip-complexity-v3
branch
from
October 7, 2026 13:36
6c66fa0 to
fe2f793
Compare
danceratopz
added a commit
to danceratopz/eip-complexity
that referenced
this pull request
Oct 7, 2026
Pin PM commit fe2f793 (ethspecs/pm#147) with the target-specific cross-EIP bonus wording, an explicit any-base-score bonus rule and the INV new-output/rework split. Revise the cross-EIP helper policy as question set 4. Regenerate the 48 v3 requests on unchanged source bytes; retain the verified v2 responses. Side-by-side tests are recorded in tmp/xeip-tuning.
danceratopz
added a commit
to danceratopz/eip-complexity
that referenced
this pull request
Oct 7, 2026
Point the PM commit link at the pull request's commit view and link the pull request itself.
danceratopz
marked this pull request as ready for review
October 7, 2026 14:20
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Background
Evaluating EIPs with Jev and the v2 complexity template exposed ambiguities in questions that should have straightforward answers. For example, the v2 assessments assigned:
These examples motivated a broader review of how the template describes testing work, prerequisites and evidence, in order to enable more reliable AI-based EIP assessments.
Changes
This PR revises the assessment language and guidance to improve automated evaluation while keeping the template usable by human reviewers. It keeps the same 28 criteria, numeric score options, tier thresholds and cross-EIP bonus formula. It revises definitions and scope, so scores are not directly interchangeable with earlier revisions.
The main changes are:
Other changes:
PAT,INVandPERF, to criterion headings and the checklist table.The companion dashboard compares v2 and v3 on identical EIP text across 48 assessments. In the displayed v3 results, the probabilities above fall to 1%, 2% and 7%, respectively. Some individual judgments remain questionable or move in undesirable directions; this is evidence for review, not a claim of improvement across every criterion. Historical human/LLM comparisons and a separate repeatability study are included, with their limitations noted.