Skip to content

Disambiguate EIP Complexity Template language for better use with LLMs - #147

Open
danceratopz wants to merge 10 commits into
mainfrom
eip-complexity-v3
Open

danceratopz wants to merge 10 commits into
mainfrom
eip-complexity-v3

Conversation

@danceratopz

@danceratopz danceratopz commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Background

Evaluating EIPs with Jev and the v2 complexity template exposed ambiguities in questions that should have straightforward answers. For example, the v2 assessments assigned:

These examples motivated a broader review of how the template describes testing work, prerequisites and evidence, in order to enable more reliable AI-based EIP assessments.

Changes

This PR revises the assessment language and guidance to improve automated evaluation while keeping the template usable by human reviewers. It keeps the same 28 criteria, numeric score options, tier thresholds and cross-EIP bonus formula. It revises definitions and scope, so scores are not directly interchangeable with earlier revisions.

The main changes are:

  • Define the scope as additional execution-layer testing work, including EL support for cross-layer proposals.
  • Assess dependent EIPs against the base fork plus their required prerequisites, including proposals for the same fork. Count the target's additional regression and interaction tests without charging it again for inherited mechanisms.
  • Make distinctions such as transaction types versus execution requests, and new contract deployment versus modification, explicit.
  • Clarify what each score requires, including the evidence needed for exceptional scores and how to handle missing information, and state that the cross-EIP bonus counts only target-specific interaction cases, at any base score.
  • Include performance validation of existing workloads after gas-price or resource-limit changes.

Other changes:

  1. Replace “anchor” with “criterion” throughout the template.
  2. Add stable abbreviations, such as PAT, INV and PERF, to criterion headings and the checklist table.
  3. Ask assessors to record the EIP/template revisions, baseline, prerequisites, sources and material assumptions; document the changes as checklist revision 3.

The companion dashboard compares v2 and v3 on identical EIP text across 48 assessments. In the displayed v3 results, the probabilities above fall to 1%, 2% and 7%, respectively. Some individual judgments remain questionable or move in undesirable directions; this is evidence for review, not a claim of improvement across every criterion. Historical human/LLM comparisons and a separate repeatability study are included, with their limitations noted.

Use the same Cryptography label in the criterion and checklist row.
Define additional execution-layer testing work as the assessment scope,
including execution-client support for cross-layer changes.

Assess dependent EIPs relative to their required prerequisites, including
same-fork proposals. Attribute inherited work to its defining EIP while
retaining target-required changes and regression or interaction tests.
Explain that this baseline sets neither activation order nor bundle cost.

Require supplied evidence, supporting sources and material gaps to be
explicit. A dependency reference does not establish its semantics, and
missing evidence does not establish either low or high complexity.

Define test families and baseline tests without assuming particular test
files exist. This separates behavioral breadth from fixture counts.

Place provenance requirements in Notes. Record source, rubric, baseline
and assumptions, plus model and instruction versions for automated use.

Define the conditional baseline as the stated base fork plus required mechanisms, including dependencies absent from the requires header; record both without assuming missing semantics.
Apply the tested revision 3 definitions and score-selection guidance,
with a smaller set of clarifications following the local studies:

- PAT excludes constructing prerequisite suites, not target-required updates.
- INV can apply alongside changed expectations and requires a distinct output.
- SGAS excludes state-access and BAL-size charges alone.
- Modified opcode semantics explicitly include emitted logs.

Retain the original v3 PAT scoring levels and PERF wording. Preserve score
domains, tier thresholds, the cross-EIP bonus and table formatting.
The replacement v3a run evaluated the substantive candidate before final
editorial cleanup. Remove study-only wording and the code-only bonus
instruction; retain UNSP in the total. Wording changes alone are not
evidence of improved assessment quality.

Resolve the EDGE exception contradiction, include added CRYP mechanisms and protocol account/state encodings, restore the SAO matrix note, and count target-specific XEIP combinations with unchanged rules. Preserve numeric levels and the bonus formula.
Update the revision number after the assessment guidance and criteria.
Explain that revision 3 disambiguates the text to improve AI assessments
while remaining usable by humans. Preserve earlier revision history and
note that scores are not directly interchangeable across revisions.

Disclose substantive interpretation choices for encoding, cross-EIP combinations and normative omissions; unchanged numeric options do not imply interchangeable scores.
Recognize renewed benchmarking of existing workloads after changes to gas
costs or resource limits, even when the underlying mechanisms are unchanged.
Attribute only the validation required by the changed capacity or bound.

Keep the PERF definition and scoring levels unchanged. Evaluate this note
as v3b, with the current committed template as the study snapshot.
Clarify that deploying a newly introduced system contract does not modify an existing system contract, even when its address already has an empty account. Assess deployment under +SC and activation work under FORK; changes to contracts in the prerequisite baseline still count under ~SC.
@danceratopz danceratopz changed the title Clarify the EIP complexity assessment criteria Disambiguate EIP Complexity Template language for better use with LLMs Oct 7, 2026
@danceratopz
danceratopz marked this pull request as draft October 7, 2026 10:48
Ask for coordinated cases the target adds, and exclude citations and tests
of the other EIP's own behavior, instead of excluding "unchanged
inheritance". The previous wording led automated helpers to reject almost
every interaction, including recorded system-contract and access-list
behavior for EIP-7928, so no assessment received a bonus.
The template did not say whether the bonus requires a base score of 3.
Make explicit that it depends only on the number of qualifying EIPs, which
matches how the bonus has been computed. Tested wording moves base
cross-EIP scores by 0.07 on average across 48 EIPs.
Define a new invariant as an output that did not exist before the target,
with examples, and assign changed costs, limits or expected values to
Patterns affecting pre-existing tests. Attribute prerequisite assertions to
the prerequisite. Drop the "score both only when both are required"
sentence, which could be read as a gate on either criterion.
danceratopz added a commit to danceratopz/eip-complexity that referenced this pull request Oct 7, 2026
Pin PM commit fe2f793 (ethspecs/pm#147) with the target-specific cross-EIP
bonus wording, an explicit any-base-score bonus rule and the INV
new-output/rework split. Revise the cross-EIP helper policy as question set
4. Regenerate the 48 v3 requests on unchanged source bytes; retain the
verified v2 responses. Side-by-side tests are recorded in tmp/xeip-tuning.
danceratopz added a commit to danceratopz/eip-complexity that referenced this pull request Oct 7, 2026
Point the PM commit link at the pull request's commit view and link the
pull request itself.
@danceratopz
danceratopz marked this pull request as ready for review October 7, 2026 14:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant