Skip to content

Latest commit

 

History

History
106 lines (61 loc) · 6.31 KB

File metadata and controls

106 lines (61 loc) · 6.31 KB

Rules

Every check alwayspass performs, what it means, why it is worth reporting, and how to fix it.

Severities: error — the case cannot report a failure, or it measures the wrong thing. warning — it works, but not as it appears to. info — a note about coverage of the analysis.

The behaviour this rests on

From promptfoo's own documentation: a model-graded assertion returns { pass, score }, and without a threshold, the pass field decides. A grader that answers { pass: true, score: 0 } is a pass.

A threshold is resolved from three places, in order:

  1. the assertion,
  2. the test case,
  3. defaultTest.

Any of them makes the score decide instead. That is why the fix for an entire suite is usually one line at the top of the file, and why nothing here is reported when a threshold exists at any level.


graded-assertion-without-threshold — error

A model-graded assertion — llm-rubric, g-eval, factuality, answer-relevance, similar and the rest — with no threshold at any of the three levels.

Why it matters. The rubric is written, the grader is called, the tokens are paid for, and then the judgement it produced is discarded in favour of a boolean it was never really asked about. A suite built this way reports green on almost any output, and it does so on every pull request — which is the worst possible place for a false assurance, because it is exactly where somebody decides a change is safe.

Fix. threshold on the assertion, on the test, or once on defaultTest.


assertion-cannot-fail — error

An assertion whose value is satisfied by every possible output.

Reported forms: contains, icontains, starts-with or contains-any of the empty string; a regex of .*, ^.*$, (.*) or empty; a javascript or python body of true, return true, 1, output.length >= 0, output !== undefined; and any assertion with threshold: 0.

Why it matters. Each of these runs, reports success, and tells you nothing. They usually arrive as a placeholder somebody meant to fill in before the pull request, and in a green report they are indistinguishable from an assertion that works.

Fix. Replace the value with one a wrong answer would fail, or delete the assertion so the case is honestly untested rather than dishonestly tested.


test-asserts-nothing — error

A test case with no assertions of its own and none inherited from defaultTest.

Why it matters. The prompt is rendered, the model is called, the tokens are spent, and the result is recorded as a pass because nothing threw. It is a test of the network. It also inflates the suite: "212 cases passing" reads as coverage, and some of those cases assert nothing at all.

Fix. Add an assertion, or move these to a separate configuration for outputs you only want to look at.


expected-output-is-in-the-input — error

A deterministic assertion expects a string of twelve characters or more that one of the test's own vars already contains.

vars:
  question: "My order ORD-99182-XY has not arrived, where is ORD-99182-XY?"
assert:
  - type: contains
    value: "ORD-99182-XY"

Why it matters. The model is handed the answer and then asked for it. The assertion passes whatever the model understands, because copying is enough, so the case measures retrieval from the prompt rather than the behaviour it was written for. This is the quietest way for a suite to look like it covers something it has never tested — and it survives review, because the test reads as though it is about extraction.

What is not reported. Expected strings shorter than twelve characters, because short strings collide by accident. Model-graded assertions, whose value is prose about the answer rather than the answer itself. And not- assertions, where the input containing the string is the whole point — asserting that a credit card number in the input is not echoed back is a real check, and the commonest shape of one.

Fix. Assert on something the input does not already contain, or change the input so the answer has to be produced rather than repeated.


assertion-weighted-to-zero — warning

An assertion with weight: 0.

Why it matters. A zero weight removes the assertion from the weighted score, so it cannot move the verdict in either direction. It still runs, and a model-graded one still costs a call to the grader on every evaluation. It is almost always a temporary silencing that outlived the reason for it.

Fix. Give it a real weight, or remove it. A muted assertion sitting in a suite reads as a check that is being made.


provider-not-pinned — warning

A provider identifier that names a moving alias rather than a version — openai:gpt-4.1 rather than openai:gpt-4.1-2025-04-14.

Why it matters. The alias points at whatever the provider currently ships. Two runs of the same commit can measure two different models, so a score that moves says nothing about whether your change was an improvement — which is the one question the suite exists to answer. It is the eval equivalent of an unpinned dependency, and more consequential, because the output is a judgement rather than a build.

Fix. Pin the version, and change it deliberately when you want to measure a new model.


Deliberate non-findings

Cases where alwayspass stays quiet on purpose.

Situation Why nothing is reported
A rubric with a threshold anywhere in the chain The score decides. That is all this tool asks for.
A not-contains matching text in the input Asserting the output does not echo the input is a real check.
A short expected string appearing in vars Short strings collide by accident; accusing them would fire everywhere.
A javascript body that is not a recognised always-true form Assumed able to fail. Guessing at arbitrary code would produce confident, wrong findings.
A rubric whose wording is vague Whether a rubric discriminates is a question for a human, or for running it.
A suite whose tests come from a file or a dataset Reported as unreadable from here, not as clean.

The bias throughout is toward under-reporting. A missed finding costs one weak test. A false positive costs the user's trust in every other finding, and then the check gets removed — leaving a suite that still cannot fail and now looks audited.