Skip to content

Add workflow validation, runner artifacts and merged reports - #32

Open
JE-Chen wants to merge 35 commits into
devfrom
feature/testpioneer-platform-improvements
Open

JE-Chen wants to merge 35 commits into
devfrom
feature/testpioneer-platform-improvements

Conversation

@JE-Chen

@JE-Chen JE-Chen commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

Summary

Implements phases 1, 2, 3 and 5 of docs/roadmap-platform-improvements.md. A workflow can be checked without running it, every run has an ID and a result, the artifacts of what failed are kept, and the runners' own reports are merged into one report.

Existing workflows run unchanged, and python -m test_pioneer -e behaves as before, including its exit status.

What is in it

Phase 1 — Contracts and validation

  • Run/artifact/result data models (test_pioneer/models/)
  • schema/testpioneer.schema.json, built from one description of the workflow keys (test_pioneer/schema/)
  • load_yaml, validate_yaml, lint_yaml, get_yaml_schema
  • python -m test_pioneer validate <files> [--format json] [--strict], exit 1 on an error, every problem with line and column; - reads the workflow from standard input, for an editor's unsaved buffer
  • Unit tests, plus a cross-check of the built-in schema validator against jsonschema

Phase 2 — Runner artifacts

  • Run ID and artifacts/<run-id>/ with one directory per runner execution
  • TEST_PIONEER_RUN_ID and TEST_PIONEER_ARTIFACT_DIR passed to every runner
  • Exit code, timing and output of each runner recorded; manifest.json and execution.log
  • Runner adapter contract (test_pioneer/runner/)
  • Artifacts of what did not pass are kept; a run that passed leaves nothing (keep_artifacts)
  • python -m test_pioneer run <file>, exit 1 when the run did not pass

Phase 3 — Report merge

  • Normalized result with the tests each runner recorded
  • Readers for the report pair of the API, web, load and GUI runners
  • report/testpioneer-report.json and .html after every run, optional JUnit XML
  • A step declares the files its runner writes with artifacts:; a failed test in a runner report fails the run
  • Integration tests with several real runner processes

Phase 5 — Documentation

  • README quick start rewritten as one complete example, in English, zh-TW and zh-CN, extracted and run as written
  • docs/validation.rst, docs/artifacts.rst, docs/reports.rst, docs/migration.rst
  • architecture.md, progress.md, docs/updates/

Phase 4 — UI / editor

  • Not in this repository. The editor and workspace belong in PyBreeze; this PR provides what they need: validate --format json (also on standard input), the schema, run with its exit status, and the report files, with a JUnit file PyBreeze's report viewer reads as it is (architecture.md §6).
  • It is not started there, for two reasons recorded in progress.md Dev #9: the commands have to be in a released test_pioneer first, and PyBreeze's report viewer and language service are on a local stack that is not pushed yet.

Also fixed on the way

  • The sample workflows under test/unit_test/ named scripts that could not be found, so the integration job passed without executing anything. They now run real scripts against a local page and fail CI when they fail.
  • A download_file step whose download failed was reported as done.
  • A second execute_yaml call in one process stopped with "job name duplicated".
  • An in-process run step's report repeated the records of earlier steps of the same runner; each record is now counted once.
  • je-mail-thunder was declared but never imported; it is no longer a dependency.

Behaviour to know about

  • New output: report/ after every run, artifacts/<run-id>/ for a run that did not pass. report_formats: [] and keep_artifacts: never turn them off.
  • execute_yaml returns a RunResult instead of None.
  • A run can be failed where it used to end silently: a parallel_run runner that exits non-zero, a failed test in a collected runner report, a failed download.
  • The runner packages exit 0 when an action fails and do not read the artifact variables yet. Until they do, a failed action is detected through the runner's report and artifacts:. What native support means is written down in docs/artifacts.rst, TestPioneer already handles a runner that has it, and progress.md TestPioneer Auto Set Env #13 lists the steps per runner. LoadDensity's dev branch already exits non-zero after a failed action; the others can do the same without a change to the shared executor.
  • A JUnit timestamp is a UTC date-time to the second without a zone suffix.

docs/migration.rst lists all of it.

Checks

  • 565 tests pass locally with every runner package installed (2 more need jsonschema, which CI's unit-test job installs). CI runs them on Python 3.10 to 3.14, and its integration job now executes the sample workflows for real.
  • mypy --strict, ruff and pylint are clean for the new modules; SonarCloud and Codacy pass.
  • The README example and the four sample workflows were executed with the real API, load and file runners.

Decisions taken in this PR

Each has an entry in docs/updates/2026-10.md with the reasoning.

  • Run output is on by default: the report after every run, artifacts for a run that did not pass (U-20261008-13).
  • je-mail-thunder is dropped, not wired up (U-20261008-10).
  • Step names are unique per workflow, not per process; opening a program under a name that is still open fails the step (U-20261008-08).
  • The sample workflows are hermetic and their scripts fail CI when the run fails (U-20261008-07).
  • The runner packages are not changed by this PR; what they need is written down (U-20261008-15, corrected by U-20261008-16).

Merging

  • The base is dev, and dev is merged into this branch, so the merge is clean.
  • The branch started from main, so it also carries main's version-bump commits: pyproject.toml goes from 0.1.33 to 0.1.37 on dev. Nothing else comes from main.
  • A merge into dev publishes test_pioneer_dev from CI.

@codacy-production

codacy-production Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Up to standards ✅

🟢 Issues 0 issues

Results:
0 new issues

View in Codacy

🟢 Metrics 1126 complexity · 2 duplication

Metric Results
Complexity 1126
Duplication 2

View in Codacy

NEW Get contextual insights on your PRs based on Codacy's metrics, along with PR and Jira context, without leaving GitHub. Enable AI reviewer
TIP This summary will be updated as you push new changes.

Phase 1 of the platform roadmap (PR #32). A workflow can now be checked
without executing it, and every problem is reported with its YAML path,
line and column.

- schema: spec.py describes the top-level keys, the step types and their
  fields once; definition.py builds the versioned JSON Schema from it and
  schema/testpioneer.schema.json is that schema published for editors.
- validation: load_yaml keeps the position of every key and value,
  validate_yaml checks syntax and schema with a built-in validator (no
  new dependency), lint_yaml adds the rules a schema cannot express.
- cli: "python -m test_pioneer validate" exits 1 on an error and has a
  JSON output for tools; "schema" prints the schema. -e/--execute_yaml
  is unchanged and execute_yaml does not call the validator.
- models: Diagnostic and ValidationResult, plus the normalized run
  result that the artifact and report phases fill in.
- runner/registry.py holds the runner-to-package table that parallel_run
  kept privately; the schema and the linter read the same table.

CI validates the bundled sample workflows and installs jsonschema so the
built-in validator is compared with the reference implementation.
Phase 2 of the platform roadmap (PR #32). parallel_run started its
runners and waited; nothing recorded how they ended, and a runner that
failed left no trace and did not change the outcome.

- execute_yaml opens a run session and returns a RunResult: each step
  and each runner execution with its status, timing and message.
- Artifacts go to <artifacts_path>/<run-id>/: one directory per runner
  execution plus testpioneer/execution.log and manifest.json. With the
  default keep_artifacts: on_failure a run that passed leaves nothing
  and a failed one keeps what did not pass.
- Each runner execution gets TEST_PIONEER_RUN_ID and
  TEST_PIONEER_ARTIFACT_DIR. A runner that ignores them is wrapped: its
  output is kept as stdout.log and stderr.log, still shown on the
  console, and its exit code is recorded. No thread is used.
- runner/adapter.py defines the adapter contract; the registry holds one
  adapter per runner tag.
- "python -m test_pioneer run" exits 1 when the run did not pass. The
  -e flag still exits 0 whatever the steps did.

A failed parallel_run runner still does not stop later steps, and
exceptions propagate as before. A handler called outside execute_yaml
records nothing. A problem while collecting artifacts becomes a warning
on the result and never changes the outcome.
Phase 3 of the platform roadmap (PR #32). Each runner wrote its own
report wherever its script said, a run had no report of its own, and a
failed action was invisible because the runners exit 0 either way.

- A step names the files its runner writes with "artifacts:". Files
  changed since the runner started are copied into its artifact
  directory; patterns that leave the working directory are refused.
- report/readers.py reads the *_success.json / *_failure.json pair of
  the API, web, load and GUI runners into test cases. Each runner's
  adapter carries the reader for its field names.
- A runner execution that ended normally but recorded a failed test is
  failed, and so are its step and the run.
- Every run writes report/testpioneer-report.json and .html; JUnit XML
  is optional. report_path and report_formats are new workflow keys,
  RunOptions fields and "run" options.
- The schema knows the new keys; the linter checks artifact patterns and
  the length of a parallel_run artifacts list.

No XML is parsed. A report that is malformed or cannot be written is a
warning on the result and never changes the outcome of the run.
SonarCloud's gate failed on two path findings from command line
arguments, and Codacy reported six new issues.

- Remove "schema -o FILE": the command only prints, and the output is
  redirected to save it.
- Keep reading the workflow that -e names, which is the purpose of the
  flag, and mark the line. _load_yaml no longer puts the content of a
  file that is not a mapping into its exception message.
- Record an in-process runner's outcome in a finally block.
- Build the XML character class of the JUnit writer from code point
  ranges; it had been saved as raw characters.
- Explain and mark the one place a runner process is started.
- Tests: exercise the entry point with runpy instead of a sub-process,
  split composite assertions and keep one call per raises block.
Phase 5 of the platform roadmap (PR #32). The quick start was three
isolated snippets. It is now one example from an empty folder to a CI
job, in the three READMEs and in the Getting Started page: project
layout, runner scripts, the workflow, validate, run, where every output
goes, a GitHub Actions job, editors and troubleshooting.

The example needs no browser and no external service: the workflow
serves a page with http.server, tests it with the API runner, then runs
the API and load runners in parallel. It was run as written with the
real runner packages.

- docs/migration.rst collects what changes for an existing project.
- test/test_readme_example.py extracts the example from each README and
  checks that the translations match and the workflow validates.
- The CI integration job uploads report/ and artifacts/.
Both branches added docs/updates/2026-10.md and changed its index, the
CI workflow and architecture.md. Each conflict keeps what both sides
added:

- docs/updates/2026-10.md: dev's five entries of 2026-10-01, then the
  five of 2026-10-08, in date order. No ID is used twice.
- docs/updates/README.md: the index lists all 23 entries, newest first,
  and the batch table counts 13 for 2026-09 and 10 for 2026-10.
- .github/workflows/ci.yml: the report upload step ends the integration
  job and the publish-dev job follows it.
- architecture.md: the workflow keys of the branch and the PyPI section
  of dev; MANIFEST.in joins the packaging files in section 8.
@JE-Chen
JE-Chen changed the base branch from main to dev October 8, 2026 04:54
download_single_file ignored what automation_file.download_file returned.
That function does not raise: it returns False for a URL it refuses and
for a transfer that fails. A step that downloaded nothing was therefore
reported as done and the workflow went on.

The step now fails, logs the URL, and stops the steps after it.
The four workflows under test/unit_test named scripts that could not be
found, needed a display, a browser and an external site, and their
scripts ignored the outcome, so the integration job passed without
executing anything.

They now serve a page with http.server and test it with the API, load
and file runners; the download sample fetches this repository's LICENSE.
Each test script exits 1 unless its run passed. The unit-test job
validates the four workflows, and the step that uploaded a video no
sample records is removed.
_validate_steps added every step name to a set that nothing emptied, so
a second execute_yaml call in the same process that reused a name
stopped with "job name duplicated" before running a step.

The set is emptied when a run validates its steps. Closing a program
looks the name up elsewhere, so a program an earlier call left open can
still be closed, and open_program now refuses a name that is still open
instead of replacing its entry.
The runner packages keep one list of records per process and write the
whole list into every report. run and run_folder call the runner inside
the TestPioneer process, so the report of a second step of the same
runner repeated the records of the first: tests were counted twice and
a failure of step 1 also failed step 2.

A report of an in-process call that starts with the records the same
runner reported last time now adds only what follows them. A report that
does not start with them means the script cleared the records, and all
of it counts. parallel_run entries are separate processes and are left
alone.
Nothing in test_pioneer imports it. It is removed from pyproject.toml
and dev.toml, from the dependency list in CLAUDE.md and from
architecture.md. uv.lock is regenerated with every other locked version
kept.

A mail step can declare the dependency again when one exists.
A runner that reads TEST_PIONEER_ARTIFACT_DIR writes a relative report
name below that directory, so the pattern a step declares no longer
matches below the working directory and the collector warned that it
matched nothing. The collector now looks for the pattern inside the
runner directory first, so one workflow works with runners of both
kinds.
The timestamp of a testsuite was the start time of the result as it is,
with milliseconds and a trailing Z. JUnit has a date-time to the second
without a zone, and readers built on datetime.fromisoformat before
Python 3.11 drop a value that ends in Z. The fraction and the Z are left
out, and a suite without a start time has no timestamp attribute.
The report is written after every run and artifacts are kept for a run
that did not pass. The entry says why this is not opt-in and how to turn
both off.
…patch

SonarCloud reported the page the samples serve for its missing doctype,
language and title, and a test for changing a shared mock by hand.
A dash in place of a file name reads the workflow from standard input,
so an editor can check a buffer that is not saved. The input is decoded
as UTF-8 whatever the console encoding is, and is reported under the
name <stdin>. progress.md says what phase 4 still waits on.
With TEST_PIONEER_ARTIFACT_DIR set, a runner writes relative output
names below that directory and exits 1 from --execute_file when an
action failed. TestPioneer already handles a runner that does. The
entry in progress.md says where that work belongs in the runner
repositories and why it is not started from here.
dataclasses.replace is typed as returning a generic dataclass instance,
which SonarCloud reported against the declared return type.
@JE-Chen JE-Chen changed the title [Draft] Plan UI, YAML validation, runner artifacts, and merged reports Add workflow validation, runner artifacts and merged reports Oct 8, 2026
@JE-Chen
JE-Chen marked this pull request as ready for review October 8, 2026 06:12
The development branches of the runners were read. LoadDensity already
exits non-zero after a failed action, and the other runners can count
failures through the reporter hook the shared executor has always had,
so no release of it is needed. The docs, architecture.md and the open
item in progress.md say so, and the item lists the steps per runner.
Native support is written in the five runner repositories, one draft
pull request each. The item now lists them and what changes here once
they are released.
@sonarqubecloud

sonarqubecloud Bot commented Oct 8, 2026

Copy link
Copy Markdown

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant