Skip to content

feat: Web console: run the e2e tests on embedded clusters, modernize them, and cover more of the console - #20544

Open
vogievetsky wants to merge 26 commits into
apache:masterfrom
vogievetsky:make_console_e2e_tests_embeded
Open

vogievetsky wants to merge 26 commits into
apache:masterfrom
vogievetsky:make_console_e2e_tests_embeded

Conversation

@vogievetsky

Copy link
Copy Markdown
Contributor

Note from the human

So the e2e tests in the web console have not been touched in a long time. On a whim I asked Claude "If I was added e2e tests today would they be written differently"... and boy did it have a lot to say.

While Claude was clauding I described what I was doing to @gianm and he said: "It would be cool if web console e2e tests used the embedded tests framework" and the rest is history...

Description

The web console's end-to-end tests ran in CI by building a full Druid distribution, starting it with script/druid, running the Playwright specs against it and stopping it. That was slow, the specs shared one long-lived cluster (so leftovers from one test could break another), and the console flows lived apart from the Druid features they exercise.

This PR runs the specs on the embedded clusters that the rest of Druid's tests use (embedded-tests), modernizes the specs themselves, and adds console coverage for features that the embedded tests already cover through the API. Each console test lives next to the embedded tests of its feature (for example s3/S3WebConsoleTest next to S3StorageTest), so the "load data from S3" flow is tested through the console too.

Running the console e2e specs on embedded clusters

  • WebConsoleTestBase (in embedded-tests, package org.apache.druid.testing.embedded.console) starts an embedded cluster with a Router, which serves the console from the web-console jar, and runs a Playwright spec of web-console/e2e-tests on it with npx playwright test <spec>. A subclass adds the resources the spec needs (S3 or Kafka containers, extensions, properties) in addResources / configureCluster and passes their settings to the spec as environment variables. The JUnit failure message carries Playwright's output; each spec writes its screenshots and traces to its own web-console/test-results/<spec>/.
  • The tests are tagged web-console: excluded from the default surefire run and run by failsafe with the new web-console-tests profile:
    ./mvnw -pl web-console install -DskipTests          # with the console built, for the Router to serve it
    ./mvnw -pl embedded-tests verify -Pweb-console-tests [-Dit.test=BatchIndexingWebConsoleTest]
  • -Dweb.console.keepAlive=true keeps a test's cluster up and prints the command to run its spec against it (with --ui / --debug as you like). -Dweb.console.port=18081 runs the specs against the dev server (npm start) instead of the bundled console.
  • Every test class gets a fresh cluster, so the specs no longer clean up after themselves there (DRUID_E2E_TEST_CLUSTER_IS_DISPOSABLE).
  • WebConsoleSpecsTest (a plain unit test) fails if a spec of web-console/e2e-tests isn't run by any *WebConsoleTest, or if one runs a spec that doesn't exist.
  • The specs still run against a quickstart cluster with npm run test-e2e (script/druid stays for that); the ones that need a resource (S3, Kafka, extra extensions, a lookup table) skip themselves there.

CI

.github/scripts/web-checks.sh installs web-console (with the console built) and then runs ./mvnw verify -pl 'embedded-tests,!web-console' -am -Pweb-console-tests -DskipUTs, which builds only the modules embedded-tests needs: no distribution, no extensions download, no separate JVMs to start. worker.yml uploads embedded-tests/target/failsafe-reports/ (the Druid and Playwright logs) and web-console/test-results/ instead of the distribution's logs, and the existing JUnit report step now shows each console test.

Modernized e2e specs

The specs were written in a Puppeteer style; they now use Playwright as intended:

  • The @playwright/test runner (playwright.config.ts) instead of Jest: fixtures instead of per-spec browser setup, a globalSetup that waits for the cluster, expect.poll / toPass instead of a home-made retry, screenshots, traces and an error-context.md on failure.
  • Locators (getByRole, form fields found by their label) instead of XPath and element handles.
  • The cluster state (tasks finishing, segments loading, supervisor and Dart query states) is checked through SQL on the system tables rather than by reading view columns by position; views are read by their column headers.
  • Test data is set up through the HTTP API rather than by shelling out to post-index-task.
  • Page objects are plain functions and types (no Object.assign data classes), with loadData(page, config) for the data loader.
  • E2E_SHOW_STEPS=true saves a screenshot of each step a test takes to e2e-tests/steps/<spec>-NNN.png, a quick way to see what a test does.
  • web-console/e2e-tests/README.md describes the tests, how to run them both ways, and how to write one.

Running them on embedded clusters also surfaced three latent races in the specs that the quickstart's timing had hidden (reading a view's table before it loaded, cancelling a query before it was sent, reading the Query view's previous result), now fixed.

New console coverage

Spec Run by What it covers
s3-ingestion.spec.ts s3.S3WebConsoleTest Data loader from S3 (an S3 container that is also deep storage)
input-formats.spec.ts indexer.InputFormatsWebConsoleTest Data loader with CSV, TSV, Parquet, ORC and Avro OCF files: the format it picks and the columns it parses
kafka-ingestion.spec.ts indexing.KafkaWebConsoleTest A Kafka supervisor through the data loader (a Kafka container), then suspend, resume and terminate from the Supervisors view
sql-ingestion.spec.ts msq.SqlIngestionWebConsoleTest The SQL data loader, and REPLACE from external data and from a datasource in the Query view
datasource-actions.spec.ts indexing.DatasourceActionsWebConsoleTest Mark a datasource's segments unused and used again, then delete them with a kill task
retention-rules.spec.ts server.RetentionRulesWebConsoleTest Editing retention rules (loadForever, then dropForever)
jdbc-lookup.spec.ts lookup.JdbcLookupWebConsoleTest Initializing lookups and adding a JDBC lookup (on a table in the Derby metadata store), then querying it
dart.spec.ts msq.DartWebConsoleTest Running a query with the Dart engine, the current Dart queries panel and its details, cancelling a running Dart query

The existing specs (tutorial-batch, reindexing, auto-compaction, multi-stage-query, cancel-query) run from indexing.BatchIndexingWebConsoleTest, compact.AutoCompactionWebConsoleTest, msq.MultiStageQueryWebConsoleTest and query.SqlQueryCancelWebConsoleTest.

Web console fixes found on the way

  • Table filters in the URL (#datasources/datasource=...) couldn't hold values with # or ? (which end the hash route) or | (which separates values), so going from a datasource with one of those in its name to its tasks or segments filtered on the wrong thing. #, ?, &, / and % are now encoded, and \| / \\ escape | / \ in a value.
  • Api.encodePath left ^ and | as is, which Jetty rejects (400 Illegal Path Character), so every action on a datasource with one of those in its name (mark unused, kill, retention rules, compaction config) failed.

Release note

Fixed the web console's handling of datasource names containing #, ?, |, ^ or \: links between views now filter on the right datasource, and actions such as marking segments unused, killing data, and editing retention rules or compaction configs no longer fail for such names.


Key changed/added classes in this PR
  • WebConsoleTestBase, WebConsoleSpecsTest and the *WebConsoleTest classes (embedded-tests)
  • embedded-tests/pom.xml (web-console-tests profile), .github/scripts/web-checks.sh, .github/workflows/worker.yml
  • web-console/e2e-tests/**, web-console/playwright.config.ts
  • TableFilter, TableFilters, Api.encodePath (web console)

This PR has:

  • been self-reviewed.
  • added documentation for new or modified features or behaviors (web-console/e2e-tests/README.md).
  • a release note entry in the PR description.
  • added Javadocs for most classes and all non-trivial methods. Linked related entities via Javadoc links.
  • added comments explaining the "why" and the intent of the code wherever would not be obvious for an unfamiliar reader.
  • added unit tests or modified existing tests to cover new code paths (table-filter.spec.ts, table-filters.spec.ts, hash-routing.spec.ts, api.spec.ts).
  • added integration tests.

vogievetsky and others added 24 commits October 9, 2026 10:12
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- run the e2e tests with @playwright/test (playwright.config.ts) instead of Jest and the
  playwright-chromium library: the page fixture replaces the per-spec browser setup, a global
  setup waits for the console, failed tests keep a screenshot and a trace, and expect().toPass()
  replaces the retry helpers
- use locators (getByRole, form groups found by their label) instead of XPath, page.$ and
  waitForSelector
- open the Datasources and Tasks views filtered to the test's datasource, so the tests work on a
  cluster with more than a page of datasources or tasks
- fail a query attempt as soon as it errors, so that it can be retried while the Broker catches up

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- wait for tasks to finish and segments to load by querying sys.tasks and sys.segments
  (util/sql.ts) instead of reading the Tasks and Datasources views by column position
- read the console's tables by their column headers (extractTableRecords), and use it for the
  remaining checks of the Datasources and Tasks views in the tutorial test

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A filter on a value with a #, ? or | in it (like a link from a datasource to its tasks or
segments) filtered on the wrong thing:
- TableFilters.toString() did not encode # and ?, which end the hash route, so the rest of the
  value was dropped; now they are encoded (and decoded) like & % and /
- | separates a filter's values and could not be escaped, so a value with a | was split into
  several; now \| and \\ stand for a literal | and \ in a value, and TableFilters.eq and the
  filter menu of a table cell filter on their value as one value

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Now that table filters in the URL can hold any value, open the Datasources and Tasks views
filtered on the test's full datasource name, rather than on the part of it that a filter could
hold.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rather than run sed and examples/bin/post-index-task (bash and Python, with the Coordinator
hard-coded on :8081), read the tutorial ingestion spec, set the datasource name (and interval) on
it, POST it through the console's service with the request fixture, and poll the task's status,
failing with the task's error if it fails. The spec's input now points at this checkout's tutorial
directory rather than one relative to where Druid runs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Api.encodePath left ^ and | as they are, which the browser sends as is and Jetty rejects with
"400 Illegal Path Character", so every action on a datasource with one of those in its name
(mark unused, kill, retention rules, compaction config) failed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The newDatasourceName fixture makes a unique datasource name and, after the test passes, shuts
down the datasource's tasks, deletes its compaction config, marks its segments unused and kills
them. After a failure the datasource is kept to look into, and the log says which.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- merge component/query into component/workbench (both drove the Query view)
- plain object types for data (rows read from views, what to set in forms) instead of classes
  using Object.assign and interface merging, so the no-unsafe-declaration-merging eslint override
  goes; union types with functions for the partitions specs and data connectors
- the data loader is a loadData(page, config) function with one flat config instead of a class
  with a config object per step

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Adds WebConsoleTestBase and CoreWebConsoleTest to embedded-tests (tag
web-console, profile web-console-tests), which start an embedded cluster
with a Router and run the Playwright specs on it, with a keep-alive mode
for working on a spec. Fixes two races in the specs that this surfaced:
reading a view's table before it loaded, and canceling a query before
it was sent.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
S3WebConsoleTest runs the new s3-ingestion spec on an embedded cluster
with an S3 container (also its deep storage and task log storage),
loading the tutorial file through the data loader's Amazon S3 connector.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
web-checks.sh installs web-console with the console built, then runs the
web-console-tests profile of embedded-tests (building only the modules
it needs) instead of building a distribution and starting it with
script/druid. worker.yml uploads the failsafe reports and the Playwright
results in place of the distribution's logs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nctionality

Splits CoreWebConsoleTest into a *WebConsoleTest per functionality
(indexing, compact, msq, query) and moves S3WebConsoleTest to the s3
package. WebConsoleTestBase stays in console, as shared infrastructure.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…y spec is run

The specs no longer delete their datasources on an embedded cluster, as
it's thrown away after the test class. WebConsoleSpecsTest checks that
every spec of web-console/e2e-tests is run by a *WebConsoleTest.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
showStep(page, title) saves a screenshot to e2e-tests/steps/<spec>-NNN.png
(git ignored) when E2E_SHOW_STEPS=true, and is called by the page objects
at each step of the data loader, each view read, query results and
dialogs, and by a fixture as the test ends. setQueryInput closes the
autocomplete that typing opens, as it covered the query.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
input-formats.spec.ts loads CSV, TSV, Parquet, ORC and Avro OCF files
through the data loader, checking the input format it picks and the
columns it parses, run by InputFormatsWebConsoleTest (with the Avro,
Parquet and ORC extensions) next to ITLocalInputSourceAllInputFormatTest.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
kafka-ingestion.spec.ts sets up a Kafka supervisor through the data
loader (as in the Kafka tutorial), then suspends, resumes and terminates
it from the Supervisors view, run by KafkaWebConsoleTest with a Kafka
container next to the other Kafka embedded tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
sql-ingestion.spec.ts loads data through the SQL data loader, then
ingests and reindexes with REPLACE in the Query view, run by
SqlIngestionWebConsoleTest next to the MSQ ingestion embedded tests.
WorkbenchOverview gets runIngestQuery, and waits for the query's own
request before its result, as the view shows the last result until then.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
datasource-actions.spec.ts marks a datasource's segments unused, used
again, then deletes them with a kill task from the Datasources view, run
by DatasourceActionsWebConsoleTest. DatasourcesOverview can show unused
datasources and run and confirm a datasource's actions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
retention-rules.spec.ts edits a datasource's retention rules from the
Datasources view (loadForever, then dropForever, which drops the data),
run by RetentionRulesWebConsoleTest. jdbc-lookup.spec.ts initializes the
lookups and adds a JDBC lookup from the Lookups view, then queries it,
run by JdbcLookupWebConsoleTest (with the table in the Derby metadata
store, as JdbcLookupTest does).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
dart.spec.ts runs a query with the Dart engine picked in the Query view,
follows it in the current Dart queries panel (with its details), and
cancels a running Dart query from the panel, run by DartWebConsoleTest.
The embedded clusters keep the reports of finished Dart queries, as
script/druid build does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… header fails

Addresses the CodeQL alert (potential input resource leak): the
FileInputStream was created inside the GZIPInputStream wrapper, whose
constructor reads the header and can throw before the try-with-resources
has a resource to close.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… to see an e2e test fail in CI

Temporary, to be reverted: the SQL data loader test of sql-ingestion.spec.ts
(run by SqlIngestionWebConsoleTest) clicks 'Start loading data', which this
renames, so it should time out looking for the button.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@FrankChen021 FrankChen021 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The current head cannot pass the new SQL ingestion E2E because the helper searches for a button label the application does not render. Additional issues can produce stale query assertions, indefinite E2E hangs, and incorrect filtering for datasource names containing |.

Reviewed 79 of 79 changed files.

Severity Findings
P0 0
P1 2
P2 2
P3 0
Total 4

This is an automated review by Codex GPT-5.6-Luna(max)

After addressing the findings or replying to the comments, you can request another review from me to trigger a new automated review.

await clickButton(destinationDialog, 'Save');
await destinationDialog.waitFor({ state: 'detached' });
await showStep(page, 'SQL data loader: Schema');
await clickButton(schemaStep.locator('.prev-next-bar'), 'Start loading data');

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Use the SQL loader's rendered submit label

Finding: The schema step renders the exact button name Start ingesting data, but this helper still searches for Start loading data. Because clickButton uses an exact accessible-name match, sql-ingestion.spec.ts times out before starting ingestion, so SqlIngestionWebConsoleTest cannot pass at this head.

Suggestion: Update the helper to the rendered label, or revert the temporary UI rename before merging.

await this.submit(query, options);

const error = this.page.locator('.execution-error-pane');
await done.or(error).waitFor({ timeout: QUERY_TIMEOUT });

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Wait for the current ingestion result

Finding: The workbench keeps the previous successful execution visible while a new query loads, so done.or(error).waitFor() can resolve on the old .ingest-success-pane. The second runIngestQuery in sql-ingestion.spec.ts can therefore return the first query's text before the new task finishes.

Suggestion: Wait for the current execution to complete or verify a query-specific state transition before reading the result pane.

new InputStreamReader(process.getInputStream(), StandardCharsets.UTF_8)
)) {
String line;
while ((line = reader.readLine()) != null) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Enforce the spec timeout while consuming output

Finding: The runner drains process.getInputStream() to EOF before calling waitFor with SPEC_TIMEOUT_MINUTES. A hung Playwright process can keep stdout open, so the timeout branch is never reached and CI can hang until an external timeout.

Suggestion: Consume output asynchronously while waiting, then destroy the process and join or close the reader when the timeout expires.

this.key = key;
this.mode = mode;
this.values = typeof value === 'string' ? value.split('|') : value;
this.values = typeof value === 'string' ? TableFilter.splitNeedle(value) : value;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Preserve literal pipes in datasource filters

Finding: String values are now parsed with splitNeedle, where an unescaped | means OR. Existing datasource links in segments-view.tsx and datasources-view.tsx still pass datasource names as strings, so a valid datasource named a|b becomes two filter values instead of one literal datasource.

Suggestion: Use the constructor's literal-value form for those datasource links and add a pipe-containing datasource regression case.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants