Skip to content

[WIP] Map State TTL - #3

Closed
ericm-db wants to merge 496 commits into
masterfrom
list-state-ttl
Closed

ericm-db wants to merge 496 commits into
masterfrom
list-state-ttl

Conversation

@ericm-db

@ericm-db ericm-db commented Apr 1, 2024

Copy link
Copy Markdown
Owner

What changes were proposed in this pull request?

Why are the changes needed?

Does this PR introduce any user-facing change?

How was this patch tested?

Was this patch authored or co-authored using generative AI tooling?

yaooqinn and others added 30 commits March 13, 2024 20:29
…ct#getCatalystType`

### What changes were proposed in this pull request?

This PR adds guidelines for mapping database timestamps to Spark SQL Timestamps through the JDBC Standard API and Spark JDBCDialects trait. The details of this PR can be viewed directly in the method descriptions, no more copying here

### Why are the changes needed?

These guidelines can help us revise the built-in jdbc datasource later without controversies. It also encourages custom dialects to follow it。

### Does this PR introduce _any_ user-facing change?

no, developer API doc changes

### How was this patch tested?

doc build

### Was this patch authored or co-authored using generative AI tooling?
no

Closes apache#45496 from yaooqinn/SPARK-47375.

Authored-by: Kent Yao <yao@apache.org>
Signed-off-by: Kent Yao <yao@apache.org>
…in IDE

### What changes were proposed in this pull request?
The pr aims to make the related Protobuf `UT` run well in IDE (IntelliJ IDEA).

### Why are the changes needed?
Facilitate developers to debug the related Protobuf `UT`.

Before:
<img width="1279" alt="image" src="https://github.com/apache/spark/assets/15246973/c00781b2-3477-4b2c-b871-ead997fda697">

After:
<img width="884" alt="image" src="https://github.com/apache/spark/assets/15246973/665fc67d-c69e-45c7-b37d-bb4ef8e72930">

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
- Manually test.
- Pass GA.

### Was this patch authored or co-authored using generative AI tooling?
No.

Closes apache#45498 from panbingkun/SPARK-47378.

Authored-by: panbingkun <panbingkun@baidu.com>
Signed-off-by: Dongjoon Hyun <dhyun@apple.com>
…an unquoted identifier and fix "IS ! NULL" et al.

### What changes were proposed in this pull request?

In this PR we propose to extend the lexing of IDENTIFIER beyond what is legitimate for unquoted identifiers to include
"plausible" identifiers. We then use the "exit" hook in the parser raise INVALID_IDENTIFIER error which is more meaningful than a syntax error.

Specifically we allow:
* general letters beyond the ASCII a-z. This will catch locale specific names
* URIs which are used for table's represented by a path.

As part of this PR we also found that rolling `NOT` and `!` into one token is a "bad idea".
We allow:

CREATE TABLE t(c1 INT ! NULL); etc.
This is clearly not intended.

! is now ONLY allowed as a boolean prefix operator.

### Why are the changes needed?

This change  improves the user experience in case of an error by returning a more meaningful error.

### Does this PR introduce _any_ user-facing change?

No

### How was this patch tested?

Existing test suite + new unit tests

### Was this patch authored or co-authored using generative AI tooling?

No

Closes apache#45470 from srielau/SPARK-47344.

Lead-authored-by: Serge Rielau <serge@rielau.com>
Co-authored-by: Wenchen Fan <cloud0fan@gmail.com>
Signed-off-by: Gengliang Wang <gengliang@apache.org>
…with transformWithState operator

### What changes were proposed in this pull request?
Add support for processing/event time based timers with `transformWithState` operator

### Why are the changes needed?
Changes are required to add event-driven timer based support for stateful streaming applications based on arbitrary state  API with the `transformWithState` operator

As part of this change - we introduce a bunch of functions that users can use within the `StatefulProcessor` logic. Using the `StatefulProcessorHandle`, users can do the following:
- register timer at a given timestamp
- delete timer at a given timestamp
- list timers

Note that all the above operations are tied to the implicit grouping key.

In terms of the implementation, we make use of additional column families to support the operations mentioned above. For registered timers, we maintain a primary index (as a col family) that keeps the mapping between the grouping key and expiry timestamp. This col family is used to add and delete timers with direct access to the key and also for listing registered timers for a given grouping key using `prefix scan`. We also maintain a secondary index that inverts the ordering of the timestamp and grouping key. We will incorporate the use of the range scan encoder for this col family in a separate PR.

Few additional constraints:
- only registered timers are tracked and occupy storage (locally and remotely)
- col families starting with `_` are reserved and cannot be used as state variables
- timers are checkpointed as before
- users have to provide a `timeoutMode` to the operator. Currently, they can choose to not register timeouts or register timeouts that are processing-time based or event-time based. However, this mode has to be declared upfront within the operator arguments.

### Does this PR introduce _any_ user-facing change?
Yes

### How was this patch tested?
Added unit tests as well as pseudo-integration tests

StatefulProcessorHandleSuite
```
13:58:42.463 WARN org.apache.spark.sql.execution.streaming.state.StatefulProcessorHandleSuite:

===== POSSIBLE THREAD LEAK IN SUITE o.a.s.sql.execution.streaming.state.StatefulProcessorHandleSuite, threads: rpc-boss-3-1 (daemon=true), shuffle-boss-6-1 (daemon=true) =====
[info] Run completed in 4 seconds, 559 milliseconds.
[info] Total number of tests run: 8
[info] Suites: completed 1, aborted 0
[info] Tests: succeeded 8, failed 0, canceled 0, ignored 0, pending 0
[info] All tests passed.
```

TransformWithStateSuite
```
13:48:41.858 WARN org.apache.spark.sql.streaming.TransformWithStateSuite:

===== POSSIBLE THREAD LEAK IN SUITE o.a.s.sql.streaming.TransformWithStateSuite, threads: QueryStageCreator-0 (daemon=true), state-store-maintenance-thread-0 (daemon=true), ForkJoinPool.commonPool-worker-4 (daemon=true), state-store-maintenance-thread-1 (daemon=true), QueryStageCreator-1 (daemon=true), rpc-boss-3-1 (daemon=true), F
orkJoinPool.commonPool-worker-3 (daemon=true), QueryStageCreator-2 (daemon=true), QueryStageCreator-3 (daemon=true), state-store-maintenance-task (daemon=true), ForkJoinPool.com...
[info] Run completed in 1 minute, 32 seconds.
[info] Total number of tests run: 20
[info] Suites: completed 1, aborted 0
[info] Tests: succeeded 20, failed 0, canceled 0, ignored 0, pending 0
[info] All tests passed.
```

### Was this patch authored or co-authored using generative AI tooling?
No

Closes apache#45051 from anishshri-db/task/SPARK-46913.

Authored-by: Anish Shrigondekar <anish.shrigondekar@databricks.com>
Signed-off-by: Jungtaek Lim <kabhwan.opensource@gmail.com>
…n `test_column_arithmetic_ops`

### What changes were proposed in this pull request?
Enable column name comparsion in `test_column_arithmetic_ops`

### Why are the changes needed?
the default column name should had been already fixed

### Does this PR introduce _any_ user-facing change?
no, test-only

### How was this patch tested?
ci

### Was this patch authored or co-authored using generative AI tooling?
no

Closes apache#45493 from zhengruifeng/test_column_negative_name.

Authored-by: Ruifeng Zheng <ruifengz@apache.org>
Signed-off-by: Ruifeng Zheng <ruifengz@apache.org>
…scription in JDBC doc

### What changes were proposed in this pull request?

Correct the preferTimestampNTZ option description in JDBC doc as per apache#45496

### Why are the changes needed?

The current doc is wrong about the jdbc option preferTimestampNTZ

### Does this PR introduce _any_ user-facing change?

No

### How was this patch tested?

Just doc change
### Was this patch authored or co-authored using generative AI tooling?

No

Closes apache#45502 from gengliangwang/ntzJdbc.

Authored-by: Gengliang Wang <gengliang@apache.org>
Signed-off-by: Gengliang Wang <gengliang@apache.org>
…nectSQLTestCase`

### What changes were proposed in this pull request?
Factor out tests from `SparkConnectSQLTestCase`

### Why are the changes needed?
for testing parallelism

### Does this PR introduce _any_ user-facing change?
no, test only

### How was this patch tested?
ci

### Was this patch authored or co-authored using generative AI tooling?
no

Closes apache#45497 from zhengruifeng/breaK_spark_connect_basic.

Authored-by: Ruifeng Zheng <ruifengz@apache.org>
Signed-off-by: Hyukjin Kwon <gurwls223@apache.org>
…yml`

### What changes were proposed in this pull request?
The pr aims to update the link `text` and `ref` in `notify_test_workflow.yml`.

### Why are the changes needed?
1.The keyword `Disabling or limiting GitHub Actions for a repository` does not exist on the page.
  <img width="1027" alt="image" src="https://github.com/apache/spark/assets/15246973/392320de-53b9-4b08-9956-afd2bb3708f9">
  If a new user encounters the above issue, `she/he` will generally click on this page to search for `Disabling or limiting GitHub Actions for a repository`  But unfortunately, the content of this page has changed and the keyword `Disabling or limiting GitHub Actions for a repository` no longer exists.  I propose `updating` it to reduce user confusion.

2.The url [`https://docs.github.com/en/github/administering-a-repository/disabling-or-limiting-github-actions-for-a-repository`](https://docs.github.com/en/github/administering-a-repository/disabling-or-limiting-github-actions-for-a-repository) has been redirected to [`https://docs.github.com/en/repositories/managing-your-repositorys-settings-and-features/enabling-features-for-your-repository/managing-github-actions-settings-for-a-repository`](https://docs.github.com/en/repositories/managing-your-repositorys-settings-and-features/enabling-features-for-your-repository/managing-github-actions-settings-for-a-repository).

### Does this PR introduce _any_ user-facing change?
Yes, only for dev.

### How was this patch tested?
Manually test.

### Was this patch authored or co-authored using generative AI tooling?
No.

Closes apache#45495 from panbingkun/minor_notify_test_workflow.

Authored-by: panbingkun <panbingkun@baidu.com>
Signed-off-by: Hyukjin Kwon <gurwls223@apache.org>
…link`

### What changes were proposed in this pull request?
The pr aims to:
- fix connect-repl `usage prompt`.
- fix docs link.

### Why are the changes needed?
Only fix bug.

- Usage prompt
1.update `enable_ssl` to `use_ssl`.
2.add `user_agent` and `session_id`
  Before:
  <img width="1025" alt="image" src="https://github.com/apache/spark/assets/15246973/db401e7f-9569-469c-90eb-17c277472b4d">

  After:
  <img width="1025" alt="image" src="https://github.com/apache/spark/assets/15246973/de7b1f8c-1d0d-4696-8f93-f72636b0d1a0">

- Docs link
  The url `https://github.com/apache/spark/blob/master/connector/connect/client/jvm/src/main/scala/org/apache/spark/sql/connect/client/SparkConnectClientParser.scala#L48` does not exist.
  Before:
  <img width="1324" alt="image" src="https://github.com/apache/spark/assets/15246973/a1dd0578-cba1-4cfe-8330-fcd48c84ca69">

  After:
  <img width="1222" alt="image" src="https://github.com/apache/spark/assets/15246973/a66467e3-652f-4288-888d-d676b3365569">

### Does this PR introduce _any_ user-facing change?
Yes, the `connect-repl` `usage prompt` and `docs link` have been corrected.

### How was this patch tested?
- Manually test.
- Pass GA.

### Was this patch authored or co-authored using generative AI tooling?
No.

Closes apache#45494 from panbingkun/SPARK-47374.

Authored-by: panbingkun <panbingkun@baidu.com>
Signed-off-by: Hyukjin Kwon <gurwls223@apache.org>
## What changes were proposed in this pull request?

apache#40755  adds a null check on the input of the child deserializer in the tuple encoder. It breaks the deserializer for the `Option` type, because null should be deserialized into `None` rather than null. This PR adds a boolean parameter to `ExpressionEncoder.tuple` so that only the user that apache#40755 intended to fix has this null check.

## How was this patch tested?

Unit test.

Closes apache#45508 from chenhao-db/SPARK-47385.

Authored-by: Chenhao Li <chenhao.li@databricks.com>
Signed-off-by: Wenchen Fan <wenchen@databricks.com>
…ean up

### What changes were proposed in this pull request?
Protobuf `abbreviate` method code clean up

### Why are the changes needed?
code clean up, to make it more easy to reuse

### Does this PR introduce _any_ user-facing change?
no

### How was this patch tested?
ci

### Was this patch authored or co-authored using generative AI tooling?
no

Closes apache#45512 from zhengruifeng/connect_abbr_cleanup.

Authored-by: Ruifeng Zheng <ruifengz@apache.org>
Signed-off-by: Ruifeng Zheng <ruifengz@apache.org>
…TAMP_TZ

### What changes were proposed in this pull request?

Regarding SPARK-47375, this PR fixes the issue of converting PG TIMESTAMP_TZ & TIME_TZ to our NTZ.

### Why are the changes needed?

bugfix

### Does this PR introduce _any_ user-facing change?

yes, as 3.5 is out, this PR add a migration guide for this.

### How was this patch tested?

new tests

### Was this patch authored or co-authored using generative AI tooling?

mo

Closes apache#45513 from yaooqinn/SPARK-47390.

Authored-by: Kent Yao <yao@apache.org>
Signed-off-by: Kent Yao <yao@apache.org>
…TZ option doc

### What changes were proposed in this pull request?

Fix a mistake in JDBC's preferTimestampNTZ option doc

### Why are the changes needed?

Fix a mistake in doc

### Does this PR introduce _any_ user-facing change?

No
### How was this patch tested?

Just doc change
### Was this patch authored or co-authored using generative AI tooling?

No

Closes apache#45510 from gengliangwang/reviseJdbcDoc.

Authored-by: Gengliang Wang <gengliang@apache.org>
Signed-off-by: Kent Yao <yao@apache.org>
### What changes were proposed in this pull request?
We can already select the desired overhead memory directly via the `spark.driver/executor.memoryOverhead` flags, however, if that flag is not present the overhead memory calculation goes as follows:

```
overhead_memory = Max(384, 'spark.driver/executor.memory' * 'spark.driver/executor.memoryOverheadFactor')

[where the 'memoryOverheadFactor' flag defaults to 0.1]
```

This PR adds two new spark configs: `spark.driver.minMemoryOverhead` and `spark.executor.minMemoryOverhead`, which can be used to override the 384Mib minimum value.

The memory overhead calculation will now be :

```
min_memory = sparkConf.get('spark.driver/executor.minMemoryOverhead').getOrElse(384)

overhead_memory = Max(min_memory, 'spark.driver/executor.memory' * 'spark.driver/executor.memoryOverheadFactor')
```

### Why are the changes needed?
There are certain times where being able to override the 384Mb minimum directly can be beneficial. We may have a scenario where a lot of off-heap operations are performed (ex: using package managers/native compression/decompression) where we don't have a need for a large JVM heap but we may still need a signficant amount of memory in the spark node.

Using the `memoryOverheadFactor` config flag may not prove appropriate, since we may not want the overhead allocation to directly scale with JVM memory, as a cost saving/resource limitation problem.

### Does this PR introduce _any_ user-facing change?
Yes, as described above, two new flags have been added to the spark config. No break of existing behaviours.

### How was this patch tested?
Added tests for 3 cases:
- If `spark.driver/executor.memoryOverhead` is set, then the new changes have no effect.
- If  `spark.driver/executor.minMemoryOverhead` is set and its value is higher than  'spark.driver/executor.memory' * 'spark.driver/executor.memoryOverheadFactor', the total memory will be the allocated JVM memory + `spark.driver/executor.minMemoryOverhead`
- If  `spark.driver/executor.minMemoryOverhead` but its value is lower than 'spark.driver/executor.memory' * 'spark.driver/executor.memoryOverheadFactor', the total memory will be the allocated JVM memory + 'spark.driver/executor.memory' * 'spark.driver/executor.memoryOverheadFactor'.

### Was this patch authored or co-authored using generative AI tooling?
No

Closes apache#45240 from jpcorreia99/jcorrreia/MinOverheadMemoryOverride.

Authored-by: jpcorreia99 <jpcorreia99@gmail.com>
Signed-off-by: Thomas Graves <tgraves@apache.org>
### What changes were proposed in this pull request?
In the PR, I propose to pass `messageParameters` by name to avoid eager instantiation.

### Why are the changes needed?
Passing `messageParameters` by value independently from `requirement` might introduce perf regression.

### Does this PR introduce _any_ user-facing change?
No, this is not a part of public API.

### How was this patch tested?
By running the affected test suite:
```
$ build/sbt "test:testOnly *QueryCompilationErrorsSuite"
```

### Was this patch authored or co-authored using generative AI tooling?
No.

Closes apache#45511 from MaxGekk/fix-SparkException-require.

Authored-by: Max Gekk <max.gekk@gmail.com>
Signed-off-by: Max Gekk <max.gekk@gmail.com>
### What changes were proposed in this pull request?

Following the guidelines of SPARK-47375, this PR supports TIMESTAMP WITH TIME ZONE for H2Dialect and maps it to TimestampType regardless of the option `preferTimestampNTZ`

https://www.h2database.com/html/datatypes.html#timestamp_with_time_zone_type

### Why are the changes needed?

H2Dialect improvement, we currently don't have a default mapping for `java.sql.Types.TIME_WITH_TIMEZONE, TIMESTAMP_WITH_TIMEZONE`

### Does this PR introduce _any_ user-facing change?

no

### How was this patch tested?

new tests

### Was this patch authored or co-authored using generative AI tooling?
no

Closes apache#45516 from yaooqinn/SPARK-47394.

Authored-by: Kent Yao <yao@apache.org>
Signed-off-by: Dongjoon Hyun <dhyun@apple.com>
### What changes were proposed in this pull request?
This PR removes the legacy test case for JDK 8 added at SPARK-34607.

### Why are the changes needed?
apache#43080 already removed `isMemberClass` with `class.isMemberClass`.
In fact, the test case doesn't need any more.

On the other hand, this test case fails in Windows operation system.
```
Internal error (java.io.FileNotFoundException): D:\Users\gja\git-forks\spark\sql\catalyst\target\scala-2.13\test-classes\org\apache\spark\sql\catalyst\encoders\ExpressionEncoderSuite$OuterLevelWithVeryVeryVeryLongClassName1$OuterLevelWithVeryVeryVeryLongClassName2$OuterLevelWithVeryVeryVeryLongClassName3$OuterLevelWithVeryVeryVeryLongClassName4$OuterLevelWithVeryVeryVeryLongClassName5$OuterLevelWithVeryVeryVeryLongClassName6$.class (文件名、目录名或卷标语法不正确。)
java.io.FileNotFoundException: D:\Users\gja\git-forks\spark\sql\catalyst\target\scala-2.13\test-classes\org\apache\spark\sql\catalyst\encoders\ExpressionEncoderSuite$OuterLevelWithVeryVeryVeryLongClassName1$OuterLevelWithVeryVeryVeryLongClassName2$OuterLevelWithVeryVeryVeryLongClassName3$OuterLevelWithVeryVeryVeryLongClassName4$OuterLevelWithVeryVeryVeryLongClassName5$OuterLevelWithVeryVeryVeryLongClassName6$.class (文件名、目录名或卷标语法不正确。)
	at java.base/java.io.FileInputStream.open0(Native Method)
	at java.base/java.io.FileInputStream.open(FileInputStream.java:216)
	at java.base/java.io.FileInputStream.<init>(FileInputStream.java:157)
	at com.intellij.openapi.util.io.FileUtil.loadFileBytes(FileUtil.java:211)
	at org.jetbrains.jps.incremental.scala.local.LazyCompiledClass.$anonfun$getContent$1(LazyCompiledClass.scala:18)
	at scala.Option.getOrElse(Option.scala:201)
	at org.jetbrains.jps.incremental.scala.local.LazyCompiledClass.getContent(LazyCompiledClass.scala:17)
	at org.jetbrains.jps.incremental.instrumentation.BaseInstrumentingBuilder.performBuild(BaseInstrumentingBuilder.java:38)
	at org.jetbrains.jps.incremental.instrumentation.ClassProcessingBuilder.build(ClassProcessingBuilder.java:80)
	at org.jetbrains.jps.incremental.IncProjectBuilder.runModuleLevelBuilders(IncProjectBuilder.java:1569)
	at org.jetbrains.jps.incremental.IncProjectBuilder.runBuildersForChunk(IncProjectBuilder.java:1198)
	at org.jetbrains.jps.incremental.IncProjectBuilder.buildTargetsChunk(IncProjectBuilder.java:1349)
	at org.jetbrains.jps.incremental.IncProjectBuilder.buildChunkIfAffected(IncProjectBuilder.java:1163)
	at org.jetbrains.jps.incremental.IncProjectBuilder$BuildParallelizer$1.run(IncProjectBuilder.java:1129)
	at com.intellij.util.concurrency.BoundedTaskExecutor.doRun(BoundedTaskExecutor.java:244)
	at com.intellij.util.concurrency.BoundedTaskExecutor.access$200(BoundedTaskExecutor.java:30)
	at com.intellij.util.concurrency.BoundedTaskExecutor$1.executeFirstTaskAndHelpQueue(BoundedTaskExecutor.java:222)
	at com.intellij.util.ConcurrencyUtil.runUnderThreadName(ConcurrencyUtil.java:218)
	at com.intellij.util.concurrency.BoundedTaskExecutor$1.run(BoundedTaskExecutor.java:210)
	at java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136)
	at java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635)
	at java.base/java.lang.Thread.run(Thread.java:842)
```

### Does this PR introduce _any_ user-facing change?
'No'.

### How was this patch tested?
N/A

### Was this patch authored or co-authored using generative AI tooling?
'No'.

Closes apache#45514 from beliefer/remove-unnecessary-test.

Authored-by: beliefer <beliefer@163.com>
Signed-off-by: Dongjoon Hyun <dhyun@apple.com>
### What changes were proposed in this pull request?
The pr aims to upgrade `RoaringBitmap` from `1.0.1` to `1.0.5`.

### Why are the changes needed?
Release notes: https://github.com/RoaringBitmap/RoaringBitmap/releases/tag/1.0.5
This version includes some bugs fixed, eg:
- fix roaringbitmap - batchiterator's advanceIfNeeded to handle run lengths of zero by
- fix RangeBitmap#between bug in full section after empty section

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
Pass GA.

### Was this patch authored or co-authored using generative AI tooling?
No.

Closes apache#45507 from panbingkun/SPARK-47384.

Authored-by: panbingkun <panbingkun@baidu.com>
Signed-off-by: Dongjoon Hyun <dhyun@apple.com>
### What changes were proposed in this pull request?

This PR aims to upgrade `gas-connector` to 2.2.20.

### Why are the changes needed?

To bring the latest updates.
- https://github.com/GoogleCloudDataproc/hadoop-connectors/releases/tag/v2.2.20
    - Add support for renaming folders using rename backend API for Hierarchical namespace buckets
    - Upgrade java-storage to 2.32.1 and upgrade the version of related dependencies
- https://github.com/GoogleCloudDataproc/hadoop-connectors/releases/tag/v2.2.10

### Does this PR introduce _any_ user-facing change?

No.

### How was this patch tested?

```
$ dev/make-distribution.sh -Phadoop-cloud
$ cd dist
$ export KEYFILE=~/.ssh/apache-spark.json
$ export EMAIL=$(jq -r '.client_email' < $KEYFILE)
$ export PRIVATE_KEY_ID=$(jq -r '.private_key_id' < $KEYFILE)
$ export PRIVATE_KEY="$(jq -r '.private_key' < $KEYFILE)"
$ bin/spark-shell \
        -c spark.hadoop.fs.gs.auth.service.account.email=$EMAIL \
        -c spark.hadoop.fs.gs.auth.service.account.private.key.id=$PRIVATE_KEY_ID \
        -c spark.hadoop.fs.gs.auth.service.account.private.key="$PRIVATE_KEY"
Setting default log level to "WARN".
To adjust logging level use sc.setLogLevel(newLevel). For SparkR, use setLogLevel(newLevel).
Welcome to
      ____              __
     / __/__  ___ _____/ /__
    _\ \/ _ \/ _ `/ __/  '_/
   /___/ .__/\_,_/_/ /_/\_\   version 4.0.0-SNAPSHOT
      /_/

Using Scala version 2.13.12 (OpenJDK 64-Bit Server VM, Java 21.0.2)
Type in expressions to have them evaluated.
Type :help for more information.
24/03/14 09:33:41 WARN NativeCodeLoader: Unable to load native-hadoop library for your platform... using builtin-java classes where applicable
Spark context Web UI available at http://localhost:4040
Spark context available as 'sc' (master = local[*], app id = local-1710434021996).
Spark session available as 'spark'.

scala> spark.read.text("gs://apache-spark-bucket/README.md").count()
val res0: Long = 124

scala> spark.read.orc("examples/src/main/resources/users.orc").write.mode("overwrite").orc("gs://apache-spark-bucket/users.orc")

scala> spark.read.orc("gs://apache-spark-bucket/users.orc").show()
+------+--------------+----------------+
|  name|favorite_color|favorite_numbers|
+------+--------------+----------------+
|Alyssa|          NULL|  [3, 9, 15, 20]|
|   Ben|           red|              []|
+------+--------------+----------------+

scala>
```

### Was this patch authored or co-authored using generative AI tooling?

No.

Closes apache#45521 from dongjoon-hyun/SPARK-47400.

Authored-by: Dongjoon Hyun <dhyun@apple.com>
Signed-off-by: Dongjoon Hyun <dhyun@apple.com>
### What changes were proposed in this pull request?

This PR aims to update `YuniKorn` docs with v1.5 for Apache Spark 4.0.0.

### Why are the changes needed?

Apache YuniKorn v1.5.0 was released on 2024-03-14 with 219 resolved JIRAs.

- https://yunikorn.apache.org/release-announce/1.5.0
    - Kubernetes version support: v1.24 ~ 1.29
    - Event streaming API
    - Web UI enhancements
    - Improved Prometheus metric grouping
    - Revamped scheduler initialization support
    - Better allocation traceability
    - REST API enhancements

I installed YuniKorn v1.5.0 on K8s 1.29 and tested manually.

**K8s v1.29**
```
$ kubectl version
Client Version: v1.29.2
Kustomize Version: v5.0.4-0.20230601165947-6ce0bf390ce3
Server Version: v1.29.1
```

**YuniKorn v1.4**
```
$ helm list -n yunikorn
NAME    	NAMESPACE	REVISION	UPDATED                             	STATUS  	CHART         	APP VERSION
yunikorn	yunikorn 	1       	2024-03-14 10:05:25.323637 -0700 PDT	deployed	yunikorn-1.5.0
```

```
$ build/sbt -Pkubernetes -Pkubernetes-integration-tests -Dspark.kubernetes.test.deployMode=docker-desktop "kubernetes-integration-tests/testOnly *.YuniKornSuite" -Dtest.exclude.tags=minikube,local,decom,r -Dtest.default.exclude.tags=
...
[info] YuniKornSuite:
[info] - SPARK-42190: Run SparkPi with local[*] (6 seconds, 893 milliseconds)
[info] - Run SparkPi with no resources (8 seconds, 801 milliseconds)
[info] - Run SparkPi with no resources & statefulset allocation (8 seconds, 809 milliseconds)
[info] - Run SparkPi with a very long application name. (9 seconds, 779 milliseconds)
[info] - Use SparkLauncher.NO_RESOURCE (8 seconds, 855 milliseconds)
[info] - Run SparkPi with a master URL without a scheme. (8 seconds, 787 milliseconds)
[info] - Run SparkPi with an argument. (8 seconds, 867 milliseconds)
[info] - Run SparkPi with custom labels, annotations, and environment variables. (8 seconds, 897 milliseconds)
[info] - All pods have the same service account by default (7 seconds, 776 milliseconds)
[info] - Run extraJVMOptions check on driver (5 seconds, 424 milliseconds)
[info] - SPARK-42474: Run extraJVMOptions JVM GC option check - G1GC (5 seconds, 876 milliseconds)
[info] - SPARK-42474: Run extraJVMOptions JVM GC option check - Other GC (4 seconds, 841 milliseconds)
[info] - SPARK-42769: All executor pods have SPARK_DRIVER_POD_IP env variable (9 seconds, 812 milliseconds)
[info] - Verify logging configuration is picked from the provided SPARK_CONF_DIR/log4j2.properties (10 seconds, 288 milliseconds)
[info] - Run SparkPi with env and mount secrets. (12 seconds, 83 milliseconds)
[info] - Run PySpark on simple pi.py example (9 seconds, 813 milliseconds)
[info] - Run PySpark to test a pyfiles example (9 seconds, 923 milliseconds)
[info] - Run PySpark with memory customization (9 seconds, 811 milliseconds)
[info] - Run in client mode. (4 seconds, 364 milliseconds)
[info] - Start pod creation from template (8 seconds, 817 milliseconds)
[info] - SPARK-38398: Schedule pod creation from template (9 seconds, 839 milliseconds)
[info] - A driver-only Spark job with a tmpfs-backed localDir volume (6 seconds, 121 milliseconds)
[info] - A driver-only Spark job with a tmpfs-backed emptyDir data volume (5 seconds, 839 milliseconds)
[info] - A driver-only Spark job with a disk-backed emptyDir volume (5 seconds, 898 milliseconds)
[info] - A driver-only Spark job with an OnDemand PVC volume (6 seconds, 239 milliseconds)
[info] - A Spark job with tmpfs-backed localDir volumes (9 seconds, 63 milliseconds)
[info] - A Spark job with two executors with OnDemand PVC volumes (8 seconds, 938 milliseconds)
[info] - PVs with local hostpath storage on statefulsets !!! CANCELED !!! (2 milliseconds)
...
[info] - PVs with local hostpath and storageClass on statefulsets !!! IGNORED !!!
[info] - PVs with local storage !!! CANCELED !!! (1 millisecond)
...
[info] Run completed in 6 minutes, 30 seconds.
[info] Total number of tests run: 27
[info] Suites: completed 1, aborted 0
[info] Tests: succeeded 27, failed 0, canceled 2, ignored 1, pending 0
[info] All tests passed.
[success] Total time: 396 s (06:36), completed Mar 14, 2024, 10:19:47 AM
```

```
$ kubectl describe pod -l spark-role=driver -n  spark-f5906efe18864a22be5397ffa30f2b30
...
Events:
  Type    Reason             Age   From      Message
  ----    ------             ----  ----      -------
  Normal  Scheduling         1s    yunikorn  spark-f5906efe18864a22be5397ffa30f2b30/spark-test-app-11c19d5f8b914e719d9d5e4333e7fe16-driver is queued and waiting for allocation
  Normal  Scheduled          1s    yunikorn  Successfully assigned spark-f5906efe18864a22be5397ffa30f2b30/spark-test-app-11c19d5f8b914e719d9d5e4333e7fe16-driver to node docker-desktop
  Normal  PodBindSuccessful  1s    yunikorn  Pod spark-f5906efe18864a22be5397ffa30f2b30/spark-test-app-11c19d5f8b914e719d9d5e4333e7fe16-driver is successfully bound to node docker-desktop
  Normal  Pulled             1s    kubelet   Container image "docker.io/kubespark/spark:dev" already present on machine
  Normal  Created            1s    kubelet   Created container spark-kubernetes-driver
  Normal  Started            1s    kubelet   Started container spark-kubernetes-driver
```

### Does this PR introduce _any_ user-facing change?

No.

### How was this patch tested?

Manual review.

### Was this patch authored or co-authored using generative AI tooling?

No.

Closes apache#45523 from dongjoon-hyun/SPARK-47401.

Authored-by: Dongjoon Hyun <dhyun@apple.com>
Signed-off-by: Dongjoon Hyun <dhyun@apple.com>
…esDialect

### What changes were proposed in this pull request?

This PR fixes a bug in SPARK-47390, we shall separate TIME from TIMESTAMP case-match branch

### Why are the changes needed?

bugfix

### Does this PR introduce _any_ user-facing change?

no

### How was this patch tested?

local test with apache#45519 merged together

### Was this patch authored or co-authored using generative AI tooling?
no

Closes apache#45522 from yaooqinn/SPARK-47390-F.

Authored-by: Kent Yao <yao@apache.org>
Signed-off-by: Dongjoon Hyun <dhyun@apple.com>
### What changes were proposed in this pull request?
The pr aims to remove some unused error classes, include:
- _LEGACY_ERROR_TEMP_3193: The code in PR(apache#44961) no longer uses this error class, but the related item in the file `error-classes.json` has not been deleted synchronously.
- _LEGACY_ERROR_TEMP_3197: The code in PR(apache#44961) no longer uses this error class, but the related item in the file `error-classes.json` has not been deleted synchronously.
- _LEGACY_ERROR_TEMP_3234: This should be the follow-up work of apache#44665
- STATE_STORE_MULTIPLE_VALUES_PER_KEY: Introduced by [SPARK-46864](apache#44883), but not actually used in codebase.
- TWS_VALUE_SHOULD_NOT_BE_NULL: Introduced by [SPARK-46864](apache#44883), but not actually used in codebase.

### Why are the changes needed?
Make the code cleaner.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
- Pass GA.
- Manually check.

### Was this patch authored or co-authored using generative AI tooling?
No.

Closes apache#45509 from panbingkun/remove_outdated_error_classes.

Authored-by: panbingkun <panbingkun@baidu.com>
Signed-off-by: Max Gekk <max.gekk@gmail.com>
…o TimestampNTZType

### What changes were proposed in this pull request?

Add a general mapping for TIME WITHOUT TIME ZONE to TimestampNTZType

### Why are the changes needed?

TIME WITHOUT TIME ZONE should be able to be configured to map to TimestampNTZType like TIMESTAMP WITHOUT TIME ZONE

### Does this PR introduce _any_ user-facing change?

yes, TIME WITHOUT TIME ZONE can be mapped to TimestampNTZType when preferTimestampNTZ is true

### How was this patch tested?

new tests

### Was this patch authored or co-authored using generative AI tooling?
no

Closes apache#45519 from yaooqinn/SPARK-47396.

Authored-by: Kent Yao <yao@apache.org>
Signed-off-by: Dongjoon Hyun <dhyun@apple.com>
### What changes were proposed in this pull request?

This PR aims to upgrade `ZooKeeper` to 3.9.2.

### Why are the changes needed?

To match the versions with the latest server-side bug fixes.
- https://zookeeper.apache.org/doc/r3.9.2/releasenotes.html

### Does this PR introduce _any_ user-facing change?

No.

### How was this patch tested?

Pass the CIs.

### Was this patch authored or co-authored using generative AI tooling?

No.

Closes apache#45524 from dongjoon-hyun/SPARK-47402.

Authored-by: Dongjoon Hyun <dhyun@apple.com>
Signed-off-by: Dongjoon Hyun <dhyun@apple.com>
### What changes were proposed in this pull request?

This PR adds the implementation of the `parse_json` expression to replace the current fake implementation. This expression parses a JSON string as a variant value. It throws an exception when the string is not a valid JSON value or when the resulting variant cannot fit into the size limit.

This feature is achieved by introducing a library`common/variant`. It includes the utility functions to build and manipulate binary-encoded variant values. It also contains a `README` that describes the variant binary format. It is intended that this library can be used outside of Spark.

Some usage examples of `parse_json`:

```
select parse_json('{"a": 1, "b": 2}');

create table variant_table as select parse_json(j) as v from json_table;

select parse_json('[');
-- will throw an exception because the input is not valid JSON

select parse_json('"' || repeat('a', 16 * 1024 * 1024) || '"');
-- will throw an exception because the variant exceeds the size limit
```

### Does this PR introduce _any_ user-facing change?

Yes, the `parse_json` expression will return an actual binary-encoded variant value rather than the original placeholder value.

### How was this patch tested?

Unit tests that validate the `parse_json` result. Negative cases where the expression fail on invalid/large JSON are also covered.

Some unit tests need to be temporarily disabled because the `toString` implementation doesn't match the `parse_json` implementation yet. I will shortly add a new `toString` implementation and re-enable them.

Closes apache#45479 from chenhao-db/parse_json.

Authored-by: Chenhao Li <chenhao.li@databricks.com>
Signed-off-by: Wenchen Fan <wenchen@databricks.com>
### What changes were proposed in this pull request?
The pr aims to upgrade scala from `2.13.12` to `2.13.13`.

### Why are the changes needed?
- The new version bring some bug fixes:
  scala/scala#10525
  scala/scala#10528

- The release notes as follows: https://github.com/scala/scala/releases/tag/v2.13.13

### Does this PR introduce _any_ user-facing change?
Yes, The `scala` version is changed from `2.13.12` to `2.13.13`.

### How was this patch tested?
- Pass GA.
- After the master is upgraded to this version `2.13.13`, we need to continue to observe.

### Was this patch authored or co-authored using generative AI tooling?
No.

Closes apache#45342 from panbingkun/SPARK-47234.

Authored-by: panbingkun <panbingkun@baidu.com>
Signed-off-by: Dongjoon Hyun <dhyun@apple.com>
### What changes were proposed in this pull request?

In MySQL, TIMESTAMP and DATETIME are different. The former is a TIMESTAMP WITH LOCAL TIME ZONE and the latter is a TIMESTAMP WITHOUT TIME ZONE

Following [SPARK-47375](https://issues.apache.org/jira/browse/SPARK-47375), MySql TIMESTAMP goes directly to TimestampType, DATETIME's mapping is decided by preferTimestampNTZ.

### Why are the changes needed?

align the guidelines for jdbc timestamps
### Does this PR introduce _any_ user-facing change?

yes,migration guide provided

### How was this patch tested?

new tests

### Was this patch authored or co-authored using generative AI tooling?

Closes apache#45530 from yaooqinn/SPARK-47406.

Authored-by: Kent Yao <yao@apache.org>
Signed-off-by: Kent Yao <yao@apache.org>
…ations

### What changes were proposed in this pull request?
Disable the ability to use collations in expressions for generated columns.

### Why are the changes needed?
Changing the collation of a column or even just changing the ICU version could lead to a differences in the resulting expression so it would be best if we simply disable it for now.

### Does this PR introduce _any_ user-facing change?

No.

### How was this patch tested?

With new unit tests.

### Was this patch authored or co-authored using generative AI tooling?

No.

Closes apache#45520 from stefankandic/disableGeneratedColumnsCollation.

Authored-by: Stefan Kandic <stefan.kandic@databricks.com>
Signed-off-by: Max Gekk <max.gekk@gmail.com>
### What changes were proposed in this pull request?

- Support java.sql.Types.NULL map to NullType
- Support makeGetter for NullType

### Why are the changes needed?

For better JDBC standard completeness

### Does this PR introduce _any_ user-facing change?

Yes, databases that support void/null type will not raise error

### How was this patch tested?

new tests
### Was this patch authored or co-authored using generative AI tooling?

no

Closes apache#45531 from yaooqinn/SPARK-47407.

Authored-by: Kent Yao <yao@apache.org>
Signed-off-by: Max Gekk <max.gekk@gmail.com>
…tils`

### What changes were proposed in this pull request?
The pr aims to move `log4j2-defaults.properties` from `core/src/main/resources/org/apache/spark` to `common/utils/src/main/resources/org/apache/spark`.
<img width="1114" alt="image" src="https://github.com/apache/spark/assets/15246973/fc931f80-615c-4cf6-8a3b-2a4434d452c9">

### Why are the changes needed?
- As the class `org.apache.spark.internal.Logging` is moved from module `core` to `common/utils`, the corresponding file `log4j2-defaults.properties` should be moved as well.
- Fix bug, as follows:
  Before:
  <img width="721" alt="image" src="https://github.com/apache/spark/assets/15246973/a5448754-354f-401f-a656-0a96ce5fcef0">

  After:
  <img width="719" alt="image" src="https://github.com/apache/spark/assets/15246973/4f0c5994-0e99-48df-95aa-9119b91c8319">

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
- Manually test.
- Pass GA.

### Was this patch authored or co-authored using generative AI tooling?
No.

Closes apache#45532 from panbingkun/SPARK-47419.

Authored-by: panbingkun <panbingkun@baidu.com>
Signed-off-by: Dongjoon Hyun <dhyun@apple.com>
@ericm-db ericm-db closed this Apr 1, 2024
ericm-db pushed a commit that referenced this pull request Mar 2, 2026
…framework

### What changes were proposed in this pull request?

Add a `--DEBUG` marker directive for the SQL golden file test framework (`SQLQueryTestSuite`). When placed on its own line before a query in an input `.sql` file, it enables a focused debug mode:

- **Selective execution**: Only `--DEBUG`-marked queries and setup commands (CREATE TABLE, INSERT, SET, etc.) are executed; all other queries are skipped.
- **Full error details**: Failed queries print the complete stacktrace to the console.
- **Golden comparison**: Results are still compared against the golden file so you can verify correctness.
- **Safety net**: The test always fails at the end with a reminder to remove `--DEBUG` markers before committing.
- **DataFrame access**: Documentation guides users to set a breakpoint in `runDebugQueries` to inspect the `DataFrame` instance for ad-hoc plan analysis.

Example usage in an input file:
```sql
CREATE TABLE t (id INT, val INT) USING parquet;
INSERT INTO t VALUES (1, 10), (2, 20);
-- this query is skipped in debug mode
SELECT count(*) FROM t;
-- this is the query I'm debugging
--DEBUG
SELECT sum(val) OVER (ORDER BY id) FROM t;
```

Example console output when running the test:
```
=== DEBUG: Query #3 ===
SQL: SELECT sum(val) OVER (ORDER BY id) FROM t
Golden answer: matches
```

When the debug query fails:
```
=== DEBUG: Query #3 ===
SQL: SELECT sum(val) OVER (ORDER BY id) FROM t
org.apache.spark.sql.AnalysisException: [ERROR_CLASS] ...
	at org.apache.spark.sql.catalyst.analysis...
	at ...
Golden answer: matches
```

### Why are the changes needed?

Debugging golden file test failures is currently painful:
1. You must run all queries even if only one needs debugging.
2. Error output is minimal (just the error class/message), with no stacktrace.
3. There is no way to access the `DataFrame` instance for plan inspection.

This change addresses all three issues with a simple, zero-config marker that can be temporarily added during development.

### Does this PR introduce _any_ user-facing change?

No. This is a test infrastructure improvement only.

### How was this patch tested?

- Manually tested with `--DEBUG` markers on both passing and failing queries in `inline-table.sql`.
- Verified backward compatibility: tests pass normally when no `--DEBUG` markers are present.
- Verified debug mode output includes full stacktraces, golden answer comparison, and the safety-fail message.

### Was this patch authored or co-authored using generative AI tooling?

Yes. cursor

Closes apache#54554 from cloud-fan/golden.

Authored-by: Wenchen Fan <wenchen@databricks.com>
Signed-off-by: Kent Yao <kentyao@microsoft.com>
ericm-db pushed a commit that referenced this pull request May 5, 2026
### What changes were proposed in this pull request?

Address the open follow-ups from [SPARK-56681](https://issues.apache.org/jira/browse/SPARK-56681) (umbrella for PATH / SPARK-56605 cleanup) in a single cleanup PR. Items #1 and #2 were already wired by SPARK-56639; this PR covers the remainder.

| # | Item | Resolution |
|---|---|---|
| #1 | `FunctionResolution.resolveProcedure` was dead code | Already wired by SPARK-56639 (no action). |
| #2 | Frozen view / SQL-function PATH wiring unfinished | Already done by SPARK-56639 (no action). |
| #3 | `AnalysisContext.resolutionPathEntries` threadlocal | Audit only: confirmed `withNewAnalysisContext` / `reset()` correctly clear it. Full removal needs a coordinated refactor to plumb the path through `RelationResolution` / `FunctionResolution` method calls; flagged as a follow-up. |
| #4 | `Analyzer.executeAndCheck` clobbers outer `SQLConf.withExistingConf` | Extracted `runWithSessionConf` helper, added `SQLConf.getExistingConfIfSet`. `executeAndCheck` and `executeSameContext` now share one path that yields to any outer scope. |
| #5 | `VariableResolution.allowUnqualifiedSessionTempVariableLookup` force-loads default catalog | Replaced the hot-path catalog read with `CatalogManager.isSystemSessionOnPath`, which inspects stored session-path entries directly. No catalog load on column resolution. |
| #6 | `DROP VARIABLE` PATH gate asymmetric with `DECLARE` / `CREATE` | Removed the gate. DDL on session variables (`DECLARE` / `CREATE` / `DROP`) always targets `system.session` directly; only DML (`SET VAR`, `SELECT x`) goes through PATH. |
| #7 | `lookupFunctionType` exception swallow too broad | Narrowed from `NonFatal` to the explicit not-found list (`NoSuchFunctionException`, `NoSuchNamespaceException`, `CatalogNotFoundException`, `FORBIDDEN_OPERATION`). Other exceptions propagate. |
| #8 | `lookupFunctionType` fan-out had wasteful `system.*` candidates | Filtered them out — `system.session`, `system.builtin`, `system.ai` are already resolved earlier in the same method. |
| #9 | Three near-duplicate path-resolution helpers | Lifted into `CatalogManager.resolutionPathEntriesForAnalysis(pinnedEntries, viewCatalogAndNamespace)`. Relation, routine, and procedure resolution all route through it. |
| #10 | Tests for the new error paths and gates | Added a DECLARE / SET VAR / DROP cycle test under non-default PATH and a struct-variable field-vs-qualified ambiguity test in `sql-session-variables.sql`. |
| #11 | `ProtoToParsedPlanTestSuite.analyzerIsolationConf` was a bare `SQLConf` | Clone `spark.sessionState.conf` and only override `PATH_ENABLED=false`, so all `sparkConf` overrides (ANSI, alias config, ...) propagate automatically. |
| Bonus | `ResolveSetVariable` hardcoded `SYSTEM.SESSION` regardless of actual PATH | `unresolvedVariableError` now takes `Seq[Seq[String]]` path entries with **required** `Origin` (no overloads). DML lookup failures (`SET VAR`, `FETCH ... INTO`) report the full SQL path as a bracketed list, byte-for-byte consistent with `UNRESOLVED_ROUTINE` and `TABLE_OR_VIEW_NOT_FOUND`. DDL name validation in `ResolveCatalogs` continues to report `[system.session]` since PATH does not apply there. Origin is plumbed through `VariableManager.set` so all error sites carry a `queryContext` pointing at the offending variable identifier (parser opt-ins via `withOrigin(identifierReference)` so the highlight is the variable name, not the whole statement). |

### Why are the changes needed?

These are the cleanup items called out on SPARK-56681 from the post-merge source review of SPARK-56605. They eliminate dead code paths, plug user-visible bugs (force-loading a misconfigured default catalog on column resolution; clobbering pinned session configs; swallowing real catalog errors as `UNRESOLVED_ROUTINE`), remove the asymmetry between DDL and DML on session variables, and make `UNRESOLVED_VARIABLE` self-consistent with the other "not found" errors.

### Does this PR introduce _any_ user-facing change?

Yes.

- **`UNRESOLVED_VARIABLE.searchPath`** is now rendered as a bracketed list. For DML lookups (`SET VAR`, `FETCH ... INTO`), the list reflects the actual SQL PATH that was consulted instead of a hardcoded `SYSTEM.SESSION`. For DDL name validation (`DECLARE` / `DROP` with a non-session namespace), the list is `[`` `system`.`session` ``]` since PATH does not apply.
- **`UNRESOLVED_VARIABLE`** now always carries a `queryContext` that highlights just the offending variable identifier (e.g. `"builtin.var1"`, `"ses.var1"`), not the whole `DECLARE` / `SET VAR` statement.
- **`DROP TEMPORARY VARIABLE`** no longer raises `UNRESOLVED_VARIABLE` when the SQL PATH does not contain `system.session`. DDL on session variables ignores PATH, matching the existing behaviour of `DECLARE OR REPLACE VARIABLE`.
- **`lookupFunctionType`** no longer swallows non–`NotFound` errors. A catalog reporting `PERMISSION_DENIED` (or similar) for a function lookup now propagates instead of silently producing `UNRESOLVED_ROUTINE`.

### How was this patch tested?

- Added `sql-session-variables.sql` regression test for the struct-variable field-vs-qualified ambiguity (`DECLARE VARIABLE session STRUCT<a INT>` → `SELECT session.a` succeeds → `DROP` → `SELECT session.a` falls through to `UNRESOLVED_COLUMN`).
- Updated `SetPathSuite`: DECLARE / SET VAR / DROP cycle under a non-default PATH; bonus test asserts the actual rendered search path and the variable-identifier `queryContext`.
- Updated `SqlScriptingExecutionSuite` for the new bracketed `searchPath` and identifier-pinned `queryContext`.
- Regenerated `sql-session-variables.sql.out` for the new error shape.
- Added `resolutionPathEntriesForAnalysis` stubs to mocked `CatalogManager` instances in `PlanResolutionSuite`, `AlignAssignmentsSuiteBase`, and `TableLookupCacheSuite`.
- Ran focused suites locally; all pass:
  - `build/sbt 'sql/testOnly *SetPathSuite *SqlScriptingExecutionSuite *ExecuteImmediateEndToEndSuite'`
  - `build/sbt 'sql/testOnly *SimpleSQLViewSuite *SQLFunctionSuite'`
  - `build/sbt 'sql/testOnly *PlanResolutionSuite *UpdateTableAlignAssignmentsSuite *MergeIntoTableAlignAssignmentsSuite'`
  - `build/sbt 'catalyst/testOnly *TableLookupCacheSuite *AnalysisSuite *AnalysisErrorSuite *LookupFunctionsSuite'`
  - `build/sbt 'sql/testOnly *FunctionQualificationSuite *RelationQualificationSuite *DataSourceV2FunctionSuite'`
  - `build/sbt 'sql/testOnly *SQLQuerySuite'`
  - `build/sbt 'connect/testOnly *ProtoToParsedPlanTestSuite'`
  - `build/sbt 'sql/testOnly *SQLQueryTestSuite -- -z sql-session-variables.sql'`
  - Full `org.apache.spark.sql.catalyst.analysis.*`, `org.apache.spark.sql.catalyst.parser.*`, and `org.apache.spark.sql.analysis.resolver.*` suites.
- `scalastyle` and `scalafmt` clean across catalyst, sql, and connect modules.

### Was this patch authored or co-authored using generative AI tooling?

Generated-by: Cursor Claude Opus 4.7

Closes apache#55647 from srielau/SPARK-56681-patch-clean-up.

Authored-by: Serge Rielau <serge@rielau.com>
Signed-off-by: Daniel Tenedorio <daniel.tenedorio@databricks.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.