Skip to content

[GLUTEN-11088][VL] Fall back CSV reader - #11190

Merged
philo-he merged 3 commits into
apache:mainfrom
jinchengchenghh:csv
Jan 19, 2026
Merged

philo-he merged 3 commits into
apache:mainfrom
jinchengchenghh:csv

Conversation

@jinchengchenghh

@jinchengchenghh jinchengchenghh commented Nov 25, 2025 •

Copy link
Copy Markdown
Contributor

Related issue: #11088

@github-actions github-actions Bot added the CORE works for Gluten Core label Nov 25, 2025
@github-actions

Copy link
Copy Markdown

Run Gluten ClickHouse CI on ARM

@github-actions

Copy link
Copy Markdown

Run Gluten ClickHouse CI on ARM

@zhouyuan

Copy link
Copy Markdown
Member

Run Gluten ClickHouse CI on x86

@zhouyuan zhouyuan changed the title [GLUTEN-11088][VL] Enable CSV suite [GLUTEN-11088][VL] Enable CSV suite in Spark-4.0 Nov 26, 2025
@github-actions

Copy link
Copy Markdown

Run Gluten ClickHouse CI on ARM

@jinchengchenghh

jinchengchenghh commented Nov 26, 2025 •

Copy link
Copy Markdown
Contributor Author

Passed the tests one time, but after rerun CSV failed by /arrow/java/dataset/src/main/cpp/jni_util.cc:79: Failed to update reservation while freeing bytes: Java Exception: java.lang.NullPointerException

@jinchengchenghh

Copy link
Copy Markdown
Contributor Author

@github-actions

Copy link
Copy Markdown

Run Gluten ClickHouse CI on ARM

@jinchengchenghh

Copy link
Copy Markdown
Contributor Author

Trigger the flaky test

2025-11-26T21:18:55.4968782Z #
2025-11-26T21:18:55.4981909Z # A fatal error has been detected by the Java Runtime Environment:
2025-11-26T21:18:55.4983467Z #
2025-11-26T21:18:55.4983914Z #  SIGSEGV (0xb) at pc=0x00007f5170655120, pid=82017, tid=107720
2025-11-26T21:18:55.4984658Z #
2025-11-26T21:18:55.4986579Z # JRE version: OpenJDK Runtime Environment 21.9 (17.0.1+12) (build 17.0.1+12-LTS)
2025-11-26T21:18:55.4996055Z # Java VM: OpenJDK 64-Bit Server VM 21.9 (17.0.1+12-LTS, mixed mode, sharing, tiered, compressed oops, compressed class ptrs, g1 gc, linux-amd64)
2025-11-26T21:18:55.4997760Z # Problematic frame:
2025-11-26T21:18:55.4998099Z # C  0x00007f5170655120
2025-11-26T21:18:55.4998416Z #
2025-11-26T21:18:55.5003611Z # Core dump will be written. Default location: Core dumps may be processed with "/lib/systemd/systemd-coredump %P %u %g %s %t 9223372036854775808 %h %d" (or dumping to /__w/incubator-gluten/incubator-gluten/gluten-ut/spark40/core.82017)
2025-11-26T21:18:55.5005320Z #
2025-11-26T21:18:55.5005948Z # An error report file with more information is saved as:
2025-11-26T21:18:55.5006859Z # /__w/incubator-gluten/incubator-gluten/gluten-ut/spark40/hs_err_pid82017.log

@zhztheplayer zhztheplayer left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems the CI is still failing

@jinchengchenghh
jinchengchenghh marked this pull request as draft November 28, 2025 13:10
@jinchengchenghh

Copy link
Copy Markdown
Contributor Author

Maybe we need to compile the arrow, flaky test may cause by platform difference

@github-actions github-actions Bot added the VELOX label Dec 23, 2025
@jinchengchenghh
jinchengchenghh marked this pull request as ready for review December 23, 2025 08:22
@github-actions

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@jinchengchenghh jinchengchenghh changed the title [GLUTEN-11088][VL] Enable CSV suite in Spark-4.0 [GLUTEN-11088][VL] Fallback CSV reader Dec 23, 2025
@github-actions

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@github-actions

github-actions Bot commented Jan 5, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@github-actions

github-actions Bot commented Jan 6, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

BloomFilterMightContainJointRewriteRule.apply(
c.session,
c.caller.isBloomFilterStatFunction()))
injector.injectPreTransform(c => ArrowScanReplaceRule.apply(c.session))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just to confirm: is CSV format no longer supported, or do we only need a fallback for Spark 40 and later versions?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

CSV format is no longer supported

@github-actions

github-actions Bot commented Jan 7, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@jinchengchenghh

Copy link
Copy Markdown
Contributor Author

Could you help approve? Thanks! @zhztheplayer

@philo-he philo-he left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@jinchengchenghh, if Arrow CSV reader is not required, can we directly use the official Apache Arrow Jar to replace the Jar locally built by developers? cc @zhouyuan

@jinchengchenghh

jinchengchenghh commented Jan 19, 2026 •

Copy link
Copy Markdown
Contributor Author

I remember there is several patches applied to arrow 15, not only csv reader related change, for arrow 18(Spark4.0), we use the official release @philo-he

@philo-he

Copy link
Copy Markdown
Member

I remember there is several patches applied to arrow 15, not only csv reader related change, for arrow 15, we use the official release @philo-he

@jinchengchenghh, do we need to remove those CSV-reader-specific patches under ep/build-velox/src/? At some time point, we may directly use official Arrow JAR if some higher Arrow version used by Gluten includes the remaining patches or the remaining patches are only related to Arrow C++, not Java.

@jinchengchenghh

Copy link
Copy Markdown
Contributor Author

This patch only fallbacks the csv reader, we does not remove all the csv related code from java code, when we decide to remove it, we will also remove the patch, I'm not sure if some customer may be interested on it.

@philo-he philo-he changed the title [GLUTEN-11088][VL] Fallback CSV reader [GLUTEN-11088][VL] Fall back CSV reader Jan 19, 2026

@philo-he philo-he left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the clarification.

@philo-he
philo-he merged commit e121903 into apache:main Jan 19, 2026
163 of 168 checks passed
@baibaichen

Copy link
Copy Markdown
Contributor

@jinchengchenghh would you please also fallback csv for spark 4.1?

@jinchengchenghh

Copy link
Copy Markdown
Contributor Author

Yes, csv fall back for all the Spark version in this PR @baibaichen

@baibaichen

Copy link
Copy Markdown
Contributor

Yes, csv fall back for all the Spark version in this PR @baibaichen

Oh, right. We also need to re-enable the CSV-related suites in Spark 4.1.

sh-shamsan added a commit to sh-shamsan/incubator-gluten that referenced this pull request Apr 14, 2026
yaooqinn added a commit to yaooqinn/gluten that referenced this pull request May 23, 2026
…config

Following the removal of ArrowConvertorRule/ArrowScanReplaceRule (already

unwired from VeloxRuleApi by PR apache#11190 "[GLUTEN-11088][VL] Fall back CSV

reader" merged 2026-01-19), the spark.gluten.sql.native.arrow.reader.enabled

config and its plumbing have no consumers:

  * GlutenConfig.enableNativeArrowReader / NATIVE_ARROW_READER_ENABLED

  * BackendSettingsApi.enableNativeArrowReadFiles (default)

  * VeloxBackend.enableNativeArrowReadFiles (override)

Test suites still set this flag (MiscOperatorSuite, GlutenCSVSuite,

GlutenReadSchemaSuite across spark35/40/41) but it has been a no-op since

PR apache#11190; CSV continues to be covered by these suites via the Spark

native CSV path. The corresponding entry in docs/Configuration.md is

also removed.

Generated-by: Claude claude-opus-4.7
yaooqinn added a commit to yaooqinn/gluten that referenced this pull request May 23, 2026
…config

Following the removal of ArrowConvertorRule/ArrowScanReplaceRule (already

unwired from VeloxRuleApi by PR apache#11190 "[GLUTEN-11088][VL] Fall back CSV

reader" merged 2026-01-19), the spark.gluten.sql.native.arrow.reader.enabled

config and its plumbing have no consumers:

  * GlutenConfig.enableNativeArrowReader / NATIVE_ARROW_READER_ENABLED

  * BackendSettingsApi.enableNativeArrowReadFiles (default)

  * VeloxBackend.enableNativeArrowReadFiles (override)

Test suites still set this flag (MiscOperatorSuite, GlutenCSVSuite,

GlutenReadSchemaSuite across spark35/40/41) but it has been a no-op since

PR apache#11190; CSV continues to be covered by these suites via the Spark

native CSV path. The corresponding entry in docs/Configuration.md is

also removed.

Generated-by: Claude claude-opus-4.7
yaooqinn added a commit to yaooqinn/gluten that referenced this pull request May 23, 2026
…config

Following the removal of ArrowConvertorRule/ArrowScanReplaceRule (already

unwired from VeloxRuleApi by PR apache#11190 "[GLUTEN-11088][VL] Fall back CSV

reader" merged 2026-01-19), the spark.gluten.sql.native.arrow.reader.enabled

config and its plumbing have no consumers:

  * GlutenConfig.enableNativeArrowReader / NATIVE_ARROW_READER_ENABLED

  * BackendSettingsApi.enableNativeArrowReadFiles (default)

  * VeloxBackend.enableNativeArrowReadFiles (override)

Test suites still set this flag (MiscOperatorSuite, GlutenCSVSuite,

GlutenReadSchemaSuite across spark35/40/41) but it has been a no-op since

PR apache#11190; CSV continues to be covered by these suites via the Spark

native CSV path. The corresponding entry in docs/Configuration.md is

also removed.

Generated-by: Claude claude-opus-4.7
yaooqinn added a commit to yaooqinn/gluten that referenced this pull request May 23, 2026
…config

Following the removal of ArrowConvertorRule/ArrowScanReplaceRule (already

unwired from VeloxRuleApi by PR apache#11190 "[GLUTEN-11088][VL] Fall back CSV

reader" merged 2026-01-19), the spark.gluten.sql.native.arrow.reader.enabled

config and its plumbing have no consumers:

  * GlutenConfig.enableNativeArrowReader / NATIVE_ARROW_READER_ENABLED

  * BackendSettingsApi.enableNativeArrowReadFiles (default)

  * VeloxBackend.enableNativeArrowReadFiles (override)

Test suites still set this flag (MiscOperatorSuite, GlutenCSVSuite,

GlutenReadSchemaSuite across spark35/40/41) but it has been a no-op since

PR apache#11190; CSV continues to be covered by these suites via the Spark

native CSV path. The corresponding entry in docs/Configuration.md is

also removed.

Generated-by: Claude claude-opus-4.7
yaooqinn added a commit to yaooqinn/gluten that referenced this pull request May 23, 2026
…config

Following the removal of ArrowConvertorRule/ArrowScanReplaceRule (already

unwired from VeloxRuleApi by PR apache#11190 "[GLUTEN-11088][VL] Fall back CSV

reader" merged 2026-01-19), the spark.gluten.sql.native.arrow.reader.enabled

config and its plumbing have no consumers:

  * GlutenConfig.enableNativeArrowReader / NATIVE_ARROW_READER_ENABLED

  * BackendSettingsApi.enableNativeArrowReadFiles (default)

  * VeloxBackend.enableNativeArrowReadFiles (override)

Test suites still set this flag (MiscOperatorSuite, GlutenCSVSuite,

GlutenReadSchemaSuite across spark35/40/41) but it has been a no-op since

PR apache#11190; CSV continues to be covered by these suites via the Spark

native CSV path. The corresponding entry in docs/Configuration.md is

also removed.

Generated-by: Claude claude-opus-4.7
yaooqinn added a commit that referenced this pull request May 25, 2026
)

* [MINOR][VL] Remove dead Arrow-CSV / Arrow-Dataset JVM code path

The ArrowCSV file format and ArrowBatchScanExec chain are unreachable:
no injection in VeloxRuleApi, no META-INF/services entry, and all
ArrowCsvScanSuite cases are @ignore'd. They were introduced as a
squash-merge byproduct in #11776 and never wired up.

Verified by compiling:
  * spark-3.5 + scala-2.12 + arrow 15.0.0-gluten (install)
  * spark-4.0 + scala-2.13 + arrow 18.1.0 (compile)

Generated-by: claude-opus-4.7

* [MINOR][VL] Remove dead Arrow dataset reader paths from ArrowUtil

makeArrowDiscovery / readArrowSchema / readArrowFileColumnNames /
readSchema(FragmentScanOptions) overloads / loadMissingColumns /
loadPartitionColumns / loadBatch in ArrowUtil have zero callers across
the repo after the previous removal of the ArrowCSV chain. Drop them
together with the now-unused imports (arrow.dataset.*, FileStatus,
URI/URLDecoder, ArrowRecordBatch, Optional, Logging, etc.).

Verified by compiling:
  * spark-3.5 + scala-2.12 (test-compile, patched arrow 15.0.0-gluten)
  * spark-4.0 + scala-2.13 (compile, pure arrow 18.1.0)

Generated-by: claude-opus-4.7

* [MINOR][VL] Remove dead spark.gluten.sql.native.arrow.reader.enabled config

Following the removal of ArrowConvertorRule/ArrowScanReplaceRule (already

unwired from VeloxRuleApi by PR #11190 "[GLUTEN-11088][VL] Fall back CSV

reader" merged 2026-01-19), the spark.gluten.sql.native.arrow.reader.enabled

config and its plumbing have no consumers:

  * GlutenConfig.enableNativeArrowReader / NATIVE_ARROW_READER_ENABLED

  * BackendSettingsApi.enableNativeArrowReadFiles (default)

  * VeloxBackend.enableNativeArrowReadFiles (override)

Test suites still set this flag (MiscOperatorSuite, GlutenCSVSuite,

GlutenReadSchemaSuite across spark35/40/41) but it has been a no-op since

PR #11190; CSV continues to be covered by these suites via the Spark

native CSV path. The corresponding entry in docs/Configuration.md is

also removed.

Generated-by: Claude claude-opus-4.7

* [MINOR][UT] Fix spotless violation in GlutenReadSchemaSuite

Generated-by: claude-opus-4.7

* [MINOR][UT] Fix spotless violation in spark40/spark41 GlutenReadSchemaSuite

Generated-by: claude-opus-4.7

* [MINOR][UT] Remove unused GlutenConfig import from GlutenCSVSuite (spark35/40/41)

Generated-by: claude-opus-4.7
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CORE works for Gluten Core VELOX

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants