[GLUTEN-11088][VL] Fall back CSV reader - #11190
Conversation
|
Run Gluten ClickHouse CI on ARM |
|
Run Gluten ClickHouse CI on ARM |
|
Run Gluten ClickHouse CI on x86 |
|
Run Gluten ClickHouse CI on ARM |
|
Passed the tests one time, but after rerun CSV failed by |
|
Run Gluten ClickHouse CI on ARM |
|
Trigger the flaky test |
zhztheplayer
left a comment
There was a problem hiding this comment.
It seems the CI is still failing
|
Maybe we need to compile the arrow, flaky test may cause by platform difference |
|
Run Gluten Clickhouse CI on x86 |
|
Run Gluten Clickhouse CI on x86 |
|
Run Gluten Clickhouse CI on x86 |
|
Run Gluten Clickhouse CI on x86 |
| BloomFilterMightContainJointRewriteRule.apply( | ||
| c.session, | ||
| c.caller.isBloomFilterStatFunction())) | ||
| injector.injectPreTransform(c => ArrowScanReplaceRule.apply(c.session)) |
There was a problem hiding this comment.
Just to confirm: is CSV format no longer supported, or do we only need a fallback for Spark 40 and later versions?
There was a problem hiding this comment.
CSV format is no longer supported
|
Run Gluten Clickhouse CI on x86 |
|
Could you help approve? Thanks! @zhztheplayer |
philo-he
left a comment
There was a problem hiding this comment.
@jinchengchenghh, if Arrow CSV reader is not required, can we directly use the official Apache Arrow Jar to replace the Jar locally built by developers? cc @zhouyuan
|
I remember there is several patches applied to arrow 15, not only csv reader related change, for arrow 18(Spark4.0), we use the official release @philo-he |
@jinchengchenghh, do we need to remove those CSV-reader-specific patches under |
|
This patch only fallbacks the csv reader, we does not remove all the csv related code from java code, when we decide to remove it, we will also remove the patch, I'm not sure if some customer may be interested on it. |
philo-he
left a comment
There was a problem hiding this comment.
Thanks for the clarification.
|
@jinchengchenghh would you please also fallback csv for spark 4.1? |
|
Yes, csv fall back for all the Spark version in this PR @baibaichen |
Oh, right. We also need to re-enable the CSV-related suites in Spark 4.1. |
This reverts commit e121903.
…config Following the removal of ArrowConvertorRule/ArrowScanReplaceRule (already unwired from VeloxRuleApi by PR apache#11190 "[GLUTEN-11088][VL] Fall back CSV reader" merged 2026-01-19), the spark.gluten.sql.native.arrow.reader.enabled config and its plumbing have no consumers: * GlutenConfig.enableNativeArrowReader / NATIVE_ARROW_READER_ENABLED * BackendSettingsApi.enableNativeArrowReadFiles (default) * VeloxBackend.enableNativeArrowReadFiles (override) Test suites still set this flag (MiscOperatorSuite, GlutenCSVSuite, GlutenReadSchemaSuite across spark35/40/41) but it has been a no-op since PR apache#11190; CSV continues to be covered by these suites via the Spark native CSV path. The corresponding entry in docs/Configuration.md is also removed. Generated-by: Claude claude-opus-4.7
…config Following the removal of ArrowConvertorRule/ArrowScanReplaceRule (already unwired from VeloxRuleApi by PR apache#11190 "[GLUTEN-11088][VL] Fall back CSV reader" merged 2026-01-19), the spark.gluten.sql.native.arrow.reader.enabled config and its plumbing have no consumers: * GlutenConfig.enableNativeArrowReader / NATIVE_ARROW_READER_ENABLED * BackendSettingsApi.enableNativeArrowReadFiles (default) * VeloxBackend.enableNativeArrowReadFiles (override) Test suites still set this flag (MiscOperatorSuite, GlutenCSVSuite, GlutenReadSchemaSuite across spark35/40/41) but it has been a no-op since PR apache#11190; CSV continues to be covered by these suites via the Spark native CSV path. The corresponding entry in docs/Configuration.md is also removed. Generated-by: Claude claude-opus-4.7
…config Following the removal of ArrowConvertorRule/ArrowScanReplaceRule (already unwired from VeloxRuleApi by PR apache#11190 "[GLUTEN-11088][VL] Fall back CSV reader" merged 2026-01-19), the spark.gluten.sql.native.arrow.reader.enabled config and its plumbing have no consumers: * GlutenConfig.enableNativeArrowReader / NATIVE_ARROW_READER_ENABLED * BackendSettingsApi.enableNativeArrowReadFiles (default) * VeloxBackend.enableNativeArrowReadFiles (override) Test suites still set this flag (MiscOperatorSuite, GlutenCSVSuite, GlutenReadSchemaSuite across spark35/40/41) but it has been a no-op since PR apache#11190; CSV continues to be covered by these suites via the Spark native CSV path. The corresponding entry in docs/Configuration.md is also removed. Generated-by: Claude claude-opus-4.7
…config Following the removal of ArrowConvertorRule/ArrowScanReplaceRule (already unwired from VeloxRuleApi by PR apache#11190 "[GLUTEN-11088][VL] Fall back CSV reader" merged 2026-01-19), the spark.gluten.sql.native.arrow.reader.enabled config and its plumbing have no consumers: * GlutenConfig.enableNativeArrowReader / NATIVE_ARROW_READER_ENABLED * BackendSettingsApi.enableNativeArrowReadFiles (default) * VeloxBackend.enableNativeArrowReadFiles (override) Test suites still set this flag (MiscOperatorSuite, GlutenCSVSuite, GlutenReadSchemaSuite across spark35/40/41) but it has been a no-op since PR apache#11190; CSV continues to be covered by these suites via the Spark native CSV path. The corresponding entry in docs/Configuration.md is also removed. Generated-by: Claude claude-opus-4.7
…config Following the removal of ArrowConvertorRule/ArrowScanReplaceRule (already unwired from VeloxRuleApi by PR apache#11190 "[GLUTEN-11088][VL] Fall back CSV reader" merged 2026-01-19), the spark.gluten.sql.native.arrow.reader.enabled config and its plumbing have no consumers: * GlutenConfig.enableNativeArrowReader / NATIVE_ARROW_READER_ENABLED * BackendSettingsApi.enableNativeArrowReadFiles (default) * VeloxBackend.enableNativeArrowReadFiles (override) Test suites still set this flag (MiscOperatorSuite, GlutenCSVSuite, GlutenReadSchemaSuite across spark35/40/41) but it has been a no-op since PR apache#11190; CSV continues to be covered by these suites via the Spark native CSV path. The corresponding entry in docs/Configuration.md is also removed. Generated-by: Claude claude-opus-4.7
) * [MINOR][VL] Remove dead Arrow-CSV / Arrow-Dataset JVM code path The ArrowCSV file format and ArrowBatchScanExec chain are unreachable: no injection in VeloxRuleApi, no META-INF/services entry, and all ArrowCsvScanSuite cases are @ignore'd. They were introduced as a squash-merge byproduct in #11776 and never wired up. Verified by compiling: * spark-3.5 + scala-2.12 + arrow 15.0.0-gluten (install) * spark-4.0 + scala-2.13 + arrow 18.1.0 (compile) Generated-by: claude-opus-4.7 * [MINOR][VL] Remove dead Arrow dataset reader paths from ArrowUtil makeArrowDiscovery / readArrowSchema / readArrowFileColumnNames / readSchema(FragmentScanOptions) overloads / loadMissingColumns / loadPartitionColumns / loadBatch in ArrowUtil have zero callers across the repo after the previous removal of the ArrowCSV chain. Drop them together with the now-unused imports (arrow.dataset.*, FileStatus, URI/URLDecoder, ArrowRecordBatch, Optional, Logging, etc.). Verified by compiling: * spark-3.5 + scala-2.12 (test-compile, patched arrow 15.0.0-gluten) * spark-4.0 + scala-2.13 (compile, pure arrow 18.1.0) Generated-by: claude-opus-4.7 * [MINOR][VL] Remove dead spark.gluten.sql.native.arrow.reader.enabled config Following the removal of ArrowConvertorRule/ArrowScanReplaceRule (already unwired from VeloxRuleApi by PR #11190 "[GLUTEN-11088][VL] Fall back CSV reader" merged 2026-01-19), the spark.gluten.sql.native.arrow.reader.enabled config and its plumbing have no consumers: * GlutenConfig.enableNativeArrowReader / NATIVE_ARROW_READER_ENABLED * BackendSettingsApi.enableNativeArrowReadFiles (default) * VeloxBackend.enableNativeArrowReadFiles (override) Test suites still set this flag (MiscOperatorSuite, GlutenCSVSuite, GlutenReadSchemaSuite across spark35/40/41) but it has been a no-op since PR #11190; CSV continues to be covered by these suites via the Spark native CSV path. The corresponding entry in docs/Configuration.md is also removed. Generated-by: Claude claude-opus-4.7 * [MINOR][UT] Fix spotless violation in GlutenReadSchemaSuite Generated-by: claude-opus-4.7 * [MINOR][UT] Fix spotless violation in spark40/spark41 GlutenReadSchemaSuite Generated-by: claude-opus-4.7 * [MINOR][UT] Remove unused GlutenConfig import from GlutenCSVSuite (spark35/40/41) Generated-by: claude-opus-4.7
Related issue: #11088