Conversation
This Arrow patch added a native CSV/dataset scan-option API (CsvFragmentScanOptions::from, DeserializeMap, mapToExpressionLiteral, etc.) that was originally consumed by Gluten's native Arrow CSV reader path. That path is gone: - apache#11190 fell CSV back to vanilla Spark - 0658e90 / 97f4638 / 9ea8290 / d206c5e removed the JVM callers, the ArrowUtil reader path, and the spark.gluten.sql.native.arrow.reader.enabled config Nothing inside Gluten or Velox now references the symbols the patch introduces, so it is dead code in the build. This change drops: - ep/build-velox/src/modify_arrow_dataset_scan_option.patch - the `patch -p1 <...>` line in dev/build-arrow.sh - the `cp` / `git add` lines in ep/build-velox/src/get-velox.sh that staged the patch into the Velox source tree Verification (Ubuntu 24.04, x86_64): - grep across the repo: 0 callers of CsvFragmentScanOptions, CsvFragmentScanOptions::from, DeserializeMap, mapToExpressionLiteral, createNative(... FragmentScanOptions), CsvConvertOptions, testCsvConvertOptions - `nm -C` on the freshly built libarrow_dataset.a (java-dist and cpp-jni) shows none of the patch's symbols are present - arrow_ep/cpp/src/arrow/dataset/file_csv.cc on disk is the unmodified upstream source -- patch was never applied - dev/buildbundle-veloxbe.sh --enable_vcpkg=ON completes BUILD SUCCESS for all 5 Spark profiles (3.3 / 3.4 / 3.5 / 4.0 / 4.1) - spark-shell on the resulting bundle, reading a CSV file, prints "GlutenFallbackReporter: Validation failed for plan: Scan csv , due to: Unsupported file format TextReadFormat" and produces a vanilla `FileScan csv` physical plan -- confirming CSV is fallback-by-design and never enters the native path the dropped patch fed Generated-by: Claude claude-opus-4.7
|
Thanks for the PR. Just to confirm – will we permanently remove the arrow based CSV reader support? cc @jinchengchenghh, @FelixYBW |
No, we may need to enable it later as it's required by one customer. If we decide to enable it, we will maintain the csv scan. FYI, @zhouyuan |
|
If the Arrow-based CSV reader is still needed, can we introduce a build option to control whether these related patches are applied, and skip the Arrow Java build accordingly (use the official Arrow JAR from Maven Central instead)? If only a few users rely on it, perhaps we could disable it by default. |
|
Closing per @FelixYBW's note above — keeping the CSV scan scaffolding for the potential future customer use case. Thanks for the context! |
|
Thanks, we are now investigating on Velox CSV reader and it seems also have a gaps vs. Spark CSV reader |
|
Just noted the csv reader is already deleted in PR #12130 Let's clean up the patch as well. |
What changes were proposed in this pull request?
Drop
ep/build-velox/src/modify_arrow_dataset_scan_option.patchand the twoshell call sites that apply / stage it:
dev/build-arrow.sh—patch -p1 < .../modify_arrow_dataset_scan_option.patchep/build-velox/src/get-velox.sh—cp+git addthat copy the patchinto the Velox source tree
Why are the changes needed?
The patch added a native CSV / dataset scan-option API
(
CsvFragmentScanOptions::from,DeserializeMap,mapToExpressionLiteral, …)that was originally consumed by Gluten's native Arrow CSV reader path.
That path is gone:
0658e906f/97f463813/9ea8290a/d206c5e20removed the JVM-sidecallers, the
ArrowUtilreader path, and thespark.gluten.sql.native.arrow.reader.enabledconfigNothing inside Gluten or Velox now references the symbols the patch
introduces, so it's dead code in the build. This PR is the final cleanup in
that chain.
How was this patch tested?
Local verification on Ubuntu 24.04 / x86_64:
grepacross the repo — 0 callers ofCsvFragmentScanOptions,CsvFragmentScanOptions::from,DeserializeMap,mapToExpressionLiteral,createNative(... FragmentScanOptions),CsvConvertOptions,testCsvConvertOptionsnm -Con freshly builtlibarrow_dataset.a(bothjava-distandcpp-jnioutputs) — none of the patch's symbols are presentarrow_ep/cpp/src/arrow/dataset/file_csv.ccon disk is the unmodifiedupstream source — the patch was never applied with this change in place
dev/buildbundle-veloxbe.sh --enable_vcpkg=ON→ BUILD SUCCESS for all 5Spark profiles (3.3 / 3.4 / 3.5 / 4.0 / 4.1)
spark-shellon the resulting bundle, reading a CSV file, printsGlutenFallbackReporter: Validation failed for plan: Scan csv , due to: Unsupported file format TextReadFormatand produces a vanillaFileScan csvphysical plan — confirming CSV is fallback-by-design andnever enters the native path the dropped patch fed.
Generated-by: Claude claude-opus-4.7