Backend
ClickHouse (only ClickHouse offloads CSV scans; Velox and Bolt fall back).
Bug description
FileSourceScanExecTransformer.getProperties builds the native text reader options from the CSV options, and it only forwards the legacy delimiter key. sep, which is Spark's canonical key and what spark.read.option("sep", ...) and OPTIONS (sep '|') set, is ignored, so the native reader splits on its default comma:
// file content: "a|1\nb|2\n"
spark.read.schema("c1 string, c2 int").option("sep", "|").csv(path).show()
// vanilla: (a, 1), (b, 2); ClickHouse native: c1 = "a|1", c2 = null
The same table read with delimiter '|' is correct. Spark's own "test aliases sep and encoding for delimiter and charset" in CSVSuite is excluded for ClickHouse.
Two related gaps in the same mapping:
- Spark decodes the delimiter with
CSVExprUtils.toDelimiterStr, so sep '\t' written as a backslash and a t is a tab. The raw string is forwarded, and the native reader splits on its first byte, the backslash.
- The native reader only uses the first byte of the delimiter, so a multi-character delimiter such as
|| or a non-ASCII one is read wrongly instead of falling back.
The ClickHouse path was traced by reading the code; I could not run the ClickHouse backend.
Gluten version
main (2681cc9)
Spark version
Spark 3.5
This issue was written with the assistance of AI (Claude Opus).
Backend
ClickHouse (only ClickHouse offloads CSV scans; Velox and Bolt fall back).
Bug description
FileSourceScanExecTransformer.getPropertiesbuilds the native text reader options from the CSV options, and it only forwards the legacydelimiterkey.sep, which is Spark's canonical key and whatspark.read.option("sep", ...)andOPTIONS (sep '|')set, is ignored, so the native reader splits on its default comma:The same table read with
delimiter '|'is correct. Spark's own "test aliases sep and encoding for delimiter and charset" inCSVSuiteis excluded for ClickHouse.Two related gaps in the same mapping:
CSVExprUtils.toDelimiterStr, sosep '\t'written as a backslash and atis a tab. The raw string is forwarded, and the native reader splits on its first byte, the backslash.||or a non-ASCII one is read wrongly instead of falling back.The ClickHouse path was traced by reading the code; I could not run the ClickHouse backend.
Gluten version
main (2681cc9)
Spark version
Spark 3.5
This issue was written with the assistance of AI (Claude Opus).