Skip to content

[CH] The CSV sep option is ignored by the native text scan #13203

Description

@LuciferYang

Backend

ClickHouse (only ClickHouse offloads CSV scans; Velox and Bolt fall back).

Bug description

FileSourceScanExecTransformer.getProperties builds the native text reader options from the CSV options, and it only forwards the legacy delimiter key. sep, which is Spark's canonical key and what spark.read.option("sep", ...) and OPTIONS (sep '|') set, is ignored, so the native reader splits on its default comma:

// file content: "a|1\nb|2\n"
spark.read.schema("c1 string, c2 int").option("sep", "|").csv(path).show()
// vanilla: (a, 1), (b, 2); ClickHouse native: c1 = "a|1", c2 = null

The same table read with delimiter '|' is correct. Spark's own "test aliases sep and encoding for delimiter and charset" in CSVSuite is excluded for ClickHouse.

Two related gaps in the same mapping:

  • Spark decodes the delimiter with CSVExprUtils.toDelimiterStr, so sep '\t' written as a backslash and a t is a tab. The raw string is forwarded, and the native reader splits on its first byte, the backslash.
  • The native reader only uses the first byte of the delimiter, so a multi-character delimiter such as || or a non-ASCII one is read wrongly instead of falling back.

The ClickHouse path was traced by reading the code; I could not run the ClickHouse backend.

Gluten version

main (2681cc9)

Spark version

Spark 3.5

This issue was written with the assistance of AI (Claude Opus).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions