Skip to content

Native CSV V2 scan parses timestamps as UTC and ignores the session timezone #6332

Description

@andygrove

Describe the bug

With spark.comet.scan.csv.v2.enabled=true, build_csv_source (native/core/src/execution/operators/csv_scan.rs:55) hands DataFusion's CsvSource the Spark schema. In that schema TimestampType is Timestamp(Microsecond, "UTC"), so arrow-csv reads a timestamp without an offset as UTC. Spark's CSV reader interprets it in the timeZone option instead, which defaults to the session timezone. CometCsvNativeScanExec builds its CSVOptions with the session timezone (CometCsvNativeScanExec.scala:92), but the CsvOptions proto has no timezone field, so the timezone never reaches the native reader. In a non-UTC session every such value is silently shifted.

Steps to reproduce

On main at 764936187, with a part-0.csv containing:

id,ts
0,2024-01-15 18:30:45
spark.conf.set("spark.comet.scan.csv.v2.enabled", "true")
spark.conf.set("spark.sql.sources.useV1SourceList", "avro,json,kafka,orc,parquet,text")
spark.conf.set("spark.sql.session.timeZone", "America/Los_Angeles")
spark.read.schema("id INT, ts TIMESTAMP").option("header", "true").csv(dir).collect()

Spark returns the instant 2024-01-16T02:30:45Z and Comet returns 2024-01-15T18:30:45Z.

Expected behavior

Timestamps are parsed in the CSV timeZone option, as Spark does. Otherwise, fall back when the schema has timestamp columns and the timezone isn't UTC.

Additional context

The config is testing-only and off by default, so this is low priority for now. It needs fixing before the scan is enabled for real workloads.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions