Describe the bug
With spark.comet.scan.csv.v2.enabled=true, build_csv_source (native/core/src/execution/operators/csv_scan.rs:55) hands DataFusion's CsvSource the Spark schema. In that schema TimestampType is Timestamp(Microsecond, "UTC"), so arrow-csv reads a timestamp without an offset as UTC. Spark's CSV reader interprets it in the timeZone option instead, which defaults to the session timezone. CometCsvNativeScanExec builds its CSVOptions with the session timezone (CometCsvNativeScanExec.scala:92), but the CsvOptions proto has no timezone field, so the timezone never reaches the native reader. In a non-UTC session every such value is silently shifted.
Steps to reproduce
On main at 764936187, with a part-0.csv containing:
id,ts
0,2024-01-15 18:30:45
spark.conf.set("spark.comet.scan.csv.v2.enabled", "true")
spark.conf.set("spark.sql.sources.useV1SourceList", "avro,json,kafka,orc,parquet,text")
spark.conf.set("spark.sql.session.timeZone", "America/Los_Angeles")
spark.read.schema("id INT, ts TIMESTAMP").option("header", "true").csv(dir).collect()
Spark returns the instant 2024-01-16T02:30:45Z and Comet returns 2024-01-15T18:30:45Z.
Expected behavior
Timestamps are parsed in the CSV timeZone option, as Spark does. Otherwise, fall back when the schema has timestamp columns and the timezone isn't UTC.
Additional context
The config is testing-only and off by default, so this is low priority for now. It needs fixing before the scan is enabled for real workloads.
Describe the bug
With
spark.comet.scan.csv.v2.enabled=true,build_csv_source(native/core/src/execution/operators/csv_scan.rs:55) hands DataFusion'sCsvSourcethe Spark schema. In that schemaTimestampTypeisTimestamp(Microsecond, "UTC"), so arrow-csv reads a timestamp without an offset as UTC. Spark's CSV reader interprets it in thetimeZoneoption instead, which defaults to the session timezone.CometCsvNativeScanExecbuilds itsCSVOptionswith the session timezone (CometCsvNativeScanExec.scala:92), but theCsvOptionsproto has no timezone field, so the timezone never reaches the native reader. In a non-UTC session every such value is silently shifted.Steps to reproduce
On
mainat764936187, with apart-0.csvcontaining:Spark returns the instant
2024-01-16T02:30:45Zand Comet returns2024-01-15T18:30:45Z.Expected behavior
Timestamps are parsed in the CSV
timeZoneoption, as Spark does. Otherwise, fall back when the schema has timestamp columns and the timezone isn't UTC.Additional context
The config is testing-only and off by default, so this is low priority for now. It needs fixing before the scan is enabled for real workloads.