Repository navigation
fix: fall back from the native CSV scan for timestamps outside UTC - #6349
Conversation
The native CSV V2 reader hands arrow-csv the Spark schema, whose TimestampType is labelled UTC, so a timestamp without an offset is parsed as UTC. Spark parses it in the CSV timeZone option, which defaults to the session timezone. Fall back when the read schema has a timestamp column and that timezone is not UTC. Closes apache#6332.
sunchao
left a comment
There was a problem hiding this comment.
Summary
- Prior state and problem: Native CSV scans interpreted offset-free timestamps as UTC, producing incorrect instants when Spark’s effective CSV timezone was non-UTC.
- Design approach: Fall back to Spark when
readDataSchemacontainsTimestampTypeand the effective timezone does not normalize to UTC. - Correctness / compatibility analysis: Verified timezone precedence and parsing against Spark 3.4.3, 3.5.9, 4.0.4, 4.1.3, and 4.2.0 sources. The rule matches Spark’s case-insensitive
timeZoneoption and session default. A standalone Spark 4.1.3 parser probe passed 15 checks covering overrides, UTC aliases, and timezone-independentTimestampNTZTypeparsing. - Key design decisions: Uses existing fallback reporting without adding native or protocol abstractions. Additional work is confined to scan planning: a schema walk and, when needed, timezone resolution. No per-row processing is added.
- Implementation sketch: Extends the CSV branch of
CometScanRuleand adds a regression test comparing results and execution paths under four timezone configurations. - Behavioral changes worth calling out: Non-UTC timestamp reads use Spark. UTC aliases and
TimestampNTZTyperetain their existing eligibility. The native CSV feature remains testing-only and disabled by default. - Suggested improvements: None at P1/P2. No introduced P1/P2 issues found within this review.
Reviewed the entire two-file diff from e1d2c11729c2fc60a5def4e87bb17e5b28df2a29 to 603b1aa44adf0c8adabb9d4ea092fc5e92d51349. The PR is non-draft. The snapshot and live discussion contained no existing reviews or comments. Routed skill: review-comet-pr; no sibling review skill applies.
Exact-head CI: Spark 3.4/3.5/4.0 JVM compile/lint checks and the Spark 4.1 build passed. No failed checks were reported. Native build and Rust tests remained in progress, with no Comet CSV suite result available yet. Spark SQL integration jobs were skipped.
Validation limits: The local probe exercises Spark’s parser, not Comet’s complete execution path. I did not build the native library or run CometCsvNativeReadSuite or the Spark SQL integration suite locally. End-to-end validation therefore remains outstanding.
Which issue does this PR close?
Closes #6332.
Rationale for this change
The native CSV V2 scan hands DataFusion's
CsvSourcethe Spark schema, whereTimestampTypeisTimestamp(Microsecond, "UTC"). As a result, arrow-csv reads a timestamp without an offset as UTC.Spark's CSV reader interprets that same value in the CSV
timeZoneoption, which defaults to the session timezone. Nothing passes that timezone to the native reader, because theCsvOptionsproto has no field for it. In a non-UTC session, every such value was therefore silently shifted. For example, inAmerica/Los_Angeles,2024-01-15 18:30:45came back as2024-01-15T18:30:45Zinstead of2024-01-16T02:30:45Z.What changes are included in this PR?
CometScanRulenow falls back to Spark when the read schema has aTimestampTypecolumn and the CSV timezone isn't UTC. The CSV timezone is thetimeZoneoption, or the session timezone when that option isn't set.TimestampNTZTypeis unaffected, because neither reader applies a timezone to it.UTC,Etc/UTC,Zand+00:00all stay native.CometCsvNativeReadSuitecovers four cases:timeZone=UTCstays nativetimeZone=Asia/Tokyofalls backThe scan is testing-only and off by default. Passing the timezone to the native reader would need a proto field and a native change to how timestamps are parsed, so this PR just makes the unsupported case fall back.
How are these changes tested?
CometCsvNativeReadSuitetests pass on Spark 4.1, including thespotlessandscalastylechecks.