Skip to content

[EPIC] Timezone handling bugs #6335

Description

@andygrove

What / Why

I audited how Comet handles timezones, starting from #2730. The model is simple and mostly sound. Spark's TimestampType is a UTC instant, so nothing is converted at the JVM/native boundary. Comet passes the raw microseconds in both directions and labels them Timestamp(Microsecond, "UTC"). TimestampNTZType is Timestamp(Microsecond, None). The session timezone never becomes part of a value. Each timezone-aware expression carries it, and it's applied inside the native kernel, or the expression runs through the codegen dispatcher with Spark's own timeZoneId. #6337 adds a contributor guide page that describes the model in more detail.

The bugs cluster where that model breaks down:

  • native expressions that emit a TimestampType value with some other label
  • session timezone IDs that the native parser can't read
  • timezone rules that come from a different database than the JVM's

Label drift is easy to miss. The scan and shuffle boundaries cast every column back to its declared type, so a test that only projects the result passes. It shows up when the result is compared, goes through a CASE, or feeds another native expression.

All of the new bugs below reproduce on main at 764936187, on Spark 3.5 and 4.1.

Bugs

Status (2026-10-05)

Six of the eight bugs are fixed on main. #6351 also replaced every getOrElse("UTC") fallback with CometTimeZone.nativeId, and #6347 turned the asserts in array_with_timezone into errors, so the section on #2730 below describes the code before those changes.

Still open besides #5633 and #6333:

The UTC fallbacks in #2730

I instrumented every timeZoneId.getOrElse("UTC") site. Then I ran the datetime, cast, SQL-file, JSON, CSV, fuzz and expression suites. About 11,700 serde calls happened across 1,020 tests, and about 900 of them arrived without a timezone. Every one of those was a cast that Spark doesn't consider timezone-sensitive: numeric casts, Comet's own nullability-widening casts, and the cast inside IntegralDivide. None was a timezone-aware expression, which fits Spark refusing to resolve one without a timezone. So the fallback isn't a correctness bug today. The helper proposed in #6329 would replace it. The fallback can't simply be removed, though, because array_with_timezone asserts a non-empty timezone even for casts that don't use one.

Related

Already documented: Python Arrow UDFs see timestamps labelled UTC rather than the session timezone, and chrono-tz's DST rules end around 2100. Not in the user guide yet: spark.sql.parquet.int96TimestampConversion=true disables Comet for the session.

Fixed earlier, same class: #2720 (SparkToColumnar labelled timestamps with the session timezone), #2649 via #4761 (the date_trunc schema mismatch, whose fix introduced the label in #6330), and #5556 (the Python runner accepts Etc/UTC for UTC).

Test gaps

The SQL-file tests use UTC, America/Los_Angeles, America/New_York, Asia/Kolkata and +05:30. None of them use Etc/UTC, or the offset and short-ID forms from #6329. Most expression tests only project their result. Adding GMT+8 to the datetime files' ConfigMatrix would have caught #6329. #6328, #6330 and #6327 need more than a timezone setting: a test that compares each native timestamp-returning expression with another timestamp, or uses it in a CASE, in both a non-UTC session and Etc/UTC.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

EPICarea:expressionsExpression evaluationarea:scanParquet scan / data readingbugSomething isn't workingcorrectnesspriority:criticalData corruption, silent wrong results, security issues

Type

No type

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions