You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
I audited how Comet handles timezones, starting from #2730. The model is simple and mostly sound. Spark's TimestampType is a UTC instant, so nothing is converted at the JVM/native boundary. Comet passes the raw microseconds in both directions and labels them Timestamp(Microsecond, "UTC"). TimestampNTZType is Timestamp(Microsecond, None). The session timezone never becomes part of a value. Each timezone-aware expression carries it, and it's applied inside the native kernel, or the expression runs through the codegen dispatcher with Spark's own timeZoneId. #6337 adds a contributor guide page that describes the model in more detail.
The bugs cluster where that model breaks down:
native expressions that emit a TimestampType value with some other label
session timezone IDs that the native parser can't read
timezone rules that come from a different database than the JVM's
Label drift is easy to miss. The scan and shuffle boundaries cast every column back to its declared type, so a test that only projects the result passes. It shows up when the result is compared, goes through a CASE, or feeds another native expression.
All of the new bugs below reproduce on main at 764936187, on Spark 3.5 and 4.1.
Six of the eight bugs are fixed on main. #6351 also replaced every getOrElse("UTC") fallback with CometTimeZone.nativeId, and #6347 turned the asserts in array_with_timezone into errors, so the section on #2730 below describes the code before those changes.
I instrumented every timeZoneId.getOrElse("UTC") site. Then I ran the datetime, cast, SQL-file, JSON, CSV, fuzz and expression suites. About 11,700 serde calls happened across 1,020 tests, and about 900 of them arrived without a timezone. Every one of those was a cast that Spark doesn't consider timezone-sensitive: numeric casts, Comet's own nullability-widening casts, and the cast inside IntegralDivide. None was a timezone-aware expression, which fits Spark refusing to resolve one without a timezone. So the fallback isn't a correctness bug today. The helper proposed in #6329 would replace it. The fallback can't simply be removed, though, because array_with_timezone asserts a non-empty timezone even for casts that don't use one.
Already documented: Python Arrow UDFs see timestamps labelled UTC rather than the session timezone, and chrono-tz's DST rules end around 2100. Not in the user guide yet: spark.sql.parquet.int96TimestampConversion=true disables Comet for the session.
Fixed earlier, same class: #2720 (SparkToColumnar labelled timestamps with the session timezone), #2649 via #4761 (the date_trunc schema mismatch, whose fix introduced the label in #6330), and #5556 (the Python runner accepts Etc/UTC for UTC).
Test gaps
The SQL-file tests use UTC, America/Los_Angeles, America/New_York, Asia/Kolkata and +05:30. None of them use Etc/UTC, or the offset and short-ID forms from #6329. Most expression tests only project their result. Adding GMT+8 to the datetime files' ConfigMatrix would have caught #6329. #6328, #6330 and #6327 need more than a timezone setting: a test that compares each native timestamp-returning expression with another timestamp, or uses it in a CASE, in both a non-UTC session and Etc/UTC.
What / Why
I audited how Comet handles timezones, starting from #2730. The model is simple and mostly sound. Spark's
TimestampTypeis a UTC instant, so nothing is converted at the JVM/native boundary. Comet passes the raw microseconds in both directions and labels themTimestamp(Microsecond, "UTC").TimestampNTZTypeisTimestamp(Microsecond, None). The session timezone never becomes part of a value. Each timezone-aware expression carries it, and it's applied inside the native kernel, or the expression runs through the codegen dispatcher with Spark's owntimeZoneId. #6337 adds a contributor guide page that describes the model in more detail.The bugs cluster where that model breaks down:
TimestampTypevalue with some other labelLabel drift is easy to miss. The scan and shuffle boundaries cast every column back to its declared type, so a test that only projects the result passes. It shows up when the result is compared, goes through a
CASE, or feeds another native expression.All of the new bugs below reproduce on
mainat764936187, on Spark 3.5 and 4.1.Bugs
timestamp_secondsreturns a TIMESTAMP_NTZ-typed array, sohour, casts and comparisons over it are wrong or fail in non-UTC sessions (critical). Fixed by fix: label native timestamp_seconds results as UTC timestamps #6341.date_trunclabels its output with the session timezone, so comparing it with another timestamp fails inEtc/UTCsessions, the default on Ubuntu and Debian images (high). Fixed by fix: keep the input's timezone label on native date_trunc results #6345.CASEandCOALESCEover timestamps panic when the branches have different Arrow timezones (high). Fixed by fix: relabel CASE and COALESCE timestamp branches instead of panicking #6347.GMT+8,ZandPSTmake casts,houranddf.show()fail (high). Fixed by fix: normalize session timezone IDs before passing them to native code #6351.timestamp_truncpanics on DST-transition timestamps in a DST timezone. Native only withallowIncompatible(high). feat: Use DataFusion date_trunc for scalar-format timestamp truncation #5956 fixes literal formats, and fix: truncate timestamps around DST transitions the way Spark does #6354 will cover per-row formats after it lands.daysis evaluated in the session timezone, whilehoursand Iceberg use UTC (low). fix: count the days partition transform in UTC #6348 is approved, and CI, including the Iceberg suites, is running.Status (2026-10-05)
Six of the eight bugs are fixed on
main. #6351 also replaced everygetOrElse("UTC")fallback withCometTimeZone.nativeId, and #6347 turned the asserts inarray_with_timezoneinto errors, so the section on #2730 below describes the code before those changes.Still open besides #5633 and #6333:
Etc/UTCand a non-UTC session.timestamp_seconds(fix: label native timestamp_seconds results as UTC timestamps #6341) anddate_trunc(fix: keep the input's timezone label on native date_trunc results #6345) got tests that compare their results, andsession_timezone_ids.sql(fix: normalize session timezone IDs before passing them to native code #6351) covers theGMT+8-style IDs.The UTC fallbacks in #2730
I instrumented every
timeZoneId.getOrElse("UTC")site. Then I ran the datetime, cast, SQL-file, JSON, CSV, fuzz and expression suites. About 11,700 serde calls happened across 1,020 tests, and about 900 of them arrived without a timezone. Every one of those was a cast that Spark doesn't consider timezone-sensitive: numeric casts, Comet's own nullability-widening casts, and the cast insideIntegralDivide. None was a timezone-aware expression, which fits Spark refusing to resolve one without a timezone. So the fallback isn't a correctness bug today. The helper proposed in #6329 would replace it. The fallback can't simply be removed, though, becausearray_with_timezoneasserts a non-empty timezone even for casts that don't use one.Related
spark.sql.session.timeZoneeverywhereAlready documented: Python Arrow UDFs see timestamps labelled
UTCrather than the session timezone, and chrono-tz's DST rules end around 2100. Not in the user guide yet:spark.sql.parquet.int96TimestampConversion=truedisables Comet for the session.Fixed earlier, same class: #2720 (
SparkToColumnarlabelled timestamps with the session timezone), #2649 via #4761 (thedate_truncschema mismatch, whose fix introduced the label in #6330), and #5556 (the Python runner acceptsEtc/UTCforUTC).Test gaps
The SQL-file tests use
UTC,America/Los_Angeles,America/New_York,Asia/Kolkataand+05:30. None of them useEtc/UTC, or the offset and short-ID forms from #6329. Most expression tests only project their result. AddingGMT+8to the datetime files'ConfigMatrixwould have caught #6329. #6328, #6330 and #6327 need more than a timezone setting: a test that compares each native timestamp-returning expression with another timestamp, or uses it in aCASE, in both a non-UTC session andEtc/UTC.