Skip to content

Bug triage results: 2026-08-17 #5385

Description

@andygrove

Triage pass over the open requires-triage queue, per the project Bug Triage Guide.

  • Date: 2026-08-17
  • Total issues processed: 33 (31 triaged, 2 skipped, 0 failed)
  • Type counts: 8 bugs, 23 enhancements
  • Priority counts applied: priority:critical 2, priority:high 0, priority:medium 3, priority:low 3
  • Guide: docs/source/contributor-guide/bug_triage.md

Labels have already been applied. A reviewer should spot-check the calls below and close this issue when satisfied; corrections should be made directly on the affected issue.

Notes on label availability and pre-existing labels:

Bugs

priority:critical

  • to_time / try_to_time: native parser rejects 'T12' and '12:30:45.' which Spark accepts (#5366)
    • Area labels: area:expressions, correctness, temporal expressions
    • Rationale: try_to_time returns NULL where Spark returns a value, and fractional digits 7–9 are truncated into a different stored nanosecond value, both with no error surfaced — the guide's "silent wrong results" case at step 1 of the decision tree. Escalated from the reporter's priority:medium (see escalations below).
  • Boolean-to-decimal cast produces invalid Decimal128 when 10^scale does not fit precision (#5334)
    • Area labels: area:expressions
    • Rationale: cast_boolean_to_decimal returns Ok with an unscaled value outside the declared precision because with_precision_and_scale validates only type metadata, so Comet emits invalid decimal data as success where Spark errors or returns NULL — data corruption, priority:critical per the guide's first decision-tree question.

priority:medium

  • Report native child-operator spill metrics in Spark task metrics for unified shuffle plans (#5382)
  • Task input metrics are unreliable when a native block mixes a native scan with a JVM input (#5336)
    • Area labels: area:scan
    • Rationale: the completion listener is registered after the iterator's close listener, so on the busy-poll path the reported bytes_scanned is stale or zero; reporter's priority:medium confirmed — broken feature, no query-result impact.
  • NegativeExpr::get_properties reports child ordering and range unchanged (#5330)
    • Area labels: area:expressions
    • Rationale: with_order is an identity function as used, so -a is reported with the child's unflipped ordering and un-negated range; the defect is demonstrated at unit level with no end-to-end query yet, so priority:medium per step 3 rather than critical (see escalations below).

priority:low

  • Nine expression benchmark rows labelled "Comet" are measuring Spark (#5372)
    • Area labels: area:expressions, test, performance
    • Rationale: the deliverable is corrected benchmark reporting, which the guide's step 4 places at priority:low as a tooling/test-only issue; the underlying ShortType fallback may be more than that (see escalations below).
  • ANSI abs integer overflow errors use Arrow-style type names instead of Spark's (#5356)
    • Area labels: area:expressions
    • Rationale: the error class matches Spark (ARITHMETIC_OVERFLOW) and only the interpolated type name differs (Int64 vs long), so nothing is broken beyond user-visible message text — cosmetic, step 4.
  • Improve CI to run enough tests when dependency upgraded (#5333)

Enhancements

  • Reduce JNI round-trips in the unified memory pools (batching, hysteresis, cheaper call path) (#5383)
    • Area labels: area:ffi, performance
    • Rationale: a performance optimization of a working code path — chunked acquire with retained slack and hysteresis — which the guide's type table places under enhancement.
  • Make native contrib scan dispatch generic (remove per-contrib arms from core Rust) (#5379)
    • Area labels: area:scan
    • Rationale: source-level tidiness, replacing #[cfg]-gated per-contrib calls with a type_url -> handler registry; the issue itself notes a default build links zero contrib symbols, so nothing is broken.
  • Move contrib scan proto messages out of core operator.proto into contrib/<name>/proto (#5378)
    • Area labels: area:scan
    • Rationale: a refactor that needs new per-contrib proto build infrastructure; the messages are inert in core today and the relocation is invisible on the wire, so there is no defect.
  • Use end to end microbenchmarks to prevent performance regressions between releases (#5373)
    • Area labels: area:ci
    • Rationale: asks for benchmark documentation, a reproducible EC2 instance type, and eventual scheduled automation — new process and tooling, not a defect.
  • Investigate: scan + shuffle-write slower than Spark on deeply nested schemas (#5355)
    • Area labels: area:scan, area:shuffle, performance
    • Rationale: classified as an enhancement because it is an explicitly unverified collection of static-analysis optimization candidates ("no benchmark numbers yet") rather than a confirmed regression from a prior Comet release; if a benchmark confirms Comet is slower than Spark on a real workload, this should be re-triaged as a priority:medium performance bug.
  • Run replace natively by default when the search string is a non-empty literal (#5354)
    • Area labels: area:expressions, performance
    • Rationale: the current dispatcher path is correct; this asks to widen the native default for a provably safe subset, which is a performance optimization.
  • Native candidate assessment: upper/lower deferred on the lack of a per-batch defer-to-dispatcher mechanism (#5353)
    • Area labels: area:expressions, performance
    • Rationale: records a missing capability — plan-time-only native/dispatcher selection with no way to defer a single batch — as a blocker for a class of optimizations; nothing behaves incorrectly today.
  • Support Apache Celeborn remote shuffle service (#5352)
    • Area labels: area:shuffle
    • Rationale: new functionality that Comet does not have today, the guide's canonical enhancement case.
  • Run rlike natively by default for patterns that are provably Java-regex equivalent (#5351)
    • Area labels: area:expressions, performance
    • Rationale: proposes a plan-time pattern analyzer so a safe subset can default to the existing native kernel; the conservative current behavior is not a defect.
  • Implement sequence natively for integral types instead of JVM codegen dispatch (#5349)
    • Area labels: area:expressions, array expressions, performance
    • Rationale: the codegen dispatcher is correct; this asks for a native implementation to remove per-row allocations, so it is a performance optimization.
  • Implement unbase64 natively instead of JVM codegen dispatch (#5347)
    • Area labels: area:expressions, performance
    • Rationale: unbase64 is already Compatible() via codegen dispatch, so behavior matches Spark; the ask is a new native kernel for the decode direction.
  • Tag fallback reasons with structured values instead of pre-formatted strings (#5346)
    • Area labels: none (user experience)
    • Rationale: a refactor of the fallback-reason representation across ~205 call sites to make reasons dedupable and queryable — an improvement to existing working reporting.
  • Bound memory when a partition has many files (#5343)
  • Broaden storage-partitioned-join tests, then enable reportPartitioning by default (#5342)
    • Area labels: area:scan
    • Rationale: adds test coverage as the precondition for flipping a config default — new coverage plus a behavior change, no existing defect.
  • Skip the merge wrapper for single-file partitions (#5341)
    • Area labels: area:scan
    • Rationale: avoids a no-op SortPreservingMergeExec for single-file partitions; correct today, just slightly wasteful, so a performance optimization.
  • Keep read concurrency in the sorted merge path (#5340)
    • Area labels: area:scan
    • Rationale: asks to benchmark the many-small-files case and add read-ahead if it regresses — prospective optimization work, with the regression not yet established.
  • Support transformed sort orders (bucket / truncate / etc.) (#5339)
    • Area labels: area:scan
    • Rationale: extends ordering reporting from identity sort keys to Iceberg transforms; the current narrow scope is documented as deliberate and always correct.
  • Merge on sort keys even when they are not selected (#5338)
    • Area labels: area:scan
    • Rationale: widens when the sorted merge can be used by reading unprojected sort columns; skipping the merge today is a correct-but-conservative fallback.
  • Sort-aware Iceberg reads: per-partition streaming merge + report ordering and grouping (v1) (#5337)
  • Add mode to run Comet planning but execute with Spark, so users can assess potential Comet coverage (#5335)
    • Area labels: none
    • Rationale: proposes a new spark.comet.planOnly.enabled dry-run mode so users can estimate coverage without execution risk — new functionality and a new config.
  • Add a fallback-invariance suite: query outcome should not change when a compatible expression is forced from native to Spark (#5328)
    • Area labels: area:expressions
    • Rationale: new test infrastructure sweeping the 293-expression registry for an invariant, replacing four hand-picked per-expression tests; classified as an enhancement even though it is expected to surface bugs, because the ask itself is coverage.
  • Comet 1.1.0 Release (September) (#5327)
    • Area labels: none
    • Rationale: a release-planning tracking issue; enhancement is the closer of the two type labels since nothing is reported broken.
  • Could Rust 1.98 new floating point optimizations help Comet? (#5326)
    • Area labels: none
    • Rationale: an open question about adopting new compiler float-math optimizations — exploratory performance work, no defect described.

Escalations to consider

  • to_time / try_to_time: native parser rejects 'T12' and '12:30:45.' which Spark accepts (#5366)
    • Escalated from the reporter's priority:medium to priority:critical. The to_time cases raise a visible error, which alone would sit at priority:medium, but try_to_time returns NULL for input Spark parses successfully and case 3 stores a different nanosecond value — both silent, which the guide's step 1 and the "correctness over crashes" principle put at critical. Narrow surface (Spark 4.1 only, unusual literals) is the counter-argument, but the guide's own critical examples include comparably narrow cast edge cases. Revert to priority:medium if the reviewer judges case 3 unobservable and the try_to_time NULL acceptable.
  • NegativeExpr::get_properties reports child ordering and range unchanged (#5330)
    • Left at priority:medium, but this is the escalation trigger to watch: the wrong ordering is reported to the optimizer, so if a plan can be shown where DataFusion drops or keeps a sort based on it, the result is silently mis-ordered output and this becomes priority:critical. The issue currently demonstrates the defect only through an expression-property unit test.
  • Nine expression benchmark rows labelled "Comet" are measuring Spark (#5372)
    • Left at priority:low as a benchmark-reporting defect, but the issue notes that Comet does support short-to-int and short-to-long casts, so if the ShortType projection failing to be replaced reproduces outside CometCastNumericToNumericBenchmark, it is a real coverage gap affecting user queries and warrants priority:medium.
  • Improve CI to run enough tests when dependency upgraded (#5333)

Skipped — needs more info

  • Bug triage results: 2026-08-11 (#5325)
    • Not a bug report or a feature request but the summary issue from the previous triage pass, so neither type label applies and no priority is meaningful. requires-triage was left in place because the skill only removes it alongside a type label. The reviewer should close it once the 2026-08-11 calls have been spot-checked; it will otherwise reappear in every future pass.
  • Bug triage results: 2026-08-03 (#5231)
    • Same as above, for the 2026-08-03 pass. This one has now been open across three triage passes, so it is worth either closing it or deciding that triage summary issues should be exempted from the requires-triage auto-label.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions