Skip to content

Bug triage results: 2026-09-14 #5922

Description

@andygrove

Triage pass over the open requires-triage queue, per the project Bug Triage Guide.

  • Date: 2026-09-14
  • Total issues processed: 77 (71 triaged, 6 skipped, 0 failed)
  • Type counts: 20 bugs, 51 enhancements
  • Priority counts applied: priority:critical 6, priority:high 3, priority:medium 10, priority:low 1
  • Guide: docs/source/contributor-guide/bug_triage.md

Labels have already been applied. A reviewer should spot-check the calls below and close this issue when satisfied; corrections should be made directly on the affected issue.

Notes on label availability and pre-existing labels:

Bugs

priority:critical

  • Native unix_timestamp returns incorrect results for negative fractional timestamps (#5896)
    • Area labels: area:expressions
    • Rationale: floor division returns a value one second lower than Spark for pre-epoch fractional timestamps with no error, which is the guide's "silent wrong results" case at step 1 of the decision tree.
  • Native Parquet scan multiplies rows for a struct with duplicate field names (#5783)
    • Area labels: area:scan, correctness, native_datafusion
    • Rationale: one output row per matching leaf instead of one per input row, with no error or warning, is silent wrong results; confirms the author's priority:critical.
  • size, arrays_zip, map_from_arrays and array_append return wrong answers for a nondeterministic child (#5781)
    • Area labels: area:expressions, correctness, array expressions, map expressions
    • Rationale: the child is serialized twice and evaluated independently, so size returns -1 where Spark returns 1 with no error, reachable in every configuration; silent wrong results per step 1.
  • Track nested-field statistics pruning and TIMESTAMP_MILLIS overflow parity in filtered scans (#5739)
  • revertToSpark erases CometIcebergWriteExec / CometNativeWriteExec because originalPlan is the node's own child (#5719)
    • Area labels: area:writer
    • Rationale: the write node disappears and a unary child is applied twice, so no data is written or the commit receives garbage rows, which is data corruption under step 1; note that it requires the opt-in spark.comet.exec.transitionRevert.enabled, so a reviewer may prefer priority:high on the "core path over experimental" principle.
  • array_distinct and array_union diverge from Spark on -0.0 for Spark versions without SPARK-54918 (#5701)
    • Area labels: area:expressions, correctness
    • Rationale: [0.0, 1.0] returned in place of Spark's [0.0, -0.0, 1.0] with no fallback is silent wrong results on 3.4, 3.5 and pre-fix 4.x patch releases; escalated from the author's priority:high (see escalations below).

priority:high

priority:medium

  • Native metrics from several plan instances in one task overwrite each other, so a coalesced scan reports only its last partition (#5879)
  • str_to_map does not honour mapKeyDedupPolicy=LAST_WIN (#5857)
    • Area labels: area:expressions
    • Rationale: Comet raises DUPLICATED_MAP_KEY where Spark succeeds, a visible functional failure with a workaround (disable the expression), so step 3 applies.
  • Compaction with native Iceberg writes leaves a dangled position delete and adds two data files where iceberg-java adds one (#5837)
    • Area labels: area:writer, area:Iceberg
    • Rationale: the rewrite produces the wrong file counts and leaves a dangling delete, a broken feature of the off-by-default native writer with the JVM writer as the workaround; step 3.
  • Native Iceberg write renders float/double partition values differently from iceberg-java, and fails with "File name too long" for large values (#5836)
    • Area labels: area:writer, area:Iceberg
    • Rationale: the write fails visibly for values outside [1e-3, 1e7) and the remaining divergence is cosmetic; broken feature on the experimental writer with a workaround, so step 3.
  • Struct-typed scalar subquery result takes the consuming projection off Comet (widened by Spark 4.2 MergeSubplans) (#5834)
    • Area labels: area:expressions, performance, spark 4.2
    • Rationale: an unnecessary fallback that cascades through the plan is a functional gap with correct results, the guide's "missing expression support" example for priority:medium; confirms the author's label.
  • Duplicate field ids inside a struct are not validated when the file schema equals the requested schema and no predicate is pushed (#5801)
    • Area labels: area:scan
    • Rationale: Comet returns rows where Spark raises the duplicate field id error, a missing validation limited to files with no key-value metadata; treated as a functional divergence rather than wrong values, see escalations.
  • Parquet Variant annotation validation is bypassed when reading as an ordinary struct (#5741)
    • Area labels: area:scan
    • Rationale: the reporter confirms an observed missing exception rather than corrupted values, so this is a functional divergence with correct data, step 3.
  • Iceberg scan falls back to Spark on IS NULL/IS NOT NULL over list/map columns (stale complex-type check) (#5731)
    • Area labels: area:scan
    • Rationale: a stale check declines scans iceberg-rust now handles, costing up to 65% on affected queries with results verified correct; a functional gap with a fallback, step 3.
  • Follow up on post-merge review of the native sequence kernel: perf crossover, unbounded per-batch allocation, and missing ceiling coverage (#5712)
    • Area labels: area:expressions, performance, test, documentation
    • Rationale: the native kernel is slower than Spark above roughly a thousand elements per row and can allocate up to 3.3 GB per batch outside the memory pool, a performance regression and a memory-safety gap with correct results; confirms the author's priority:medium.
  • Re-enable the two ignored CometAggregateSuite metric tests after the DataFusion 55 peak_mem_used change (#5703)
    • Area labels: area:aggregation, test
    • Rationale: two metric assertions are disabled after an upstream accounting change, a test failure per the type table; the author's priority:medium was left in place although the guide's default for test-only failures is priority:low.

priority:low

  • Signed-zero array test fixtures claim the literal case agrees with Spark when it does not (#5702)
    • Area labels: area:expressions, test
    • Rationale: misleading comments in test fixtures are a test-only, cosmetic issue, step 4; confirms the author's label.

Enhancements

  • Fuse OR-connected literal contains predicates on the same input column (#5912)
    • Area labels: area:expressions, performance
    • Rationale: a performance optimization proposal for a set of already-native predicates.
  • Support ARRAY in CometSparkToColumnarExec (#5911)
    • Area labels: area:ffi, array expressions
    • Rationale: new type support at the Spark-to-Comet conversion boundary.
  • Avoid temporary heap allocation when copying off-heap UTF8String into Arrow (#5910)
    • Area labels: area:ffi, performance
    • Rationale: a performance optimization of the row-to-Arrow writer.
  • Preserve dictionary encoding in Arrow IPC inputs to PyArrow UDFs (#5906)
  • perf: shuffle reader and writer performance review (native, JVM columnar, reader, Celeborn) (#5905)
    • Area labels: area:shuffle, performance
    • Rationale: a ranked list of performance improvements, not a defect report.
  • Reduce per-row sorting allocations in multi-entry map_sort normalization (#5900)
    • Area labels: area:expressions
    • Rationale: a performance optimization of map normalization with bit-identical output.
  • Publish nightly SNAPSHOT jars to repository.apache.org (#5899)
    • Area labels: area:ci
    • Rationale: a new scheduled workflow, a CI/tooling addition.
  • Support Iceberg tables on HDFS in the native scan (#5894)
    • Area labels: area:scan
    • Rationale: new storage scheme support in the native Iceberg scan, blocked on upstream iceberg-rust.
  • chore: Clear remaining Scala 2.13 warnings under the strict-warnings profile (#5893)
    • Area labels: area:ci
    • Rationale: build hygiene work extending the strict-warnings CI job to a 2.13 profile.
  • Support Spark-compatible duplicate Parquet field resolution (#5884)
  • Preserve primitive Iceberg pruning beside unsupported complex null conjuncts (#5883)
    • Area labels: area:scan
    • Rationale: the reporter states it is a performance limitation with correct results, so it is a pruning improvement.
  • Decide how the Delta Lake contrib jar is built, versioned, and published in releases (#5882)
    • Area labels: area:ci, documentation, question
    • Rationale: a release packaging and tooling decision, not a defect.
  • ci: next set of jobs to move from the PR tier to the merge queue tier (#5870)
    • Area labels: area:ci
    • Rationale: a CI tiering proposal.
  • Try and discourage use of git push --force (#5861)
    • Area labels: none
    • Rationale: a contributor guide documentation change.
  • Adding a New Operator guide omits the equals/hashCode requirement for Comet plan operators (#5832)
    • Area labels: documentation
    • Rationale: a documentation addition to the contributor guide.
  • Add a guard test that every CometNativeExec constructor parameter participates in equals (#5831)
    • Area labels: test
    • Rationale: a new guard test; the live omission it would flag on CometBroadcastExchangeExec is noted as not producing wrong results.
  • Reduce CI time (Comet is consuming 50% of DataFusion CI, which is a lot) (#5830)
    • Area labels: area:ci
    • Rationale: a CI cost reduction request.
  • Support Spark 4 EmptyRelationExec as a native input (#5819)
    • Area labels: none
    • Rationale: new operator support; no spark 4 label exists in the repository.
  • feat: support Spark HyperLogLog sketch functions (hll_sketch_agg, hll_union_agg, hll_sketch_estimate, hll_union) (#5814)
    • Area labels: area:aggregation, area:expressions
    • Rationale: new aggregate and scalar expression support.
  • Add Java-generated Rust parity fixtures for regex crate upgrades (#5813)
    • Area labels: area:expressions
    • Rationale: new test coverage for the regex whitelist.
  • Measure CometRegex whitelist coverage on representative regex patterns (#5812)
    • Area labels: area:expressions
    • Rationale: a measurement task to prioritize follow-up work.
  • Add grammar-based differential fuzzing for CometRegex (#5811)
    • Area labels: area:expressions
    • Rationale: new fuzz-test infrastructure.
  • shuffle_bench should disable parquet view types so its input matches what Comet produces (#5793)
    • Area labels: area:shuffle
    • Rationale: a benchmark fidelity fix in tooling that does not affect production behavior.
  • Reuse the decoded schema across shuffle blocks instead of re-parsing it per block (#5792)
    • Area labels: area:shuffle
    • Rationale: a performance optimization of the shuffle reader.
  • Read local shuffle blocks directly in native code instead of copying them through a JVM stream (#5791)
    • Area labels: area:shuffle
    • Rationale: a performance optimization of the shuffle reader.
  • Return native shuffle partition offsets over JNI instead of through a temp index file (#5790)
    • Area labels: area:shuffle
    • Rationale: a performance optimization of the shuffle writer.
  • Enable TopK dynamic filter pushdown into native Parquet scans (#5775)
    • Area labels: area:scan
    • Rationale: a new pushdown optimization.
  • Support dynamic filter pushdown from ungrouped MIN/MAX into native Parquet scans (#5774)
    • Area labels: area:scan, area:aggregation
    • Rationale: a new pushdown optimization.
  • perf: skip calendar reconstruction in datetime extraction (hour/minute/second, dayofweek/weekday) (#5770)
    • Area labels: area:expressions
    • Rationale: a performance optimization with bit-identical output.
  • Fuse the SerializeFromObject / MapElements / DeserializeToObject sandwich into a Comet projection instead of falling back (#5710)
    • Area labels: area:expressions, performance
    • Rationale: new operator coverage via the codegen dispatcher.
  • [EPIC] DataFusion 55 upgrade follow-ups (#5709)
    • Area labels: EPIC
    • Rationale: a tracking epic whose linked bugs are triaged individually.
  • Slice arrays before normalizing to avoid allocating over the whole backing buffer (#5708)
    • Area labels: area:expressions
    • Rationale: an allocation optimization; carries an author-applied priority:medium that the guide reserves for bugs.
  • Establish whether the DataFusion 55 TPC-H off-heap bump is real, and changelog it if so (#5704)
    • Area labels: area:aggregation
    • Rationale: an investigation and changelog task; carries an author-applied priority:medium that the guide reserves for bugs.
  • perf: coalesce small shuffle-read batches before native sort (#5686)
    • Area labels: area:shuffle
    • Rationale: a performance proposal, currently parked after an end-to-end A/B did not reproduce the win.
  • Investigate: why Comet shuffle files can be larger than Spark's, and whether byte-based block sizing would fix it (#5678)
    • Area labels: area:shuffle
    • Rationale: an investigation request into shuffle block sizing.
  • Document a capability table for the native Delta reader (#5661)
    • Area labels: area:scan
    • Rationale: a documentation addition.
  • Deepen v2-checkpoint regression coverage for the native Delta scan (#5660)
    • Area labels: area:scan
    • Rationale: additional test coverage.
  • Scope object store credentials to execution payloads and selected authorities in the native Delta scan (#5659)
    • Area labels: area:scan
    • Rationale: contract hardening with no demonstrated leak.
  • Move cloud compatibility gating into a shared helper beside the object store configuration code (#5658)
    • Area labels: area:scan
    • Rationale: a refactor.
  • Overlap DV bitmap and Parquet footer reads during native Delta scan preparation (#5657)
    • Area labels: area:scan
    • Rationale: a latency optimization.
  • Preserve compressed bitmap containers when decoding native-format deletion vectors (#5656)
    • Area labels: area:scan
    • Rationale: a decoding optimization.
  • Support splitting large DV-backed files in the native Delta scan (#5655)
    • Area labels: area:scan
    • Rationale: new parallelism support.
  • [EPIC] Production-quality native Iceberg writes (#5649)
    • Area labels: area:writer, area:Iceberg, EPIC
    • Rationale: a tracking epic whose linked bugs are triaged individually.
  • Account the native Iceberg writer's buffers in Comet's memory pool (#5648)
    • Area labels: area:writer, area:Iceberg
    • Rationale: a memory accounting audit and improvement.
  • Add a native Iceberg write benchmark (#5647)
    • Area labels: area:writer, area:Iceberg
    • Rationale: new benchmark tooling.
  • Failure-injection tests for native Iceberg writes (#5646)
    • Area labels: area:writer, area:Iceberg
    • Rationale: new test coverage.
  • Adopt a pin-bump policy for the iceberg-rust dependency (#5645)
    • Area labels: area:Iceberg
    • Rationale: a process and dependency policy.
  • Enable the split-operator plan and native Iceberg writes by default (#5644)
    • Area labels: area:writer, area:ci, area:Iceberg
    • Rationale: graduation criteria for a default flip.
  • Decide keep-or-lift for each remaining native Iceberg write eligibility restriction (#5643)
    • Area labels: area:writer, area:Iceberg
    • Rationale: a scoping decision on the writer's fallback surface.
  • Feed Spark's write path zero-copy row views instead of materializing UnsafeRow (#5625)
    • Area labels: performance
    • Rationale: a performance optimization of the JVM write path; area:writer covers the native writer, which this does not touch.
  • Arrow-based clients pay a redundant Arrow -> UnsafeRow -> Arrow round-trip (#5616)
    • Area labels: performance
    • Rationale: a performance proposal whose clean fix is upstream in Spark.

Escalations to consider

  • array_distinct and array_union diverge from Spark on -0.0 for Spark versions without SPARK-54918 (#5701)
    • Escalated from the author's priority:high to priority:critical: the guide's first decision-tree question is whether the bug can cause silent wrong results, and it does, with no error and no fallback.
  • Duplicate field ids inside a struct are not validated when the file schema equals the requested schema and no predicate is pushed (#5801)
    • Labelled priority:medium as a missing validation. If the positional read can attribute one leaf's values to the other field rather than simply succeeding where Spark errors, that is silent wrong results and it should move to priority:critical.
  • revertToSpark erases CometIcebergWriteExec / CometNativeWriteExec because originalPlan is the node's own child (#5719)
    • Labelled priority:critical on the data-corruption rule. It is only reachable behind the off-by-default spark.comet.exec.transitionRevert.enabled, so the reviewer may prefer priority:high under the "core path over experimental" principle.
  • Native metrics from several plan instances in one task overwrite each other (#5879)

Skipped — needs more info

  • Backport candidates for 1.0.x: triage of every PR merged since branch-1.0 was cut (#5815)
    • A release-management triage record rather than a bug or an enhancement; a reviewer should decide whether it needs a type label at all or whether requires-triage can simply be dropped.
  • Bug triage results: 2026-08-31 (#5601)
    • A previous triage summary awaiting reviewer close-out; not a bug or an enhancement.
  • Bug triage results: 2026-08-24 (#5454)
    • A previous triage summary awaiting reviewer close-out; note that pass applied no labels and its recommendations may still need applying.
  • Bug triage results: 2026-08-17 (#5385)
    • A previous triage summary awaiting reviewer close-out; not a bug or an enhancement.
  • Bug triage results: 2026-08-11 (#5325)
    • A previous triage summary awaiting reviewer close-out; not a bug or an enhancement.
  • Bug triage results: 2026-08-03 (#5231)
    • A previous triage summary awaiting reviewer close-out; not a bug or an enhancement.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions