Skip to content

Codegen dispatcher: run Spark's null guard for a Scala UDF inside the kernel #6704

Description

@andygrove

What is the problem the feature request solves?

Spark's HandleNullInputsForUDF rule wraps a Scala UDF with a primitive parameter over a nullable column as if(isnull(c), null, f(knownnotnull(c))), with one IsNull per such parameter joined by Or. Comet's serde sends only the ScalaUDF to the JVM codegen dispatcher, and the If runs natively as a DataFusion CASE. DataFusion evaluates a one-branch CASE with an ELSE by filtering the batch for each branch and merging the results, so the guard costs more than the call it protects.

@mbutrovich measured this on #6697 (comment). For SELECT max(f(c)) over 4M rows of a bigint column with a tenth of the rows null, at batch size 8192, about 13 ms of the 21 ms gap between a dispatched (x: Long) => x + 1 and a vectorized UDF is the guard. Putting the vectorized UDF under the same IF raises its cost from 7 ms to 20 ms.

Describe the potential solution

Recognize the exact shape HandleNullInputsForUDF produces, an If whose condition is IsNull checks on the UDF's arguments, whose true branch is a null literal, and whose false branch is the ScalaUDF with those arguments wrapped in KnownNotNull. Route the whole If through CometScalaUDF.emitJvmCodegenDispatch, so the null check becomes a branch in the kernel's loop instead of a CASE over the batch. Matching only that shape keeps the results unchanged.

Additional context

The measurements, the benchmark source and the analysis are in the comment linked above.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions