What is the problem the feature request solves?
A Scala UDF with a boxed parameter, such as (x: java.lang.Long) => ..., converts each value through Spark's input encoder. The JIT leaves the deserializer's Projection.apply as a virtual call, so every row runs a projection into a GenericInternalRow. The codegen dispatcher compiles Spark's own ScalaUDF.doGenCode, so it inherits the same per-row conversion.
@mbutrovich measured it on #6697 (comment). On one 8192-row batch, the generated kernel for (x: java.lang.Long) costs 15.6 to 19.2 ns per row, against 1.9 to 2.0 for (x: Long). End to end over 4M rows at batch size 8192, the boxed UDF costs 70 ms above max(c) without the function, against 28 ms for the primitive one.
Describe the potential solution
Have the dispatcher's kernel convert a boxed primitive parameter directly: a null check and a boxing of the primitive value, with no encoder projection. It must do exactly what Spark's input encoder does for each type it handles, so it should start with the boxed primitives (java.lang.Long, Integer, Double and the rest), where that is easy to show, and leave other encoders as they are.
Additional context
The measurements and the benchmark source are in the comment linked above. Spark's whole-stage codegen pays the same cost, so this is an improvement over Spark rather than a regression in Comet.
What is the problem the feature request solves?
A Scala UDF with a boxed parameter, such as
(x: java.lang.Long) => ..., converts each value through Spark's input encoder. The JIT leaves the deserializer'sProjection.applyas a virtual call, so every row runs a projection into aGenericInternalRow. The codegen dispatcher compiles Spark's ownScalaUDF.doGenCode, so it inherits the same per-row conversion.@mbutrovich measured it on #6697 (comment). On one 8192-row batch, the generated kernel for
(x: java.lang.Long)costs 15.6 to 19.2 ns per row, against 1.9 to 2.0 for(x: Long). End to end over 4M rows at batch size 8192, the boxed UDF costs 70 ms abovemax(c)without the function, against 28 ms for the primitive one.Describe the potential solution
Have the dispatcher's kernel convert a boxed primitive parameter directly: a null check and a boxing of the primitive value, with no encoder projection. It must do exactly what Spark's input encoder does for each type it handles, so it should start with the boxed primitives (
java.lang.Long,Integer,Doubleand the rest), where that is easy to show, and leave other encoders as they are.Additional context
The measurements and the benchmark source are in the comment linked above. Spark's whole-stage codegen pays the same cost, so this is an improvement over Spark rather than a regression in Comet.