Skip to content

Codegen dispatcher: convert boxed primitive Scala UDF parameters without the input encoder #6706

Description

@andygrove

What is the problem the feature request solves?

A Scala UDF with a boxed parameter, such as (x: java.lang.Long) => ..., converts each value through Spark's input encoder. The JIT leaves the deserializer's Projection.apply as a virtual call, so every row runs a projection into a GenericInternalRow. The codegen dispatcher compiles Spark's own ScalaUDF.doGenCode, so it inherits the same per-row conversion.

@mbutrovich measured it on #6697 (comment). On one 8192-row batch, the generated kernel for (x: java.lang.Long) costs 15.6 to 19.2 ns per row, against 1.9 to 2.0 for (x: Long). End to end over 4M rows at batch size 8192, the boxed UDF costs 70 ms above max(c) without the function, against 28 ms for the primitive one.

Describe the potential solution

Have the dispatcher's kernel convert a boxed primitive parameter directly: a null check and a boxing of the primitive value, with no encoder projection. It must do exactly what Spark's input encoder does for each type it handles, so it should start with the boxed primitives (java.lang.Long, Integer, Double and the rest), where that is easy to show, and leave other encoders as they are.

Additional context

The measurements and the benchmark source are in the comment linked above. Spark's whole-stage codegen pays the same cost, so this is an improvement over Spark rather than a regression in Comet.

Activity

  1. self-assigned this
    on Oct 6, 2026
  2. added a commit that references this issue on Oct 6, 2026
    62c5d44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions