What is the problem the feature request solves?
Two dictionary-handling paths hand-roll what arrow's cast kernel supports natively:
native/spark-expr/src/conversion_funcs/cast.rs dict_from_values (explicitly "copied from datafusion common scalar/mod.rs") manufactures a degenerate dictionary with one key per row and no deduplication, then the cast recurses into it.
native/core/src/parquet/parquet_support.rs parquet_convert_array has a branch that manually unpacks Dictionary(Int32, Utf8/LargeUtf8) via recursive value conversion plus take, or rebuilds a dictionary by hand.
Describe the potential solution
- In
cast.rs, invert the order: Spark-cast the plain array to the dictionary's value type first, then arrow::compute::cast to Dictionary(key, value). Arrow's cast_to_dictionary handles key creation and dedups values via dictionary builders. The Spark-owned semantics stay in the value cast.
- In
parquet_support.rs, the dictionary branch can likely be deleted: since the value type is restricted to Utf8/LargeUtf8, the recursion can only hit identity or the generic can_cast_types fallthrough, and arrow's cast_with_options natively supports both Dictionary -> T (unpack via take) and Dictionary -> Dictionary (cast values).
Additional context
- Output representation differs in
cast.rs (deduplicated dictionary instead of one-entry-per-row), which is logically the same array; downstream consumers treat dictionaries generically.
- Before deleting the parquet branch, verify
can_cast_types(Dictionary(..), to) covers the same combinations, since unsupported combos currently fall through to returning the original array.
Found during an audit of native code that replicates existing arrow-rs kernels.
What is the problem the feature request solves?
Two dictionary-handling paths hand-roll what arrow's cast kernel supports natively:
native/spark-expr/src/conversion_funcs/cast.rsdict_from_values(explicitly "copied from datafusion common scalar/mod.rs") manufactures a degenerate dictionary with one key per row and no deduplication, then the cast recurses into it.native/core/src/parquet/parquet_support.rsparquet_convert_arrayhas a branch that manually unpacksDictionary(Int32, Utf8/LargeUtf8)via recursive value conversion plustake, or rebuilds a dictionary by hand.Describe the potential solution
cast.rs, invert the order: Spark-cast the plain array to the dictionary's value type first, thenarrow::compute::casttoDictionary(key, value). Arrow'scast_to_dictionaryhandles key creation and dedups values via dictionary builders. The Spark-owned semantics stay in the value cast.parquet_support.rs, the dictionary branch can likely be deleted: since the value type is restricted to Utf8/LargeUtf8, the recursion can only hit identity or the genericcan_cast_typesfallthrough, and arrow'scast_with_optionsnatively supports bothDictionary -> T(unpack via take) andDictionary -> Dictionary(cast values).Additional context
cast.rs(deduplicated dictionary instead of one-entry-per-row), which is logically the same array; downstream consumers treat dictionaries generically.can_cast_types(Dictionary(..), to)covers the same combinations, since unsupported combos currently fall through to returning the original array.Found during an audit of native code that replicates existing arrow-rs kernels.