Skip to content

feat: encode the remaining types Pydantic's JSON mode knows - #1966

Closed
shcheklein wants to merge 1 commit into
mainfrom
feat/json-encoder-extra-types
Closed

shcheklein wants to merge 1 commit into
mainfrom
feat/json-encoder-extra-types

Conversation

@shcheklein

Copy link
Copy Markdown
Contributor

A model field holding one of several ordinary types cannot be stored:

class Doc(dc.DataModel):
    source: pathlib.PurePosixPath
    ttl: datetime.timedelta

class Batch(dc.DataModel):
    docs: list[Doc]

dc.read_values(b=[Batch(docs=[Doc(source=..., ttl=...)])]).save("b")
# TypeError: Object of type PurePosixPath is not JSON serializable

Why it depends on where the model sits

Two things can turn a model into storable JSON, and they know different types:

converts with knows
flatten model_dump() — python mode leaves Python objects for this encoder
warehouse model_dump(mode="json") Pydantic converts everything it knows

Which one runs is decided by the container:

list[Model]        flatten     -> python mode -> this encoder
tuple[Model, ...]  warehouse   -> JSON mode   -> Pydantic

So the same field works in one shape and fails in the other. Measured, for a declared field type:

                json mode (warehouse)   python mode + this encoder
timedelta       'PT1M30S'               TypeError
PurePosixPath   '/a/b'                  TypeError
IPv4Address     '1.2.3.4'               TypeError
plain enum      1                       TypeError
set[int]        [1, 2]                  TypeError
datetime/UUID/str enum                  same
ndarray         PydanticSerializationError   [1, 2]

This PR closes five of those rows, plus IPv6Address. Each is written the way Pydantic's JSON mode writes it, and the tests assert that equivalence against model_dump(mode="json") rather than against a fixed string, so they stay honest if Pydantic changes.

Scope

All six raise today, so nothing that currently writes changes. Eight of the nine new cases fail against main.

Decimal is deliberately left out: it already writes, as a number where Pydantic writes a string (1.2 vs "1.20"), so aligning it would change stored data rather than add to it. That belongs in its own change.

This is a prerequisite for making the two converters agree — with the type gap closed, which one runs stops changing what gets written.

datachain.json teaches ujson the types it cannot write on its own, and covered
datetime, date, time, UUID, numpy and bytes. Pydantic's JSON mode covers more
than that, so a value that arrives still holding its Python type has nowhere to
go and the write fails:

    TypeError: Object of type PurePosixPath is not JSON serializable

Which of the two converts a model decides whether that happens, and that is
settled by the shape the model sits in rather than by anything about the value:
a model inside a list is dumped in Pydantic's python mode, leaving its fields as
Python objects for this encoder, while the same model inside a tuple reaches the
warehouse live and is dumped in JSON mode instead.

Adds timedelta, PurePath, IPv4Address, IPv6Address, plain Enum and set, each
written the way Pydantic's JSON mode writes it, asserted against model_dump for
every one rather than against a fixed string. All six raise today, so nothing
that currently writes changes.

Decimal is left alone: it already writes, as a number where Pydantic writes a
string, so aligning it would change stored data rather than add to it.
@cloudflare-workers-and-pages

Copy link
Copy Markdown

Deploying datachain with  Cloudflare Pages  Cloudflare Pages

Latest commit: 44a4199
Status: ✅  Deploy successful!
Preview URL: https://e45487c1.datachain-2g6.pages.dev
Branch Preview URL: https://feat-json-encoder-extra-type.datachain-2g6.pages.dev

View logs

@codecov

codecov Bot commented Aug 30, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@shcheklein

Copy link
Copy Markdown
Contributor Author

Closing — superseded by a different direction.

This taught the encoder the types Pydantic's JSON mode already knows, so that python mode could be used everywhere. We are going the other way instead: use Pydantic's JSON mode everywhere, which covers those types natively and makes this unnecessary.

What that direction needs instead is numpy normalized before the dump, since Pydantic refuses ndarray where this encoder handles it. Measured, that costs about 1 ms on top of the tolist() that has to happen anyway.

Keeping the measurement here for whoever picks it up — declared field types, JSON mode versus python mode plus this encoder:

json mode python mode + encoder
timedelta 'PT1M30S' TypeError
PurePosixPath '/a/b' TypeError
IPv4Address '1.2.3.4' TypeError
plain enum 1 TypeError
set[int] [1, 2] TypeError
Decimal '1.20' 1.2
ndarray PydanticSerializationError [1, 2]

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant