Skip to content

FalkorDB: design pass for where predicate push-down (property flattening / dual-write) #51

Description

@yoheinakajima

Problem

where predicates are the one read surface that still always evaluates in Python. The v1.2 query hooks (#43/#45) pushed structural filters — type scans, relation lookups, neighborhoods, whole pattern chains — down into FalkorDB, so their cost now scales with result size. But Graph.objects(where=...), View object filters, and pattern WHERE clauses evaluate over the pushed-down candidate set in Python, because the structured payload is stored as a single JSON data property on the node. Cypher can't index or compare into it, so a query like

graph.objects(type="claim", where={"confidence": {">": 0.8}})

pushes type="claim" down but then fetches and JSON-decodes every claim to keep the few above 0.8. On a projection with many objects of the filtered type, cost scales with the type population, not the result — the same shape of gap #43 closed for structural queries.

This was deliberately scoped out twice: #43's out-of-scope note named "native flattening/dual-write for where push-down" as a separate future effort, and CONTRACT v1.2 #4 locked the split rule (structural filters push down; the where language stays in Python) while naming this design pass as the follow-on. This issue is that follow-on. Because it deliberately revisits a locked rule, it is design-first: no implementation lands before the questions below are answered and locked in a CONTRACT amendment.

Current workaround

Narrow the candidate set with what already pushes down (type=, objects_in_types, relation scoping, pattern chains) and accept the Python filter over the remainder. This is correct and often fine — the v1.2 hooks mean the Python side only sees candidates, not the whole graph — but it leaves indexable comparisons (confidence > 0.8, status == "open") unindexed on exactly the backend adopters chose for scale.

Proposed approach (the design questions to answer)

1. Which subset of the where language translates faithfully?
The language is small: dotted paths, bare literals meaning equality, and {op: value} with ops > < >= <= == != in / not in (plus, in pattern WHERE, AND / NOT / NOT EXISTS composition). A faithful translation must preserve some deliberately Python-flavored semantics:

  • missing path ⇒ None, and ordered comparisons on None are False (Cypher's NULL propagation happens to align, but needs conformance coverage, not assumption);
  • in / not in are Python membership over strings and lists — Cypher IN doesn't do substring matching, so string-membership either translates to CONTAINS (strings only) or stays in Python;
  • equality across JSON types (numbers vs strings, nested dicts/lists) must match Python == exactly — nested dict/list equality likely stays in Python.

The plausible v1-subset: top-level scalar fields with == != > < >= <= and list-membership in. Everything outside the subset stays in Python over the pushed-down remainder — partial push-down (translate what translates, post-filter the rest) rather than all-or-nothing.

2. Opt-in per-type vs global flattening? Flattening writes data keys as native node properties so Cypher can index and compare them. Global flattening is simple but pays write cost and property-key sprawl for every type, including ones nobody filters; per-type opt-in (e.g. the pack's ObjectType declaring indexable fields, or a store-level flatten={"claim": ["confidence", "status"]}) keeps cost where the value is, at the price of a config surface and asymmetric query plans. Leaning per-type opt-in, but this is exactly the decision the amendment must lock.

3. What does dual-write cost, and who pays it? Flattened properties are a second copy of (part of) data: every object.created / patch.applied write updates both. Costs to quantify in the design pass: write amplification per event, property-count limits and index maintenance in FalkorDB, collision rules when a data key shadows a structural property (id, type are reserved), and re-flattening on config change. The projection being disposable (CONTRACT v1.2 #1) is the escape hatch: changing the flattening config never migrates data — re-replay the log into the new shape, same as the v1.2 #3 layout change.

4. What does GraphStoreConformance assert? The v1.2 #4/#5 rule stands: Python defaults define semantics; a backend override MUST return identical results. The suite would grow a where-parity section: for each operator × (present / missing / None-valued / type-mismatched path) × (flattened / unflattened type), the pushed-down result set equals the Python evaluator's, byte for byte. The subset boundary is also asserted: predicates outside the translatable subset must still be answered correctly (via the post-filter path), never silently dropped.

Out of scope

  • Any implementation before the four answers above are locked in a CONTRACT amendment (the v1.2 Build v0.7 tools + Cypher pattern subscriptions #4 split rule is being deliberately revisited, so it needs its own amendment, not a silent in-PR change).
  • Pattern-WHERE composition (AND / NOT / NOT EXISTS) push-down — the object-query subset comes first; pattern WHERE inherits whatever subset is locked.
  • Other backends. The design should be expressible through GraphStore hooks so Neo4j/Postgres-graph backends (CONTRACT v1.2 Build v0.8 Postgres + observability + operator CLI #5) can adopt it, but FalkorDB is the reference implementation.
  • Changing the in-memory default's behavior in any way.

Provenance: @dudizimber drove the v1.2 GraphStore arc end-to-end (#38/#41/#43/#45 → PRs #39/#46); this is that arc's named follow-on (CONTRACT v1.2 #4). Filed in the same problem → workaround → approach → out-of-scope shape as that series.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    design/rfcDesign-first work; no implementation before the decision recordenhancementNew feature or requestperformanceMeasured performance or scale workstatus: ready-for-prShape agreed; implementation contribution welcome

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions