Skip to content

perf: split keygen scratch from runtime arena to reduce persistent memory #22

Description

@cedoor

Problem

compute_arena_bytes sizes Context's persistent self.arena to the max across all operations, including keygen:

max(keygen_encrypt, keygen_prepare, fhe_prepare, encrypt, decrypt, eval_0..9)

keygen_encrypt (bdd_key_encrypt_sk_tmp_bytes) and keygen_prepare (prepare_bdd_key_tmp_bytes) are one-time costs — they only run during Context::keygen(). If either term dominates (likely, given that encrypting a BDD key involves hundreds of GGSW ciphertexts), every Context permanently carries scratch memory that no runtime operation ever needs.

Proposed fix

Allocate a temporary local arena inside keygen() sized only for the keygen phase, and remove both keygen terms from compute_arena_bytes so self.arena is sized for runtime only:

// Inside keygen()
let keygen_bytes = {
    let e = self.module.bdd_key_encrypt_sk_tmp_bytes(&self.params.bdd_layout);
    let p = self.module.prepare_bdd_key_tmp_bytes(&self.params.bdd_layout);
    e.max(p)
};
let mut keygen_arena = scratch::new_arena(keygen_bytes);

bdd_key.encrypt_sk(..., scratch::borrow(&mut keygen_arena));
bdd_key_prepared.prepare(..., scratch::borrow(&mut keygen_arena));
// keygen_arena dropped here

compute_arena_bytes then only considers runtime operations:

[fhe_prepare, encrypt, decrypt, eval].into_iter().max().unwrap_or(0)

Tradeoff

One extra heap allocation per keygen() call. Since keygen already takes seconds of CPU time, this allocation cost is noise.

Before implementing

Add a quick size probe (e.g. in a test or example) to confirm the keygen terms actually dominate. If the ratio is large (e.g. ≥ 2×), the split is worth it; if they're similar, the gain is marginal.

Related code

  • src/context.rs — compute_arena_bytes (line 617), Context::keygen (line 304)
  • src/scratch.rs — new_arena / borrow

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

performanceRelated to memory, CPU, or allocation optimizations

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions