Skip to content

Performance Optimization for Large Grids (Vectorization & Caching) - #66

Merged
marota merged 6 commits into
mainfrom
perf/vectorization-optim-v1
Mar 29, 2026
Merged

marota merged 6 commits into
mainfrom
perf/vectorization-optim-v1

Conversation

@marota

@marota marota commented Mar 29, 2026 •

Copy link
Copy Markdown
Collaborator

PR: Performance Optimization for Large Grids (Vectorization & Caching)

Summary

This PR implements critical backend performance optimizations for ExpertAssist, focusing on the observation and simulation pipeline. For large grids (e.g., France 10k+ branches), these changes provide a 4x speedup in manual action simulation latency.

Key Changes

1. Vectorized Simulation Pipeline (RecommenderService)

Replaced high-overhead Python loops with NumPy vectorized operations in several key areas:

  • care_mask & Overload Detection: Achieved a 1,100x speedup (12.17s -> 0.01s) by eliminating per-branch attribute lookups and rho-balancing loops.
  • Branch Flow Extraction (_get_network_flows): Optimized result extraction from pypowsybl networks, resulting in a 13x speedup.
  • Flow Delta Computation (_compute_deltas): Vectorized the terminal-aware delta logic, providing a 47x speedup.

2. Observation Caching

Implemented an internal cache for get_obs() results during the manual action simulation loop. This eliminates redundant data retrieval and property access overhead, improving the simulation body latency by another ~600ms.

3. SVG Boosting (Frontend)

Maintained and refined the "Phase 1" SVG boosting utility in standalone_interface.html. This ensures that labels, nodes, and flow arrows remain legible on large grids by dynamically scaling them based on the diagram's native resolution.

Optimization Investigations (NAD Reduction)

As part of this work, we extensively investigated dynamic SVG reduction strategies:

  • Strategy 1 (Voltage Filtering): Evaluated but discarded due to insufficient payload reduction (~11%).
  • Strategy 3 (Viewport-Based Subsets): Implemented as a prototype achieving 50x payload reduction. However, it was ultimately discarded and reverted in this PR due to complexities in coordinate synchronization and the loss of global highlighting integrity (overloads/impacts).
    Detailed findings on these investigations are documented in docs/nad_optimization.md.

Performance Benchmarks

Metric Before After Improvement
Core Mask Loop 12.17s 0.01s 1,100x
Flow Extraction 0.82s 0.06s 13x
Delta Computation 0.47s 0.01s 47x
Total Simulation Latency ~16.5s ~4.0s 4x

Verification

Verification was performed using scripts/profile_diagram_perf.py, which measures backend latency across N, N-1, and Manual Action scenarios on the full French grid.

Test Consolidation & Performance Verification

The performance optimization logic is now fully covered by a consolidated and expanded test suite.

  • Passing Tests: All 241 backend tests are passing in the CI environment.
  • Improved Coverage:
    • `test_vectorized_monitoring.py`: Dedicated validation of NumPy-based mask/threshold logic and legacy mock coercion.
    • `test_cache_synchronization.py`: Verification of observation cache invalidation during contingency/variant changes.
    • `test_performance_budgets.py`: Automated SLA checks ensuring the logic layer remains under budget.
  • Performance Budget:
    • Grid Size: 2,000 lines (Simulated Large Network).
    • Logic Latency: ~6.5ms (Logic Budget: 50ms).
  • Maintainability: Removed 200+ lines of redundant, mock-heavy tests from `test_recommender_service.py` in favor of specialized logic tests.
    EOF

@marota
marota merged commit d53acf4 into main Mar 29, 2026
2 checks passed
marota pushed a commit that referenced this pull request Apr 30, 2026
Reconciliation of section 2 (0.5.0):
- Drop misattributed PRs that were actually pre-rebrand: save/reload
  (#49/#52), MW Start (#62), interaction-logging (#64), SLD highlights
  (#63), load shedding initial integration (#61). All now properly
  cited in section 1.5–1.6.
- Disambiguate App.tsx refactor history: PR #56 (hooks, 2100 → 800,
  pre-rebrand) vs PR #74 (components, 1000 → 650, 0.5.0) vs PR #75
  (memoization Phase 2, same LoC).
- Add accurate 0.5.0 PRs: #66 (vectorization w/ benchmark table),
  #69/#70/#71 (UI polish), #72 (curtailment), #73 (loads_p/gens_p
  format + configurable MW), #74/#75 (App.tsx decomposition),
  #78 (PST tap re-simulation), #84/#86/#87/#90 (detachable tabs).
- Add a recap table summarizing what's truly new in 0.5.0.

Diagrams added (Mermaid, GitHub-rendered):
- Gantt timeline of all 4 phases (top of doc).
- High-level architecture (frontend / backend / data).
- Two-step analysis sequence diagram (section 1.6).
- App.tsx LoC evolution flow (section 2.4).
- Backend mixin decomposition before/after PR #104/#106 (section 3).
- PyPSA-EUR pipeline flowchart (section 4).

https://claude.ai/code/session_01Pg7fuCUG2edfm5PyHS6SbN
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant