Projects
- h100-inference-control-plane: async FP8 serving on a single H100. Throughput 680 to 3,400 tok/s at 64-way concurrency, p99 latency 4.5s to 1.2s.
- expense-bench: computer-use agent environment with dual graders. A final-state grader accepts 75.5% of runs vs 56.7% for a path-aware grader, so 1 in 4 apparent successes reward-hacked.
- per_core_lock_free_event_bus: C++ SPSC ring buffer with cache-line alignment. ~24M msgs/sec at 2.5µs p99.
Open source
- NVIDIA/cuCollections: allocator lifetime fix in storage deleters, lookup test coverage.
- NVIDIA/NemoClaw: sandbox readiness races, container secret exposure, CLI test determinism.
Previously SWE II at Intuit (AI Revenue Intelligence) and TPM intern at Google.



