MSc student @ MIPT & YSDA · Research Intern @ Yandex
Theoretical ML. Cost bounds for attention sinks and efficient attention, dynamics and stability of Transformer blocks, Bayesian deep learning.
- 📄 Sink vs. Diagonal Attention: Sharpened Cost Bounds and Comparison Regimes — AXIOM Workshop @ NeurIPS 2026
- 📄 Jacobian Analysis of a Recurrent Transformer Block — AINL 2026
- 📦 Bensemble — Bayesian deep learning library (JOSS, in review)
Applying to PhD programs in theoretical ML, start fall 2027.


