Computer vision and AI engineer based in Tokyo. I have been building vision systems since 1998 — through three fairly different technology generations — and I still write code every day.
HOUJUN Co., Ltd. (Tokyo) — founder and technical lead. houjun.dev
Medical signal processing — early work on automated analysis of biological signals, in collaboration with a university hospital. My background is in life sciences and clinical medicine, which is where this started.
Large-scale video infrastructure — real-time analysis across many concurrent camera streams: detection, tracking, and event extraction under hard latency and reliability constraints. Systems that had to keep running, not just demo well.
Robot vision and motion control — perception and control for industrial robots: calibration, pose estimation, and closing the loop between what the camera sees and what the arm does.
Generative AI and on-device inference — current focus. Quantization, local VLM/LLM serving, and the engineering trade-offs that show up when a model has to run on constrained hardware instead of a datacenter GPU.
Getting vision-language and vision-language-action models to run reliably outside the lab — quantization trade-offs, tail latency in closed control loops, and how these systems actually degrade when the sensor data stops being clean.
Most of what I know about that last part came from twenty years of deployed systems rather than from benchmarks.
Two iOS applications, designed and built end to end — from model and backend through to App Store release:
- Mind Craft Fish
- Cherry Tempo
Chinese (native) · Japanese (JLPT N2) · English (working)
Based in Tokyo. Open to conversations about computer vision, robotics perception, and edge inference roles.