I build reinforcement-learning environments and evaluation harnesses that train and measure frontier models, spanning agentic coding agents, tool use, and reasoning. The environments pose real, verifiable tasks in sandboxed execution with automated reward and verification (RLHF and reinforcement learning from verifiable rewards), so models can be trained and benchmarked at scale against ground truth.
The tasks run deep into ML systems and GPU work: CUDA, Triton and Pallas kernels, distributed training (FSDP, ZeRO, NCCL), optimizer internals, and high-throughput inference (vLLM), alongside full-stack software engineering and infrastructure.
Applied machine learning: ICL Fit
I also build production machine-learning systems for surgery. ICL Fit, which I co-founded and engineer, predicts the best implantable-lens size and post-operative vault for an individual eye. It learns from anterior-segment imaging, biometry, and real surgical outcomes, combining computer-vision and tabular models with feature engineering, model calibration, and uncertainty estimation, and it is deployed as a tool surgeons use in clinic.
Medical-AI evaluation
As a physician, I focus on the part where my two fields meet: evaluating medical AI. I build rubric-based health benchmarks in which frontier models are graded against physician-written criteria for accuracy, completeness, safety, and communication, with model-based graders validated against physician judgment. It is the same problem as the clinic: measuring performance honestly against the standard of an expert.