Shuvom Sadhuka

newme.png

I am a PhD student at MIT CSAIL. I am grateful to be supported by the Hertz Fellowship and NSF GRFP and to be advised by Bonnie Berger.

I am interested in building reliable AI systems, especially during evaluation. Some representative projects include:

  1. E-valuator: we built a statistical method (using sequential hypothesis testing) to stop agent trajectories early when the agent is incorrect.
  2. SSME: we built a method to include both labeled and unlabeled samples during (simultaneous) evaluation of multiple ML models.

I am also especially interested in applications to societally-impactful problems such as healthcare. During my PhD, I have interned at Abridge on AI evals research with Alex Chouldechova and Michael Oberst and at Genentech on early stopping of agents with Hanchen Wang.

Prior to my PhD, I studied CS and Statistics at Harvard. I enjoy writing, and have written/blogged about fellowship applications, biomedical data privacy, AI evals, and old papers, among other things.

Quick links:

latest posts

Mar 30, 2026 Revenge of the Worst Case
Jan 22, 2026 Reading List
Feb 11, 2025 Measuring Entropy