Shuvom Sadhuka

newme.png

I am a PhD student at MIT CSAIL. I am grateful to be supported by the Hertz Fellowship and NSF GRFP and to be advised by Bonnie Berger.

I am interested in building reliable, safe, and aligned AI systems. I am particularly interested in evaluation. Some representative projects include:

  1. E-valuator: we built a statistical method (using sequential hypothesis testing) to stop agent trajectories early when the agent is incorrect.
  2. SSME: we built a method to include both labeled and unlabeled samples during (simultaneous) evaluation of multiple ML models.

I am also especially interested in applications to societally-impactful problems such as healthcare. During my PhD, I have interned at Abridge on clinical AI evals research with Alex Chouldechova and Michael Oberst and at Genentech on early stopping of agents with Hanchen Wang.

Prior to my PhD, I studied CS and Statistics at Harvard. I enjoy writing, and have written/blogged about fellowship applications, human preference datasets, AI evals and old papers, among other things.

Quick links:

latest posts

Sep 13, 2026 Who are we aligning to?
Mar 30, 2026 Revenge of the Worst Case
Jan 22, 2026 Reading List