As Head of Agentic QC at Turing, I lead quality control for LLM-based agents: defining what “correct” behaviour means, identifying ambiguity in task instructions, detecting misalignment between instructions and verifiers, assessing test coverage, and designing evaluations that reveal meaningful failure modes rather than merely confirm success on easy cases. I also work on benchmarking and evaluating models and agents, most recently ML4Science-Bench, a benchmark of end-to-end machine-learning tasks for science (NeurIPS 2026 AI for Science Workshop, Oral).
My research background is central to that work — a decade on information theory, generalization and alignment, including KL-regularized RLHF and best-of-n methods (NeurIPS 2025; ICLR 2026) and off-policy evaluation (ICML 2025 Spotlight). I have applied the same discipline in practice as an AI Data Scientist at HSBC, where I developed FraudTransformer (CAI 2026), and as a Research Associate at The Alan Turing Institute.
I am interested in roles in agentic AI where rigorous evaluation, benchmarking, alignment and inference-time control are essential to building trustworthy agentic systems.