I am a PhD candidate at the National University of Singapore, where I started in August 2024, and an Associate in Research at the Yong Loo Lin School of Medicine. I work on generative AI — on systems that can discover things, and on the conditions under which they quietly stop being able to.
Two questions take up most of my time. The first is the AI scientist: agents that propose and test hypotheses themselves, and whether what comes out is a real finding or a fluent restatement of the training data. The second is model collapse: what happens to a model's diversity once it begins training on machine-generated data. They turn out to be the same question approached from opposite ends.
Medicine is where I test this, because it is where a lost rare case is not an abstract metric. My recent work finds that as AI-generated content accumulates in medical records, rare but critical findings vanish from what models produce and phenotypes converge on the generic.
Before Singapore I was at Duke University, first for a master's in Electrical & Computer Engineering and then as an Associate in Research in the Department of Radiology, working on bone MRI segmentation and CT-based body composition. Earlier projects covered privacy-preserving federated learning and proactive deepfake detection.
Current focus
-
AI scientists
Systems that generate and test hypotheses on their own, rather than executing a plan a human wrote out in advance. The part I care about is whether the output is a genuine finding or a restatement of something already in the training data.
-
Self-evolving AI systems
Agents that rewrite their own scaffolding — prompts, tools, the criteria they are judged by — and improve across iterations. The open question is what keeps such a loop improving instead of drifting.
-
Model collapse and synthetic-data contamination
What happens to a generative model's diversity once it trains on machine-generated output. This is the failure mode underneath the creativity question: as models recycle their own output, the rare cases go first and everything converges on the generic.
Selected publications
-
AI-generated Data Contamination Erodes Pathological Variability and Diagnostic Reliability
Preprint, medRxiv / arXiv:2601.12946, 2026
-
Safety Challenges of AI in Medicine in the Era of Large Language Models
Preprint, arXiv:2409.18968, 2024
-
SuperMask: Generating High-resolution Object Masks from Multi-view, Unaligned Low-resolution MRIs
MIDL 2023
Contact
- Email hongyu.h@nus.edu.sg
- Scholar Google Scholar profile
- Where Yong Loo Lin School of Medicine, NUS, Singapore