Hongyu He
Portrait of Hongyu He

Hongyu He

PhD candidate, National University of Singapore

AI scientists · self-evolving systems · model collapse

I am a PhD candidate at the National University of Singapore, where I started in August 2024, and an Associate in Research at the Yong Loo Lin School of Medicine. I work on generative AI — on systems that can discover things, and on the conditions under which they quietly stop being able to.

Two questions take up most of my time. The first is the AI scientist: agents that propose and test hypotheses themselves, and whether what comes out is a real finding or a fluent restatement of the training data. The second is model collapse: what happens to a model's diversity once it begins training on machine-generated data. They turn out to be the same question approached from opposite ends.

Medicine is where I test this, because it is where a lost rare case is not an abstract metric. My recent work finds that as AI-generated content accumulates in medical records, rare but critical findings vanish from what models produce and phenotypes converge on the generic.

Before Singapore I was at Duke University, first for a master's in Electrical & Computer Engineering and then as an Associate in Research in the Department of Radiology, working on bone MRI segmentation and CT-based body composition. Earlier projects covered privacy-preserving federated learning and proactive deepfake detection.

Current focus

  1. AI scientists

    Systems that generate and test hypotheses on their own, rather than executing a plan a human wrote out in advance. The part I care about is whether the output is a genuine finding or a restatement of something already in the training data.

  2. Self-evolving AI systems

    Agents that rewrite their own scaffolding — prompts, tools, the criteria they are judged by — and improve across iterations. The open question is what keeps such a loop improving instead of drifting.

  3. Model collapse and synthetic-data contamination

    What happens to a generative model's diversity once it trains on machine-generated output. This is the failure mode underneath the creativity question: as models recycle their own output, the rare cases go first and everything converges on the generic.

More on the research →

Selected publications

  • AI-generated Data Contamination Erodes Pathological Variability and Diagnostic Reliability

    Hongyu He, Shaowen Xiang, Ye Zhang, Yingtao Zhu, Jin Zhang, Yunyi Lu, Hao Deng, Emily Alsentzer, Yun Liu, Qingyu Chen, Kun-Hsing Yu, Andrew Marshall, Tingting Chen, Srinivas Anumasa, Daniel Ebner, Dean Ho, Kee Yuan Ngiam, Ching-Yu Cheng, Dianbo Liu

    Preprint, medRxiv / arXiv:2601.12946, 2026

  • Safety Challenges of AI in Medicine in the Era of Large Language Models

    Xiaoye Wang, Nicole Xi Zhang, Hongyu He, Trang Nguyen, Kun-Hsing Yu, Hao Deng, Cynthia Brandt, Danielle S. Bitterman, Ling Pan, Ching-Yu Cheng, James Zou, Dianbo Liu

    Preprint, arXiv:2409.18968, 2024

  • SuperMask: Generating High-resolution Object Masks from Multi-view, Unaligned Low-resolution MRIs

    Hanxue Gu, Hongyu He, Roy Colglazier, Jordan Axelrod, Robert French, Maciej A. Mazurowski

    MIDL 2023

All publications →

Contact