D-Scope: Decomposing and Steering Diffusion Transformers with Sparse Autoencoders
Uses sparse autoencoders to decompose and steer diffusion transformers.
I am an incoming postdoctoral fellow at The Chinese University of Hong Kong (CUHK), where I will work with Prof. Weiyang Liu and Prof. Pheng-Ann Heng. I am also a Pivotal AI Safety Research Fellow, mentored by Dr. Peter Hase of Schmidt Sciences and Stanford University.
I completed my PhD in Electronic and Computer Engineering at HKUST in July 2026 as a Hong Kong PhD Fellowship awardee, advised by Prof. Xiaomeng Li. I also spent part of my PhD at the University of Cambridge as a visiting student, supervised by Prof. Adrian Weller and Prof. Weiyang Liu.
Before my PhD, I earned a Bachelor of Advanced Computing (Honours) at The Australian National University (ANU), specializing in Intelligent Systems. My honours research was supervised by Senior Prof. Amanda Barnard and Dr. Amanda Parker. I also held research internships at the National University of Singapore and the MIT CSAIL Computational Connectomics Group, where I worked with Prof. Lu Mi and Prof. Hao Wang.
I welcome research collaborations and would be happy to hear from you if our interests align.
* indicates equal contribution. Underlined names indicate equal advising.
Uses sparse autoencoders to decompose and steer diffusion transformers.
Combines prompt-driven lesion localization with concept-grounded reasoning for interpretable medical report generation.
Improves interpretability and robustness of Concept Bottleneck Models under domain shifts through adversarial training.
A unified framework for concept-based prediction, concept correction, and fine-grained interpretations based on conditional probabilities.
Unifies concept-based generation, conditional interpretation, concept debugging, intervention, and imputation under a joint energy-based formulation.
Recognizes positive SARS cases across numerical and categorical datasets, with explanatory rules generated for clearer interpretability.
Studies faithful-first reasoning, planning, and acting for multimodal large language models.
Combines composition, crystal structure, and radial distribution features for interpretable material property prediction and more stable out-of-distribution generalization.
Integrates attribute tokens with generative pre-trained vision-language models for medical image understanding.
Connects molecular composition and biological reactivity in reservoir data, yielding additional interpretability beyond predictive performance.
Integrates irradiation experiments and provides insights into estuarine dissolved organic matter transformation.
Conference version: The 6th Xiamen Symposium on Marine Environmental Sciences, Best Poster Award.
Uses efficient dynamic augmentation to reduce overfitting in limited medical image segmentation settings.