PhD in Applied Mathematics
The Hong Kong Polytechnic University
I study representation learning through distributional and generative perspectives. My focus is self-supervised learning and its connection to generative learning; I also study the theory of language models.

I am pursuing my PhD in the Department of Applied Mathematics at The Hong Kong Polytechnic University, advised by Prof. Defeng Sun and Prof. Houduo Qi. I received my bachelor’s degree from Northwest University and my master’s degree from Wuhan University, where I was advised by Prof. Yuling Jiao. Prof. Jiao continues to advise me during my PhD, and we work closely together.
Bringing Generative Learning to Representation Learning
The starting point is the distribution of useful representations: can we learn an encoder by matching its outputs to a geometric reference? Distribution Matching (DM) develops this perspective; Flow-Based Distribution Matching (FBDM) extends it through flow matching.
01 · The perspective
Distribution Matching
Recasting self-supervised learning as distribution matching opens the door to generative learning tools for representation learning.
02 · The flow-based extension
Flow-Based Distribution Matching
What if a generative flow could learn representations rather than generate samples?
News
- Sep 2026 · FBDM, a flow-based extension of our distribution-matching approach to self-supervised representations, is now available as a preprint.
- Aug 2026 · We substantially revised Distribution Matching, including its title and the connection between generative and representation learning.
Publications & Preprints
Authors are listed alphabetically by surname in all publications.
Learning a Flow to Self-Supervised Representations
What if a generative flow could learn representations rather than generate samples?
Bringing Generative Learning to Representation Learning: Self-Supervised Transfer Learning as Distribution Matching
Recasting self-supervised learning as distribution matching opens the door to generative learning tools for representation learning.
Beyond the Prompt in Large Language Models: Comprehension, In-Context Learning, and Chain-of-Thought
We develop a theoretical framework to model and understand zero-shot prediction, in-context learning, and chain-of-thought reasoning. We seek to explain how demonstrations and intermediate reasoning steps can improve performance along the progression from zero-shot prediction to in-context learning and chain-of-thought.
Adv-SSL: Adversarial Self-Supervised Representation Learning with Theoretical Guarantees
We introduce a minimax approach to debias existing self-supervised learning methods. This adversarial formulation improves downstream performance while helping establish theoretical guarantees for the learned representations.
Talks & Presentations
Recent presentations on distribution matching and self-supervised representation learning include JCSDS 2026 in Guiyang, EAC-ISBA 2026 in Kunming, and a poster at NeurIPS 2025 in San Diego.
Awards
NeurIPS 2025 Scholar Award
Travel Award, Young Statisticians Association Annual Meeting, 2025
Third International Symposium on Statistical Theory and Applications, Doctoral Forum. Selected for an oral presentation and awarded travel support.
Open-source Projects
mlimpl
Readable implementations of machine learning algorithms for learning, prototyping, and reproducible research.