About Me

I am an assistant professor at Department of Computer Science, George Mason University since Fall 2021. Before that I was a postdoc at Rafik B. Hariri Institute at Boston University from 2020-2021, hosted by Francesco Orabona. I received my Ph.D. at Department of Computer Science, The University of Iowa in August 2020, under the advise of Tianbao Yang. Before that I studied at Institute of Natural Sciences and School of Mathematical Sciences at Shanghai Jiao Tong University. I have also spent time working at industrial research labs, such as IBM research AI and Alibaba DAMO Academy. Here is my Google Scholar Citations.

I am looking for self-motivated PhD students (fully-funded) with strong mathematical ablities and (or) programming skills to solve challenging machine learning problems elegantly with mathematical analysis and empirical studies. The main technical skills we need include mathematical optimization, statistical learning theory, algorithms, and deep learning. If you are interested, please drop me an email with your CV and transcript, and apply our PhD program here. Undergrad and graduate student visitors are also welcome. This link provides an overview of our fast-growing GMU CS department.

Research

Our research addresses the fundamental problem of efficiency in machine learning, with a focus on the interplay of optimization and learning. We approach this problem at three connected levels: algorithms, data, and systems. Across these levels, we connect mathematical principles and provable guarantees with practical algorithms and implementations for modern machine learning problems, such as efficient foundation model training and inference.

  • Algorithm Level: Computation and Communication Efficiency. We study how optimization and learning dynamics determine the computational and communication costs of learning. This includes understanding large-step gradient descent [COLT'26] and the complexity limits of adaptive methods [ICLR'25], as well as when local updates can accelerate distributed learning [ICML'25]. We also develop efficient algorithms for hierarchical optimization, including bilevel problems [ICLR'24 Spotlight] and minimax problems [JMLR'21]. For foundation models, we develop non-Euclidean optimization methods that improve robustness to learning-rate tuning [ICML'26].

  • Data Level: Statistical Efficiency. We study how optimization shapes generalization and feature learning [NeurIPS'21, ICML'24], and how to use training data more effectively. Our work includes bilevel optimization for coreset selection in continual learning [NeurIPS'23], and data selection for foundation model pretraining that accounts for the long-term influence of training examples [ICML'26].

  • Systems Level: Hardware Efficiency. We study how algorithm design and hardware characteristics jointly determine the efficiency of machine learning systems. Our goal is to translate mathematical insights into practical gains in runtime, memory efficiency, and scalability. We connect I/O complexity analysis with efficient algorithms and GPU kernel implementations for foundation model workloads, including exact higher-order differentiation through attention [NeurIPS'26].

  • Recent News

    • (Sep 2026) One paper about I/O-efficient exact backward-over-backward for softmax attention (FlashBoB) was accepted by NeurIPS 2026. This is our group's first work on GPU kernel optimization. Congratulations to my students Anthony and Michael! This is Anthony's first paper, completed during his first year as a PhD student in our group.
    • (May 2026) One paper was accepted by COLT 2026! We show that, for low-dimensional separable data, gradient descent with arbitrarily large step sizes on logistic regression converges to arbitrarily accurate solutions in only a constant number of iterations. This is a surprising and exciting result for us. Congratulations to my student Michael!
    • (April 2026) Two papers were accepted by ICML 2026. Topics include data selection in large language model pretraining and non-Euclidean gradient descent optimization methods. Congratulations to my students Michael, Jie and Rui!
    • (April 2026) My very first PhD student, Michael Crawshaw, successfully defended his PhD dissertation on the mathematical foundations of distributed optimization for machine learning. His disseration won the outstanding dissertation award in the computer science department. He will join Flatiron Institute as a postdoc after graduation. Congratulations to Dr. Crawshaw!
    • (Jan 2026) One paper about bilevel optimization under uniform convexity was accepted by ICLR 2026. Congratulations to my students Yuman, Xiaochuan and Jie!
    • More News

    Recent Selected Publications [Full List]

    Last update: 09-25-2026