About Me
I am an assistant professor at Department of Computer Science, George Mason University since Fall 2021. Before that I was a postdoc at Rafik B. Hariri Institute at Boston University from 2020-2021, hosted by Francesco Orabona. I received my Ph.D. at Department of Computer Science, The University of Iowa in August 2020, under the advise of Tianbao Yang. Before that I studied at Institute of Natural Sciences and School of Mathematical Sciences at Shanghai Jiao Tong University. I have also spent time working at industrial research labs, such as IBM research AI and Alibaba DAMO Academy. Here is my Google Scholar Citations.
I am looking for self-motivated PhD students (fully-funded) with strong mathematical ablities and (or) programming skills to solve challenging machine learning problems elegantly with mathematical analysis and empirical studies. The main technical skills we need include mathematical optimization, statistical learning theory, algorithms, and deep learning. If you are interested, please drop me an email with your CV and transcript, and apply our PhD program here. Undergrad and graduate student visitors are also welcome. This link provides an overview of our fast-growing GMU CS department.
I am looking for self-motivated PhD students (fully-funded) with strong mathematical ablities and (or) programming skills to solve challenging machine learning problems elegantly with mathematical analysis and empirical studies. The main technical skills we need include mathematical optimization, statistical learning theory, algorithms, and deep learning. If you are interested, please drop me an email with your CV and transcript, and apply our PhD program here. Undergrad and graduate student visitors are also welcome. This link provides an overview of our fast-growing GMU CS department.
Research
Our research addresses the fundamental problem of efficiency in machine learning, with a focus on the interplay of optimization and learning. We approach this problem at three connected levels: algorithms, data, and systems. Across these levels, we connect mathematical principles and provable guarantees with practical algorithms and implementations for modern machine learning problems, such as efficient foundation model training and inference.
Algorithm Level: Computation and Communication Efficiency. We study how optimization and learning dynamics determine the computational and communication costs of learning. This includes understanding large-step gradient descent [COLT'26] and the complexity limits of adaptive methods [ICLR'25], as well as when local updates can accelerate distributed learning [ICML'25]. We also develop efficient algorithms for hierarchical optimization, including bilevel problems [ICLR'24 Spotlight] and minimax problems [JMLR'21]. For foundation models, we develop non-Euclidean optimization methods that improve robustness to learning-rate tuning [ICML'26].
Data Level: Statistical Efficiency. We study how optimization shapes generalization and feature learning [NeurIPS'21, ICML'24], and how to use training data more effectively. Our work includes bilevel optimization for coreset selection in continual learning [NeurIPS'23], and data selection for foundation model pretraining that accounts for the long-term influence of training examples [ICML'26].
Systems Level: Hardware Efficiency. We study how algorithm design and hardware characteristics jointly determine the efficiency of machine learning systems. Our goal is to translate mathematical insights into practical gains in runtime, memory efficiency, and scalability. We connect I/O complexity analysis with efficient algorithms and GPU kernel implementations for foundation model workloads, including exact higher-order differentiation through attention [NeurIPS'26].
Recent News
- (Sep 2026) One paper about I/O-efficient exact backward-over-backward for softmax attention (FlashBoB) was accepted by NeurIPS 2026. This is our group's first work on GPU kernel optimization. Congratulations to my students Anthony and Michael! This is Anthony's first paper, completed during his first year as a PhD student in our group.
- (May 2026) One paper was accepted by COLT 2026! We show that, for low-dimensional separable data, gradient descent with arbitrarily large step sizes on logistic regression converges to arbitrarily accurate solutions in only a constant number of iterations. This is a surprising and exciting result for us. Congratulations to my student Michael!
- (April 2026) Two papers were accepted by ICML 2026. Topics include data selection in large language model pretraining and non-Euclidean gradient descent optimization methods. Congratulations to my students Michael, Jie and Rui!
- (April 2026) My very first PhD student, Michael Crawshaw, successfully defended his PhD dissertation on the mathematical foundations of distributed optimization for machine learning. His disseration won the outstanding dissertation award in the computer science department. He will join Flatiron Institute as a postdoc after graduation. Congratulations to Dr. Crawshaw!
- (Jan 2026) One paper about bilevel optimization under uniform convexity was accepted by ICLR 2026. Congratulations to my students Yuman, Xiaochuan and Jie!
- More News
Recent Selected Publications [Full List]
-
#: supervised student author, *: equal contribution (alphabetical order)
-
(New! ) FlashBoB: I/O-Efficient Exact Backward-over-Backward for Softmax Attention
Anthony Givans#, Michael Crawshaw#, Mingrui Liu.
40th Conference on Neural Information Processing Systems, 2026. (NeurIPS 2026)
-
(New! ) Tight Bounds for Logistic Regression with Large Stepsize Gradient Descent in Low Dimension
Michael Crawshaw#, Mingrui Liu.
In 39th Annual Conference on Learning Theory, 2026. (COLT 2026)
-
(New! ) BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining
Jie Hao#, Rui Yu#, Wei Zhang, Huixia Wang, Jie Xu, Mingrui Liu.
In 14th International Conference on Machine Learning, 2026. (ICML 2026)
-
(New! ) An Exploration of Non-Euclidean Gradient Descent: Muon and its Many Variants
Michael Crawshaw#, Chirag Modi, Mingrui Liu, Robert M. Gower.
In 14th International Conference on Machine Learning, 2026. (ICML 2026)
-
(New! ) Bilevel Optimization with Lower-Level Uniform Convexity: Theory and Algorithm
Yuman Wu#, Xiaochuan Gong#, Jie Hao#, Mingrui Liu.
In 14th International Conference on Learning Representations, 2026. (ICLR 2026)
-
Adaptive Algorithms with Sharp Convergence Rates for Stochastic Hierarchical Optimization
Xiaochuan Gong#, Jie Hao#, Mingrui Liu.
39th Conference on Neural Information Processing Systems, 2025. (NeurIPS 2025)
-
Constant Stepsize Local GD for Logistic Regression: Acceleration by Instability
Michael Crawshaw#, Blake Woodworth, Mingrui Liu.
Proceedings of 42th International Conference on Machine Learning, 2025. (ICML 2025)
-
Local Steps Speed Up Local GD for Heterogeneous Distributed Logistic Regression
Michael Crawshaw#, Blake Woodworth, Mingrui Liu.
In 13th International Conference on Learning Representations, 2025. (ICLR 2025)
-
Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness
Michael Crawshaw#, Mingrui Liu.
In 13th International Conference on Learning Representations, 2025. (ICLR 2025)
-
Federated Learning under Periodic Client Participation and Heterogeneous Data: A New Communication-Efficient Algorithm and Analysis
Michael Crawshaw#, Mingrui Liu.
In Advances in Neural Information Processing Systems 37, 2024. (NeurIPS 2024)
-
An Accelerated Algorithm for Stochastic Bilevel Optimization under Unbounded Smoothness
Xiaochuan Gong#, Jie Hao#, Mingrui Liu.
In Advances in Neural Information Processing Systems 37, 2024. (NeurIPS 2024)
-
Provable Benefits of Local Steps in Heterogeneous Federated Learning for Neural Networks: A Feature Learning Perspective
Yajie Bao#, Michael Crawshaw#, Mingrui Liu.
Proceedings of 41th International Conference on Machine Learning, 2024. (ICML 2024)
-
A Nearly Optimal Single Loop Algorithm for Stochastic Bilevel Optimization under Unbounded Smoothness
Xiaochuan Gong#, Jie Hao#, Mingrui Liu.
Proceedings of 41th International Conference on Machine Learning, 2024. (ICML 2024)
-
Bilevel Optimization under Unbounded Smoothness: A New Algorithm and Convergence Analysis
Jie Hao#, Xiaochuan Gong#, Mingrui Liu.
In 12th International Conference on Learning Representations, 2024. (ICLR 2024) (Spotlight, 5% acceptance rate)
-
Federated Learning with Client Subsampling, Data Heterogeneity, and Unbounded Smoothness: A New Algorithm and Lower Bounds
Michael Crawshaw#, Yajie Bao#, Mingrui Liu.
In Advances in Neural Information Processing Systems 36, 2023. (NeurIPS 2023)
-
Global Convergence Analysis of Local SGD for Two-layer Neural Network without Overparameterization
Yajie Bao#, Amarda Shehu, Mingrui Liu.
In Advances in Neural Information Processing Systems 36, 2023. (NeurIPS 2023)
-
Bilevel Coreset Selection in Continual Learning: A New Formulation and Algorithm
Jie Hao#, Kaiyi Ji, Mingrui Liu.
In Advances in Neural Information Processing Systems 36, 2023. (NeurIPS 2023)
-
EPISODE: Episodic Gradient Clipping with Periodic Resampled Corrections for Federated Learning with Heterogeneous Data
Michael Crawshaw#, Yajie Bao#, Mingrui Liu.
In 11th International Conference on Learning Representations, 2023. (ICLR 2023)
-
A Communication-Efficient Distributed Gradient Clipping Algorithm for Training Deep Neural Networks
Mingrui Liu, Zhenxun Zhuang, Yunwen Lei, Chunyang Liao.
In Advances in Neural Information Processing Systems 35, 2022. (NeurIPS 2022) (Spotlight, 5% acceptance rate)
-
Robustness to Unbounded Smoothness of Generalized SignSGD
Michael Crawshaw#*, Mingrui Liu*, Francesco Orabona*, Wei Zhang*, Zhenxun Zhuang*.
In Advances in Neural Information Processing Systems 35, 2022. (NeurIPS 2022)
-
Will Bilevel Optimizers Benefit from Loops
Kaiyi Ji, Mingrui Liu, Yingbin Liang, Lei Ying.
In Advances in Neural Information Processing Systems 35, 2022. (NeurIPS 2022) (Spotlight, 5% acceptance rate)
-
Fast Composite Optimization and Statistical Recovery in Federated Learning
Yajie Bao#, Michael Crawshaw#, Shan Luo, Mingrui Liu.
Proceedings of 39th International Conference on Machine Learning, 2022. (ICML 2022)
-
Understanding AdamW through Proximal Methods and Scale-Freeness
Zhenxun Zhuang, Mingrui Liu, Ashok Cutkosky, Francesco Orabona.
Transactions on Machine Learning Research, 2022. (TMLR 2022)
-
On the Initialization for Convex-Concave Min-max Problems
Mingrui Liu, Francesco Orabona.
Algorithmic Learning Theory, 2022. (ALT 2022)
-
On the Last Iterate Convergence of Momentum Methods
Xiaoyu Li, Mingrui Liu, Francesco Orabona.
Algorithmic Learning Theory, 2022. (ALT 2022)
-
Generalization Guarantee of SGD for Pairwise Learning
Yunwen Lei, Mingrui Liu, Yiming Ying.
Advances in Neural Information Processing Systems 34, 2021. (NeurIPS 2021)
Last update: 09-25-2026
