Zhi-ning Liu
你好 / Hello / 안녕하세요 / こんにちは / Bonjour / Hola (and more)!
Building agentic, self-improving AI systems to automate industrial R&D at scale. More broadly, my research spans reliable LLM and VLM reasoning, automated data engineering and data mining. Published 30+ papers at leading venues with 1,000+ citations. Alongside these efforts, I build and maintain research-driven open-source software and projects with 4K+ GitHub stars and 200K+ downloads.
- Agent Harness (, )
- Agentic Reasoning & Planning (, )
- Self-Improving Agents (, )
- Long-Context Reasoning (, )
- Multimodal Grounding (, )
- Safety & Alignment (, , )
- Data Optimization (, , )
- Data Interpretability (, , )
- Long-Tail Learning (, , )
- Model Fusion (, , )
- Time Series (, , )
- Graph Learning (, )
News
Education

University of Illinois Urbana-Champaign

Jilin University
Experience

Amazon
Microsoft Research
Publications
Publications, preprints, and submissions, sorted by year. See the full list on Google Scholar.
-
ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning
Conference on Empirical Methods in Natural Language Processing (EMNLP 2026 Findings)
-
AdaFuse: Adaptive Ensemble Decoding for Large Language Models
64th Annual Meeting of the Association for Computational Linguistics (ACL 2026 Main)
-
Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents
64th Annual Meeting of the Association for Computational Linguistics (ACL 2026 Main)
-
Do VLMs Have a Moral Backbone? A Study on the Fragile Morality of Vision-Language Models
64th Annual Meeting of the Association for Computational Linguistics (ACL 2026 Findings)
-
Code as Agent Harness
arXiv Preprint (2026)
-
Continual Low-Rank Adapters for LLM-Based Generative Recommender Systems
International Conference on Learning Representations (ICLR 2026)
-
Seeing but Not Believing: Probing the Disconnect between Visual Attention and Answer Correctness in VLMs
International Conference on Learning Representations (ICLR 2026)
-
PlanetAlign: A Comprehensive Python Library for Benchmarking Network Alignment
International Conference on Learning Representations (ICLR 2026)
-
Language in the Flow of Time: Time-Series-Paired Texts Weaved into a Unified Temporal Narrative
International Conference on Learning Representations (ICLR 2026)
-
Mixture of Sequence: Theme-Aware Mixture-of-Experts for Long-Sequence Recommendation
The ACM Web Conference (WWW 2026 Oral)
-
ReMix: Reinforcement Routing for Mixtures of LoRAs in LLM Finetuning
LLA Workshop at ICLR (LLA 2026)
-
Flow Matching Meets Biology and Life Science: A Survey
npj Artificial Intelligence (2026)
-
TSAQA: Time Series Analysis Question and Answering Benchmark
Fifth Workshop on Generation, Evaluation and Metrics (GEM 2026)
-
Agentic Reasoning for Large Language Models
arXiv Preprint (2026)
-
ClimateBench-M: A Multi-Modal Climate Data Benchmark with a Simple Generative Method
34th ACM International Conference on Information and Knowledge Management (CIKM 2025)
-
Not All Voices Are Rewarded Equally: Probing and Repairing Reward Models across Human Diversity
Conference on Empirical Methods in Natural Language Processing (EMNLP 2025 Findings)
-
Hierarchical LoRA MoE for Efficient CTR Model Scaling
arXiv Preprint (2025)
-
LLM-RecG: A Semantic Bias-Aware Framework for Zero-Shot Sequential Recommendation
19th ACM Conference on Recommender Systems (RecSys 2025)
-
Breaking Silos: Adaptive Model Fusion Unlocks Better Time Series Forecasting
42nd International Conference on Machine Learning (ICML 2025)
-
MORALISE: A Structured Benchmark for Moral Alignment in Visual Language Models
arXiv Preprint (2025)
-
CATS: Mitigating Correlation Shift for Multivariate Time Series Classification
arXiv Preprint (2025)
-
SelfElicit: Your Language Model Secretly Knows Where Is the Relevant Evidence
63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025 Main)
-
CLIMB: Class-Imbalanced Learning Benchmark on Tabular Data
Conference on Neural Information Processing Systems (NeurIPS 2025)
-
Matcha: Mitigating Graph Structure Shifts with Test-Time Adaptation
International Conference on Learning Representations (ICLR 2025)
-
THeGCN: Temporal Heterophilic Graph Convolutional Network
arXiv Preprint (2024)
-
BACKTIME: Backdoor Attacks on Multivariate Time Series Forecasting
Conference on Neural Information Processing Systems (NeurIPS 2024 Spotlight)
-
WAPITI: A Watermark for Finetuned Open-Source LLMs
arXiv Preprint (2024)
-
AIM: Attributing, Interpreting, Mitigating Data Unfairness
30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2024)
-
Graph Mixup on Approximate Gromov–Wasserstein Geodesics
41st International Conference on Machine Learning (ICML 2024)
-
Class-Imbalanced Graph Learning without Class Rebalancing
41st International Conference on Machine Learning (ICML 2024)
-
Group Fairness via Group Consensus
ACM Conference on Fairness, Accountability, and Transparency (FAccT 2024)
-
Ensuring User-Side Fairness in Dynamic Recommender Systems
The ACM Web Conference (WWW 2024)
-
Hierarchical Multi-Marginal Optimal Transport for Network Alignment
38th AAAI Conference on Artificial Intelligence (AAAI 2024)
-
Taming Over-Smoothing Representation on Heterophilic Graphs
Information Sciences (2023)
-
Web-Based Long-Term Spine Treatment Outcome Forecasting
29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2023)
-
UADB: Unsupervised Anomaly Detection Booster
39th IEEE International Conference on Data Engineering (ICDE 2023)
-
A Survey of Explainable Graph Neural Networks for Cyber Malware Analysis
IEEE International Conference on Big Data (BigData 2022)
-
Towards Unified Representations of Knowledge Graph and Expert Rules for Machine Learning and Reasoning
2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (AACL 2022 Main)
-
Handling Inter-Class and Intra-Class Imbalance in Class-Imbalanced Learning
arXiv Preprint (2021)
-
IMBENS: Ensemble Class-Imbalanced Learning in Python
arXiv Preprint (2021)
-
Self-Paced Ensemble for Highly Imbalanced Massive Data Classification
36th IEEE International Conference on Data Engineering (ICDE 2020)
-
MESA: Boost Ensemble Imbalanced Learning with Meta-Sampler
Conference on Neural Information Processing Systems (NeurIPS 2020)
-
Inference Scaling of LLM Ensembling: Bridging Token Spaces with Token Translation
Google Scholar Entry
Awards & Honors
-
C.W. Gear Outstanding Graduate Student
University of Illinois Urbana-Champaign
Top 0.2%*1 recipient in UIUC CS department -
C.L. and Jane Liu Award
University of Illinois Urbana-Champaign
Top 0.3%*2 recipients in UIUC CS department -
Top 10 Graduate Student (University Highest Honor)
Jilin University
Top 0.04%*10 of 28,373 graduate students -
National Scholarship
Ministry of Education of China
Top 0.2%Nationwide -
National Scholarship
Ministry of Education of China
Top 0.2%Nationwide
* Estimated using official enrollment data at the time: UIUC CS PhD students (644); Jilin University graduate students (28,373).
Gallery
Fun Facts
In Chinese, "Zhi Ning" has a gently feminine feeling: Zhi (芷) means fragrant herb, and Ning (宁) means peace and tranquility. Because of this name, some friends once expected to meet a cute girl before seeing me in person, and were mildly disappointed when reality arrived.
I enjoy nearly every kind of video game, although I am not necessarily good at them: shooters, strategy, 4X, role playing games, roguelikes, racing games, and more. Some of my favorites include Battlefield, Civilization, Stellaris, GTA, The Witcher, DiRT, Homeworld, Metro, BioShock, and Borderlands.
Making things look nice and organized makes me happy. That's why I make good paper figures, tables, and website. I also have a sleek desktop setup. In another life, I might be a designer or a professional organizer. Unfortunately, the former may soon be replaced by AI, while the latter probably still has some time.