“Everything should be made as simple as possible, but not any simpler.”
Zhi-ning Liu
Hi! I’m building agentic AI systems for industrial R&D at scale, with a focus on:
- Super-Long-Horizon Agents: developing harness to orchestrate hundreds of agents for weeks-long R&D tasks involving thousands of compute jobs;
- Self-Improving Agents: enabling agents to learn from self/human feedback and reuse accumulated knowledge on new tasks; and
- Agent Observability: designing infrastructure for traceable, auditable, and resumable long-horizon agent task execution.
Beyond these areas, I work on reliable LLM/VLM reasoning, data-centric AI, and automated data mining. I was fortunate to be advised by Prof. Hanghang Tong (Ph.D. at UIUC) and Prof. Yi Chang (M.Eng. at JLU). I also work closely with researchers from industry research labs including Amazon Science, Microsoft Research, Meta, and IBM Research. I’m always happy to chat and discuss potential collaborations. Reach me via: WeChat (ZhiningLiuCS) or zhining.liu AT outlook.com.
- Agent Harness (, )
- Agentic Reasoning & Planning (, )
- Self-Improving Agents (, )
- Long-Context Reasoning (, )
- Multimodal Grounding (, )
- Safety & Alignment (, , )
- Data Optimization (, , )
- Data Interpretability (, , )
- Long-Tail Learning (, , )
- Model Fusion (, , )
- Time Series (, , )
- Graph Learning (, )
News
Education

University of Illinois Urbana-Champaign

Jilin University
Experience

Amazon

Microsoft Research
Publications
Publications, preprints, and submissions, sorted by year. See the full list on Google Scholar.
-
ReContext: Recursive Evidence Replay as LLM Harness for Long-Context ReasoningConference on Empirical Methods in Natural Language Processing (EMNLP 2026 Findings) -
AdaFuse: Adaptive Ensemble Decoding for Large Language Models64th Annual Meeting of the Association for Computational Linguistics (ACL 2026 Main) -
Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents64th Annual Meeting of the Association for Computational Linguistics (ACL 2026 Main) -
Do VLMs Have a Moral Backbone? A Study on the Fragile Morality of Vision-Language Models64th Annual Meeting of the Association for Computational Linguistics (ACL 2026 Findings) -
-
Continual Low-Rank Adapters for LLM-Based Generative Recommender SystemsInternational Conference on Learning Representations (ICLR 2026) -
Seeing but Not Believing: Probing the Disconnect between Visual Attention and Answer Correctness in VLMsInternational Conference on Learning Representations (ICLR 2026) -
PlanetAlign: A Comprehensive Python Library for Benchmarking Network AlignmentInternational Conference on Learning Representations (ICLR 2026) -
Language in the Flow of Time: Time-Series-Paired Texts Weaved into a Unified Temporal NarrativeInternational Conference on Learning Representations (ICLR 2026) -
Mixture of Sequence: Theme-Aware Mixture-of-Experts for Long-Sequence RecommendationThe ACM Web Conference (WWW 2026 Oral) -
ReMix: Reinforcement Routing for Mixtures of LoRAs in LLM FinetuningLLA Workshop at ICLR (LLA 2026) -
Flow Matching Meets Biology and Life Science: A Surveynpj Artificial Intelligence (2026) -
TSAQA: Time Series Analysis Question and Answering BenchmarkFifth Workshop on Generation, Evaluation and Metrics (GEM 2026) -
-
ClimateBench-M: A Multi-Modal Climate Data Benchmark with a Simple Generative Method34th ACM International Conference on Information and Knowledge Management (CIKM 2025) -
Not All Voices Are Rewarded Equally: Probing and Repairing Reward Models across Human DiversityConference on Empirical Methods in Natural Language Processing (EMNLP 2025 Findings) -
-
LLM-RecG: A Semantic Bias-Aware Framework for Zero-Shot Sequential Recommendation19th ACM Conference on Recommender Systems (RecSys 2025) -
Breaking Silos: Adaptive Model Fusion Unlocks Better Time Series Forecasting42nd International Conference on Machine Learning (ICML 2025) -
MORALISE: A Structured Benchmark for Moral Alignment in Visual Language ModelsarXiv Preprint (2025) -
CATS: Mitigating Correlation Shift for Multivariate Time Series ClassificationarXiv Preprint (2025) -
SelfElicit: Your Language Model Secretly Knows Where Is the Relevant Evidence63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025 Main) -
-
Matcha: Mitigating Graph Structure Shifts with Test-Time AdaptationInternational Conference on Learning Representations (ICLR 2025) -
-
BACKTIME: Backdoor Attacks on Multivariate Time Series ForecastingConference on Neural Information Processing Systems (NeurIPS 2024 Spotlight) -
-
AIM: Attributing, Interpreting, Mitigating Data Unfairness30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2024) -
Graph Mixup on Approximate Gromov–Wasserstein Geodesics41st International Conference on Machine Learning (ICML 2024) -
Class-Imbalanced Graph Learning without Class Rebalancing41st International Conference on Machine Learning (ICML 2024) -
Group Fairness via Group ConsensusACM Conference on Fairness, Accountability, and Transparency (FAccT 2024) -
Ensuring User-Side Fairness in Dynamic Recommender SystemsThe ACM Web Conference (WWW 2024) -
Hierarchical Multi-Marginal Optimal Transport for Network Alignment38th AAAI Conference on Artificial Intelligence (AAAI 2024) -
Taming Over-Smoothing Representation on Heterophilic GraphsInformation Sciences (2023) -
Web-Based Long-Term Spine Treatment Outcome Forecasting29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2023) -
UADB: Unsupervised Anomaly Detection Booster39th IEEE International Conference on Data Engineering (ICDE 2023) -
A Survey of Explainable Graph Neural Networks for Cyber Malware AnalysisIEEE International Conference on Big Data (BigData 2022) -
Towards Unified Representations of Knowledge Graph and Expert Rules for Machine Learning and Reasoning2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (AACL 2022 Main) -
Handling Inter-Class and Intra-Class Imbalance in Class-Imbalanced LearningarXiv Preprint (2021) -
-
Self-Paced Ensemble for Highly Imbalanced Massive Data Classification36th IEEE International Conference on Data Engineering (ICDE 2020) -
MESA: Boost Ensemble Imbalanced Learning with Meta-SamplerConference on Neural Information Processing Systems (NeurIPS 2020) -
Inference Scaling of LLM Ensembling: Bridging Token Spaces with Token TranslationGoogle Scholar Entry
Awards & Honors
-
C.W. Gear Outstanding Graduate Student
University of Illinois Urbana-Champaign
Top 0.2%*1 recipient in UIUC CS department -
C.L. and Jane Liu Award
University of Illinois Urbana-Champaign
Top 0.3%*2 recipients in UIUC CS department -
Top 10 Graduate Student (University Highest Honor)
Jilin University
Top 0.04%*10 of 28,373 graduate students -
National Scholarship
Ministry of Education of China
Top 0.2%Nationwide -
National Scholarship
Ministry of Education of China
Top 0.2%Nationwide
* Estimated using official enrollment data at the time: UIUC CS PhD students (644); Jilin University graduate students (28,373).
Gallery
Fun Facts
In Chinese, "Zhi Ning" has a gently feminine feeling: Zhi (芷) means fragrant herb, and Ning (宁) means peace and tranquility. Because of this name, some friends once expected to meet a cute girl before seeing me in person, and were mildly disappointed when reality arrived.
I enjoy nearly every kind of video game, although I am not necessarily good at them: shooters, strategy, 4X, role playing games, roguelikes, racing games, and more. Some of my favorites include Battlefield, Civilization, Stellaris, GTA, The Witcher, DiRT, Homeworld, Metro, BioShock, and Borderlands.
Making things look nice and organized makes me happy. That's why I make good paper figures, tables, and website. I also have a sleek desktop setup. In another life, I might be a designer or a professional organizer. Unfortunately, the former may soon be replaced by AI, while the latter probably still has some time.