“Everything should be made as simple as possible, but not any simpler.”
Zhi-ning Liu
Applied Scientist @ Amazon | PhD @ University of Illinois Urbana-Champaign (UIUC)
Hi! I’m building agentic AI systems for industrial R&D at scale, with a focus on:
- Super-Long-Horizon Agents: developing harness to orchestrate hundreds of agents for weeks-long R&D tasks involving thousands of compute jobs;
- Self-Improving Agents: enabling agents to learn from self/human feedback and reuse accumulated knowledge on new tasks; and
- Agent Observability: designing infrastructure for traceable, auditable, and resumable long-horizon agent task execution.
More broadly, my research spans reliable LLM and VLM reasoning, data-centric AI, and automated data mining.
- Agent Harness (, )
- Agentic Reasoning & Planning (, )
- Self-Improving Agents (, )
- Long-Context Reasoning (, )
- Multimodal Grounding (, )
- Safety & Alignment (, , )
- Data Optimization (, , )
- Data Interpretability (, , )
- Long-Tail Learning (, , )
- Model Fusion (, , )
- Time Series (, , )
- Graph Learning (, )
News
Aug 2026
🎓 Ph.D. unlocked! Successfully defended on August 21st : )
Jul 2026
🎉 COLM'26: 2 papers accepted to COLM 2026.
May 2026
🏆 Award: 2026 C.W. Gear Outstanding Graduate Student (1 in UIUC CS Dept.).
May 2026
🎉 ICML'26: One paper accepted to ICML 2026.
Apr 2026
🎉 ACL'26: 3 papers (2 Main, 1 Findings) accepted to ACL 2026.
Jan 2026
🇧🇷 ICLR'26: 4 papers accepted to ICLR 2026. See you in Brazil!
Oct 2025
👀 VLM Perception: VLMs can see the image, but still may not use it. [PDF]
Sep 2025
May 2025
🏆 Award: Honored to receive the 2025 C.L. and Jane Liu Award! (2 in UIUC CS Dept.) (school website)
May 2025
💼 Intern@Amazon: Back to the Bay Area again.
Jan 2025
🎉 ICLR'25: One paper on test-time adaptation for graph structural shift accepted. [PDF]
May 2024
💼 Intern@Amazon: Starting my Applied Scientist Internship in the Bay Area.
Mar 2024
⚖️ FAccT'24: Group Fairness via Group Consensus, with Eunice Chan. [PDF]
May 2023
🏥 KDD'23: Web-based Long-term Spine Treatment Outcome Forecasting, with Hangting Ye. [PDF]
May 2023
💼 Intern@Amazon: Starting my Applied Scientist Internship in Seattle.
Mar 2022
🎓 Starting Ph.D.@UIUC: I will join Prof. Hanghang Tong's group at UIUC in Fall 2022.
Jan 2022
Apr 2020
📦 Open-source: Awesome-Imbalanced-Learning, a curated list of imbalanced learning resources.
Oct 2019
Jul 2019
🎓 Graduation@JilinU: Received my B.Sc. from Tang Aoqing Honors Program in Science, Jilin University.
Sep 2018
💼 Intern@Microsoft: Starting my internship at Microsoft Research Asia. Supervisors: Dr. Jiang Bian and Dr. Wei Cao.
Education

University of Illinois Urbana-Champaign
Ph.D. in Computer Science · 2022 - 2026

Jilin University
M.Eng. in Computer Science · 2019 - 2022
School of Artificial Intelligence · Advisor: Prof. Yi Chang
B.Sc. in Computer Science · 2015 - 2019
Tang-Aoqing Honors Program (Top ~0.5% at matriculation)
Experience

Amazon
Applied Scientist, Palo Alto, CA, Jul 2026 - Present
Self-improving agent harness for industrial R&D at scale.
Applied Scientist Internships
2025, Palo Alto, CA (8 months): Reliable VLM Reasoning -> , ,
2024, Palo Alto, CA (4 months): RAG-Enhanced LLM Reasoning ->
2023, Seattle, WA (4 months): User query understanding
Microsoft Research
Research Intern, Beijing, Aug 2018 - June 2019 (9 months)
Extreme class-imbalanced learning -> ,
Publications
Publications, preprints, and submissions, sorted by year. See the full list on Google Scholar.
-
ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning
Conference on Empirical Methods in Natural Language Processing (EMNLP 2026 Findings)
-
AdaFuse: Adaptive Ensemble Decoding for Large Language Models
64th Annual Meeting of the Association for Computational Linguistics (ACL 2026 Main)
-
Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents
64th Annual Meeting of the Association for Computational Linguistics (ACL 2026 Main)
-
Do VLMs Have a Moral Backbone? A Study on the Fragile Morality of Vision-Language Models
64th Annual Meeting of the Association for Computational Linguistics (ACL 2026 Findings)
-
Code as Agent Harness
arXiv Preprint (2026)
-
Continual Low-Rank Adapters for LLM-Based Generative Recommender Systems
International Conference on Learning Representations (ICLR 2026)
-
Seeing but Not Believing: Probing the Disconnect between Visual Attention and Answer Correctness in VLMs
International Conference on Learning Representations (ICLR 2026)
-
PlanetAlign: A Comprehensive Python Library for Benchmarking Network Alignment
International Conference on Learning Representations (ICLR 2026)
-
Language in the Flow of Time: Time-Series-Paired Texts Weaved into a Unified Temporal Narrative
International Conference on Learning Representations (ICLR 2026)
-
Mixture of Sequence: Theme-Aware Mixture-of-Experts for Long-Sequence Recommendation
The ACM Web Conference (WWW 2026 Oral)
-
ReMix: Reinforcement Routing for Mixtures of LoRAs in LLM Finetuning
LLA Workshop at ICLR (LLA 2026)
-
Flow Matching Meets Biology and Life Science: A Survey
npj Artificial Intelligence (2026)
-
TSAQA: Time Series Analysis Question and Answering Benchmark
Fifth Workshop on Generation, Evaluation and Metrics (GEM 2026)
-
Agentic Reasoning for Large Language Models
arXiv Preprint (2026)
-
ClimateBench-M: A Multi-Modal Climate Data Benchmark with a Simple Generative Method
34th ACM International Conference on Information and Knowledge Management (CIKM 2025)
-
Not All Voices Are Rewarded Equally: Probing and Repairing Reward Models across Human Diversity
Conference on Empirical Methods in Natural Language Processing (EMNLP 2025 Findings)
-
Hierarchical LoRA MoE for Efficient CTR Model Scaling
arXiv Preprint (2025)
-
LLM-RecG: A Semantic Bias-Aware Framework for Zero-Shot Sequential Recommendation
19th ACM Conference on Recommender Systems (RecSys 2025)
-
Breaking Silos: Adaptive Model Fusion Unlocks Better Time Series Forecasting
42nd International Conference on Machine Learning (ICML 2025)
-
MORALISE: A Structured Benchmark for Moral Alignment in Visual Language Models
arXiv Preprint (2025)
-
CATS: Mitigating Correlation Shift for Multivariate Time Series Classification
arXiv Preprint (2025)
-
SelfElicit: Your Language Model Secretly Knows Where Is the Relevant Evidence
63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025 Main)
-
CLIMB: Class-Imbalanced Learning Benchmark on Tabular Data
Conference on Neural Information Processing Systems (NeurIPS 2025)
-
Matcha: Mitigating Graph Structure Shifts with Test-Time Adaptation
International Conference on Learning Representations (ICLR 2025)
-
THeGCN: Temporal Heterophilic Graph Convolutional Network
arXiv Preprint (2024)
-
BACKTIME: Backdoor Attacks on Multivariate Time Series Forecasting
Conference on Neural Information Processing Systems (NeurIPS 2024 Spotlight)
-
WAPITI: A Watermark for Finetuned Open-Source LLMs
arXiv Preprint (2024)
-
AIM: Attributing, Interpreting, Mitigating Data Unfairness
30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2024)
-
Graph Mixup on Approximate Gromov–Wasserstein Geodesics
41st International Conference on Machine Learning (ICML 2024)
-
Class-Imbalanced Graph Learning without Class Rebalancing
41st International Conference on Machine Learning (ICML 2024)
-
Group Fairness via Group Consensus
ACM Conference on Fairness, Accountability, and Transparency (FAccT 2024)
-
Ensuring User-Side Fairness in Dynamic Recommender Systems
The ACM Web Conference (WWW 2024)
-
Hierarchical Multi-Marginal Optimal Transport for Network Alignment
38th AAAI Conference on Artificial Intelligence (AAAI 2024)
-
Taming Over-Smoothing Representation on Heterophilic Graphs
Information Sciences (2023)
-
Web-Based Long-Term Spine Treatment Outcome Forecasting
29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2023)
-
UADB: Unsupervised Anomaly Detection Booster
39th IEEE International Conference on Data Engineering (ICDE 2023)
-
A Survey of Explainable Graph Neural Networks for Cyber Malware Analysis
IEEE International Conference on Big Data (BigData 2022)
-
Towards Unified Representations of Knowledge Graph and Expert Rules for Machine Learning and Reasoning
2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (AACL 2022 Main)
-
Handling Inter-Class and Intra-Class Imbalance in Class-Imbalanced Learning
arXiv Preprint (2021)
-
IMBENS: Ensemble Class-Imbalanced Learning in Python
arXiv Preprint (2021)
-
Self-Paced Ensemble for Highly Imbalanced Massive Data Classification
36th IEEE International Conference on Data Engineering (ICDE 2020)
-
MESA: Boost Ensemble Imbalanced Learning with Meta-Sampler
Conference on Neural Information Processing Systems (NeurIPS 2020)
-
Inference Scaling of LLM Ensembling: Bridging Token Spaces with Token Translation
Google Scholar Entry
Awards & Honors
-
C.W. Gear Outstanding Graduate Student
University of Illinois Urbana-Champaign
Top 0.2%*1 recipient in UIUC CS department -
C.L. and Jane Liu Award
University of Illinois Urbana-Champaign
Top 0.3%*2 recipients in UIUC CS department -
Top 10 Graduate Student (University Highest Honor)
Jilin University
Top 0.04%*10 of 28,373 graduate students -
National Scholarship
Ministry of Education of China
Top 0.2%Nationwide -
National Scholarship
Ministry of Education of China
Top 0.2%Nationwide
* Estimated using official enrollment data at the time: UIUC CS PhD students (644); Jilin University graduate students (28,373).
Gallery
Fun Facts
🌿 My Name
In Chinese, "Zhi Ning" has a gently feminine feeling: Zhi (芷) means fragrant herb, and Ning (宁) means peace and tranquility. Because of this name, some friends once expected to meet a cute girl before seeing me in person, and were mildly disappointed when reality arrived.
In Chinese, "Zhi Ning" has a gently feminine feeling: Zhi (芷) means fragrant herb, and Ning (宁) means peace and tranquility. Because of this name, some friends once expected to meet a cute girl before seeing me in person, and were mildly disappointed when reality arrived.
🎮 Games
I enjoy nearly every kind of video game, although I am not necessarily good at them: shooters, strategy, 4X, role playing games, roguelikes, racing games, and more. Some of my favorites include Battlefield, Civilization, Stellaris, GTA, The Witcher, DiRT, Homeworld, Metro, BioShock, and Borderlands.
I enjoy nearly every kind of video game, although I am not necessarily good at them: shooters, strategy, 4X, role playing games, roguelikes, racing games, and more. Some of my favorites include Battlefield, Civilization, Stellaris, GTA, The Witcher, DiRT, Homeworld, Metro, BioShock, and Borderlands.
🎨 Making Things
Making things look nice and organized makes me happy. That's why I make good paper figures, tables, and website. I also have a sleek desktop setup. In another life, I might be a designer or a professional organizer. Unfortunately, the former may soon be replaced by AI, while the latter probably still has some time.
Making things look nice and organized makes me happy. That's why I make good paper figures, tables, and website. I also have a sleek desktop setup. In another life, I might be a designer or a professional organizer. Unfortunately, the former may soon be replaced by AI, while the latter probably still has some time.