About
I am a Ph.D. student in Electrical and Computer Engineering at the University of California, San Diego, advised by Prof. Pengtao Xie. Previously, I was a Research Scientist Intern at Meta. I received my B.S. in Mathematics and Physics from the Yingcai Honors College at the University of Electronic Science and Technology of China.
My research focuses on AI agents and LLM reasoning.
Research
I build AI agents and develop methods to steer and evaluate LLM reasoning, with a primary focus on improving capabilities at inference time.
AI Agents. Through the AIBuildAI series, I build autonomous agents for machine learning engineering (MLE), which take a task description and data and build the AI model end to end. I develop their harnesses and workflows, memory and knowledge bases, and model routing for efficient execution. My current work extends these systems to post-training, enabling agents to set up infrastructure and run multi-node training, a step toward recursive self-improvement (RSI). Looking ahead, I am interested in large-scale multi-agent collaboration and knowledge bases.
Reasoning Steering & Evaluation. I study how to guide reasoning, evaluate intermediate steps and final responses, and select models under compute constraints. My work spans metacognitive control, process reward modeling, model routing, and LLM-as-a-judge evaluation. I also explore training methods, including agentic reinforcement learning and on-policy distillation, to improve agent and reasoning capabilities.
01
Machine Learning Engineering Agents
AIBuildAI Series
Agent harnesses & workflows
Memory & knowledge bases
Model routing for cost & token efficiency
Post-training Agentin progress
Infrastructure setup & multi-node training, toward recursive self-improvement (RSI)
Next
Large-scale multi-agent collaborationLarge-scale knowledge bases
02
Reasoning Steering & Evaluation
Guiding, selecting, and judging
Reasoning control
Steer reasoning and allocate test-time effort
Process rewards
DreamPRM · DreamPRM-1.5 · DreamPRM-Code
Evaluate intermediate steps and guide solution selection
Model selection & allocation
Balance capability and compute
Evaluation & judgment
LLM-judge subjectivity at scale and how to steer it
Exploring
Agentic RL · On-policy distillation
News
Wrapped up a wonderful research scientist internship at Meta! JudgeProfile is the work from my internship.
Started my research scientist internship at Meta!
AIBuildAI v2 is out!
SCOPE is accepted by ICML 2026!
AIBuildAI ranks No.1 on MLE-bench!
UCSD Today News covers our work DreamPRM!
We built a project page for SCOPE.
I will join Meta as a research scientist intern in Summer 2026!
Starting my PhD at UCSD.
Selected Publications
View All →JudgeProfile: Understanding and Steering Subjectivity in LLM Judges
Qi Cao, Kangning Liu, Xuan Kan, Shunwen Tan, Yang Pei, Dake Chen, Yatai Ji, Zixuan Ye, Yuanpeng Tu, Daniel Li, Junbiao Tang, Pengtao Xie, Zihao He
Arxiv Preprint
Studying LLM-judge subjectivity at scale with SubjectiveSet: 50K response pairs, 21 judges from 0.8B to 2T parameters, 87 attributes, and 1.6M inference calls. Judges share a hidden consensus in perception but differ in how they prioritize attributes, and reweighting attributes steers them to a target standard.
LLMs Know When They Know, but Do Not Act on It: A Metacognitive Harness for Test-time Scaling
Qi Cao, Yufan Wang, Peijia Qin, Shuhao Zhang, Pengtao Xie
Arxiv Preprint
A training-free metacognitive harness that turns LLMs' pre-solve feeling-of-knowing and post-solve judgment-of-learning signals into an explicit test-time control interface, boosting a fixed Claude Sonnet-4.6 across text, code, and multimodal benchmarks.
AIBuildAI: An AI agent that automatically builds AI models
Ruiyi Zhang†, Peijia Qin†, Qi Cao†, Li Zhang†, Pengtao Xie
ICML 2026 Workshop on AI as a Tool for Mathematics, Computer Science, and Machine Learning (ICML Workshop)
We present AIBuildAI, an AI agent that automatically builds AI models, with the goal of solving general AI tasks in an end-to-end manner.
Models Under SCOPE: Scalable and Controllable Routing via Pre-hoc Reasoning
Qi Cao†, Shuhao Zhang†, Ruizhe Zhou, Ruiyi Zhang, Peijia Qin, Pengtao Xie
The Forty-Third International Conference on Machine Learning (ICML)
SCOPE, a model routing framework that predicts how accurate and how expensive each model will be before running it, allowing users to control cost-accuracy trade-offs and naturally handle new models.
DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning
Qi Cao, Ruiyi Wang, Ruiyi Zhang, Sai Ashish Somayajula, Pengtao Xie
The Thirty-Ninth Annual Conference on Neural Information Processing Systems (NeurIPS)
A multimodal Process Reward Model (PRM) trained with domain-reweighting. Top 1 method on MathVista, MMMU & R-Bench-V.
Bidomain Modeling Paradigm for Pansharpening
Junming Hou†, Qi Cao†, Ran Ran, Che Liu, Junling Li, Liang-jian Deng
Proceedings of the 31st ACM international conference on multimedia (ACM MM)
We propose BiPan, a bidomain pansharpening framework that models band-specific local spectral features and global spatial details in the Fourier domain, achieving state-of-the-art performance by better handling spectral diversity and MS image degradation.
Zero-shot Semi-supervised Learning for Pansharpening
Qi Cao, Liang-Jian Deng, Wu Wang, Junming Hou, Gemine Vivone
Information Fusion
Zero-shot pansharpening (ZS-Pan) only requires a single pair of PAN/LRMS images. Any pansharpening network can take the ZS-Pan as a plug-and-play module. A two-phase three-component semi-supervised model is designed for ZS-Pan.
