Qi Cao

Qi Cao

PhD Student

University of California, San Diego

Research Interests

AI Agents
LLM Reasoning
Inference-time Scaling

About

I am a Ph.D. student in Electrical and Computer Engineering at the University of California, San Diego, advised by Prof. Pengtao Xie. Previously, I was a Research Scientist Intern at Meta. I received my B.S. in Mathematics and Physics from the Yingcai Honors College at the University of Electronic Science and Technology of China.

My research focuses on AI agents and LLM reasoning.

Research

I build AI agents and develop methods to steer and evaluate LLM reasoning, with a primary focus on improving capabilities at inference time.

AI Agents. Through the AIBuildAI series, I build autonomous agents for machine learning engineering (MLE), which take a task description and data and build the AI model end to end. I develop their harnesses and workflows, memory and knowledge bases, and model routing for efficient execution. My current work extends these systems to post-training, enabling agents to set up infrastructure and run multi-node training, a step toward recursive self-improvement (RSI). Looking ahead, I am interested in large-scale multi-agent collaboration and knowledge bases.

Reasoning Steering & Evaluation. I study how to guide reasoning, evaluate intermediate steps and final responses, and select models under compute constraints. My work spans metacognitive control, process reward modeling, model routing, and LLM-as-a-judge evaluation. I also explore training methods, including agentic reinforcement learning and on-policy distillation, to improve agent and reasoning capabilities.

01

Machine Learning Engineering Agents

AIBuildAI Series

  1. AIBuildAI

    Agent harnesses & workflows

  2. AIBuildAI-2

    Memory & knowledge bases

  3. AIBuildAI-2.5

    Model routing for cost & token efficiency

  4. Post-training Agentin progress

    Infrastructure setup & multi-node training, toward recursive self-improvement (RSI)

  5. Next

    Large-scale multi-agent collaborationLarge-scale knowledge bases

02

Reasoning Steering & Evaluation

Guiding, selecting, and judging

News

2026-09

Wrapped up a wonderful research scientist internship at Meta! JudgeProfile is the work from my internship.

2026-06

Started my research scientist internship at Meta!

2026-05

AIBuildAI v2 is out!

2026-05

SCOPE is accepted by ICML 2026!

2026-03

DeepTech, ASI, and AIxiv cover our work AIBuildAI!

2026-03

AIBuildAI ranks No.1 on MLE-bench!

2026-02

UCSD Today News covers our work DreamPRM!

2026-02

We built a project page for SCOPE.

2025-12

I will join Meta as a research scientist intern in Summer 2026!

2024-09

Starting my PhD at UCSD.

Selected Publications

View All →

JudgeProfile: Understanding and Steering Subjectivity in LLM Judges

Qi Cao, Kangning Liu, Xuan Kan, Shunwen Tan, Yang Pei, Dake Chen, Yatai Ji, Zixuan Ye, Yuanpeng Tu, Daniel Li, Junbiao Tang, Pengtao Xie, Zihao He

Arxiv Preprint

Studying LLM-judge subjectivity at scale with SubjectiveSet: 50K response pairs, 21 judges from 0.8B to 2T parameters, 87 attributes, and 1.6M inference calls. Judges share a hidden consensus in perception but differ in how they prioritize attributes, and reweighting attributes steers them to a target standard.

LLMs Know When They Know, but Do Not Act on It: A Metacognitive Harness for Test-time Scaling

Qi Cao, Yufan Wang, Peijia Qin, Shuhao Zhang, Pengtao Xie

Arxiv Preprint

A training-free metacognitive harness that turns LLMs' pre-solve feeling-of-knowing and post-solve judgment-of-learning signals into an explicit test-time control interface, boosting a fixed Claude Sonnet-4.6 across text, code, and multimodal benchmarks.

AIBuildAI: An AI agent that automatically builds AI models

Ruiyi Zhang†, Peijia Qin†, Qi Cao†, Li Zhang†, Pengtao Xie

ICML 2026 Workshop on AI as a Tool for Mathematics, Computer Science, and Machine Learning (ICML Workshop)

We present AIBuildAI, an AI agent that automatically builds AI models, with the goal of solving general AI tasks in an end-to-end manner.

Models Under SCOPE: Scalable and Controllable Routing via Pre-hoc Reasoning

Qi Cao†, Shuhao Zhang†, Ruizhe Zhou, Ruiyi Zhang, Peijia Qin, Pengtao Xie

The Forty-Third International Conference on Machine Learning (ICML)

SCOPE, a model routing framework that predicts how accurate and how expensive each model will be before running it, allowing users to control cost-accuracy trade-offs and naturally handle new models.

DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning

Qi Cao, Ruiyi Wang, Ruiyi Zhang, Sai Ashish Somayajula, Pengtao Xie

The Thirty-Ninth Annual Conference on Neural Information Processing Systems (NeurIPS)

Spotlight @ Multimodal Algorithmic Reasoning Workshop

A multimodal Process Reward Model (PRM) trained with domain-reweighting. Top 1 method on MathVista, MMMU & R-Bench-V.

Bidomain Modeling Paradigm for Pansharpening

Junming Hou†, Qi Cao†, Ran Ran, Che Liu, Junling Li, Liang-jian Deng

Proceedings of the 31st ACM international conference on multimedia (ACM MM)

Oral

We propose BiPan, a bidomain pansharpening framework that models band-specific local spectral features and global spatial details in the Fourier domain, achieving state-of-the-art performance by better handling spectral diversity and MS image degradation.

Zero-shot Semi-supervised Learning for Pansharpening

Qi Cao, Liang-Jian Deng, Wu Wang, Junming Hou, Gemine Vivone

Information Fusion

Zero-shot pansharpening (ZS-Pan) only requires a single pair of PAN/LRMS images. Any pansharpening network can take the ZS-Pan as a plug-and-play module. A two-phase three-component semi-supervised model is designed for ZS-Pan.