Xiaoyu Zhu

Xiaoyu Zhu

Machine Learning Researcher, Apple

Ph.D. in Artificial Intelligence, Carnegie Mellon University

Hi! I’m a Machine Learning Researcher at Apple, where I work on multimodal agentic post-training, with a focus on visual reasoning and long-horizon agents. I obtained my Ph.D. in Artificial Intelligence from Carnegie Mellon University under the supervision of Prof. Alex Hauptmann.

My research explores how models can learn more with less human supervision by discovering structure in observations, predicting what comes next, and turning generative knowledge into understanding. My long-term goal is to build intelligent agents that continually learn and improve through experience with minimal human intervention.

News

  • [08/2026] Check our latest technical report: Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning.
  • [03/2026] One paper has been accepted by ICML 2026, and selected for an oral presentation at the CVPR 2026 Workshop.
  • [12/2024] Completed my Ph.D. in Artificial Intelligence at Carnegie Mellon University, and started at Apple as a Machine Learning Researcher.
  • [07/2024] One paper has been accepted by ECCV 2024.
  • [02/2023] One paper has been accepted by CVPR 2023, and selected for an oral presentation at the SPIE 2024.

Selected Publications

Figure 1: three post-training paradigms for proactive video reasoning

Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

Zhu, X., Deng, X., Taddewadikar, S., Mondal, A. K., Jiang, Z., Liebelt, J.

We combine next-embedding prediction on unlabeled video with multimodal reasoning, teaching models to anticipate future events without the inference cost of generating intermediate images.

Press Coverage: [Synced, 机器之心] [X, Zhihu Frontier]

Figure 1: stress-testing VLMs with targeted textual perturbations

On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs

Zhao, R., Shah, A., Zhu, X., Deng, X., Jiang, Z., Yang, Y., Liebelt, J., Mondal, A. K.

ICML 2026, CVPR 2026 Workshop (Oral)

We uncover an accuracy-faithfulness trade-off in reinforcement learning, showing why improvements in multimodal reasoning must be assessed through robustness and reasoning consistency, not just final-answer accuracy.

Diff2Scene

Open Vocabulary 3D Semantic Segmentation with Text-to-Image Diffusion Models

Zhu, X., Zhou, H., Xing, P., Zhao, L., Xu, H., Liang, J., Hauptmann, A., Liu, T., Gallagher, A.

ECCV 2024

We turn knowledge learned through image generation into supervision for 3D understanding, distilling diffusion-based representations into an open-vocabulary scene model without human-labeled 3D training data.

STMT

STMT: A Spatial-Temporal Mesh Transformer for MoCap-Based Action Recognition

Zhu, X., Huang, P.*, Liang, J.*, De Melo, C., Hauptmann, A.

CVPR 2023, SPIE 2024 (Oral)

We teach a mesh Transformer to understand human motion by reconstructing masked vertices and predicting future frames, learning from unlabeled motion sequences to improve downstream action recognition.

CIVANet

Weakly Supervised 3D Semantic Segmentation Using Cross-Image Consensus and Inter-Voxel Affinity Relations

Zhu, X., Chen, J., Zeng, X., Liang, J., Li, C., Behpour, S., Xu, M.

ICCV 2021

We learn dense 3D segmentation from image-level labels, using cross-image consensus and inter-voxel affinities to recover object regions and boundaries without voxel-level human annotations.

MSNet

MSNet: A Multilevel Instance Segmentation Network for Natural Disaster Damage Assessment in Aerial Videos

Zhu, X., Liang, J., Hauptmann, A.

WACV 2021

We use temporal correspondences in unlabeled video as self-supervision to learn cross-frame similarity and refine detection confidence, improving aerial disaster damage assessment without additional annotations for this refinement.

Won the Automated Streams Analysis for Public Safety Challenge with a $30k prize; the damage assessment system was successfully tested on Hurricane Laura. Press coverage: Carnegie Mellon University News CBS Pittsburgh Microsoft News Science Daily AZO Robotics

See all publications →

Education

Ph.D. in Artificial Intelligence (Language and Information Technologies)
Carnegie Mellon University, School of Computer Science · Pittsburgh, PA

Advisor: Prof. Alexander Hauptmann · Informedia Lab

M.S. in Artificial Intelligence and Innovation
Carnegie Mellon University, School of Computer Science · Pittsburgh, PA

Experience

Apple Inc.
Machine Learning Researcher
Carnegie Mellon University, School of Computer Science
Research Assistant
Meta GenAI
AI Research Scientist Intern
Google DeepMind & Google Research
Student Researcher