Advisor: Prof. Alexander Hauptmann · Informedia Lab
Xiaoyu Zhu
Machine Learning Researcher, Apple
Ph.D. in Artificial Intelligence, Carnegie Mellon University
Hi! I’m a Machine Learning Researcher at Apple, where I work on multimodal agentic post-training, with a focus on visual reasoning and long-horizon agents. I obtained my Ph.D. in Artificial Intelligence from Carnegie Mellon University under the supervision of Prof. Alex Hauptmann.
My research explores how models can learn more with less human supervision by discovering structure in observations, predicting what comes next, and turning generative knowledge into understanding. My long-term goal is to build intelligent agents that continually learn and improve through experience with minimal human intervention.
News
- [08/2026] Check our latest technical report: Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning.
- [03/2026] One paper has been accepted by ICML 2026, and selected for an oral presentation at the CVPR 2026 Workshop.
- [12/2024] Completed my Ph.D. in Artificial Intelligence at Carnegie Mellon University, and started at Apple as a Machine Learning Researcher.
- [07/2024] One paper has been accepted by ECCV 2024.
- [02/2023] One paper has been accepted by CVPR 2023, and selected for an oral presentation at the SPIE 2024.
Selected Publications
Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
We combine next-embedding prediction on unlabeled video with multimodal reasoning, teaching models to anticipate future events without the inference cost of generating intermediate images.
Press Coverage: [Synced, 机器之心] [X, Zhihu Frontier]
On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs
ICML 2026, CVPR 2026 Workshop (Oral)
We uncover an accuracy-faithfulness trade-off in reinforcement learning, showing why improvements in multimodal reasoning must be assessed through robustness and reasoning consistency, not just final-answer accuracy.
Open Vocabulary 3D Semantic Segmentation with Text-to-Image Diffusion Models
ECCV 2024
We turn knowledge learned through image generation into supervision for 3D understanding, distilling diffusion-based representations into an open-vocabulary scene model without human-labeled 3D training data.
STMT: A Spatial-Temporal Mesh Transformer for MoCap-Based Action Recognition
CVPR 2023, SPIE 2024 (Oral)
We teach a mesh Transformer to understand human motion by reconstructing masked vertices and predicting future frames, learning from unlabeled motion sequences to improve downstream action recognition.
Weakly Supervised 3D Semantic Segmentation Using Cross-Image Consensus and Inter-Voxel Affinity Relations
ICCV 2021
We learn dense 3D segmentation from image-level labels, using cross-image consensus and inter-voxel affinities to recover object regions and boundaries without voxel-level human annotations.
MSNet: A Multilevel Instance Segmentation Network for Natural Disaster Damage Assessment in Aerial Videos
WACV 2021
We use temporal correspondences in unlabeled video as self-supervision to learn cross-frame similarity and refine detection confidence, improving aerial disaster damage assessment without additional annotations for this refinement.
Won the
Automated Streams Analysis for Public Safety Challenge
with a $30k prize;
the damage assessment system was successfully tested on
Hurricane Laura.
Press coverage:
