Advisor: Prof. Alexander Hauptmann · Informedia Lab
Xiaoyu Zhu
Machine Learning Researcher, Apple
Ph.D. in Artificial Intelligence, Carnegie Mellon University
Greetings! I’m a Machine Learning Researcher at Apple, where I work on multimodal agentic post-training, with a focus on image/video reasoning and long-horizon visual agents. I obtained my Ph.D. in Artificial Intelligence from Carnegie Mellon University under the supervision of Prof. Alexander Hauptmann.
News
- [08/2026] Check out our latest technical report: Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning.
- [03/2026] One paper has been accepted by ICML 2026, and selected for an oral presentation at the CVPR 2026 Workshop.
- [12/2024] Completed my Ph.D. in Artificial Intelligence at Carnegie Mellon University, and started at Apple as a Machine Learning Researcher.
- [07/2024] One paper has been accepted by ECCV 2024.
- [02/2023] One paper has been accepted by CVPR 2023, and selected for an oral presentation at the SPIE 2024.
Selected Publications
Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs
ICML 2026
Open Vocabulary 3D Semantic Segmentation with Text-to-Image Diffusion Models
ECCV 2024
STMT: A Spatial-Temporal Mesh Transformer for MoCap-Based Action Recognition
CVPR 2023
Weakly Supervised 3D Semantic Segmentation Using Cross-Image Consensus and Inter-Voxel Affinity Relations
ICCV 2021
MSNet: A Multilevel Instance Segmentation Network for Natural Disaster Damage Assessment in Aerial Videos
WACV 2021
Won the
Automated Streams Analysis for Public Safety Challenge
with a $30k prize;
the damage assessment system was successfully tested on
Hurricane Laura.
Press coverage:
