Haichao ZhangI am a fifth-year Ph.D. candidate in Computer Engineering at Northeastern University, where I am part of the SMILE Lab and the Physical AI Research Initiative (PAIR), fortunate to be advised by Professor Yun Raymond Fu (Member of the Academy of Europe; Fellow of ACM, IEEE, AAAI, and AAAS). My research focuses on generative and multimodal AI for computer vision, spanning Vision-Language Models (VLMs), video understanding and generation, world models, trajectory prediction, and embodied AI. I am particularly interested in connecting these areas to build token-efficient video intelligence, generative models for visual content creation, and intelligent systems capable of understanding, predicting, and interacting with the physical world. Prior to my Ph.D., I received my M.Sc. degree from Zhejiang University (ZJU). During my graduate studies, I also conducted research as a visiting student and remote research intern with The Chinese University of Hong Kong (CUHK) and the University of California, San Diego (UCSD). Beyond academia, I have worked on both fundamental and applied AI research across several industry research teams. Most recently, I interned with Google AI Edge Applied Research and Google DeepMind (Summer 2026), where I worked on efficiency planning for generative AI and collaborated on VLM evaluation for UI coding agents. Previously, I was a Research Scientist Intern at Meta Reality Labs Research (Fall 2025), working on world models and VLMs; a Ph.D. Research Intern with LinkedIn Video AI (Summer 2025), focusing on VLMs and recommendation; and an Applied Scientist Intern at Amazon AWS AI Lab (Summer 2024), working on video understanding and video large language models. Earlier, I was a research intern at Tencent (Spring 2021), where I worked on generative models for images and videos.
I am currently seeking a full-time Research Scientist position starting in January 2027 and am also open to research collaborations. If you are interested in working together, please feel free to reach out!
Email Me / Twitter / LinkedIn / GitHub / Hugging Face / GoogleScholar / CV |
|
News2026.06: I joined internship at Google AI Edge Applied Research (Core ML), working on an efficiency agent for generative AI. I also collaborated part-time with Google DeepMind on VLM evaluation for UI coding agents. 2026.03: ThinkJEPA was featured by Turing Post as one of 14 JEPA milestones. 2026.03: ThinkJEPA is now available on arXiv: paper, code, and Hugging Face data cache. 2026.06: DIVE-Bench was accepted to ECCV 2026! 2026.02: Featured in Northeastern College of Engineering Spotlight (profile/interview): COE Spotlight feature. 2026.02: Out-of-Sight Embodied Agents (journal extension of OOSTraj) has been accepted by IEEE TPAMI. 2026.02: LinkedOut (the first-ever MLLM-based video recommender) has been accepted to CVPR 2026 Findings Track. 2025.09: Our paper VQToken (Extreme Token Reduction) has been accepted to NeurIPS 2025. 2025.08: I joined Meta Reality Labs Research as a Research Scientist Intern. 2025.05: I joined LinkedIn Video AI as a Research Intern in Video GenAI. 2024.05: I joined Amazon AWS AI Labs as an Applied Scientist Intern this summer. 2024.02: Our paper Out-of-Sight Trajectory Prediction has been accepted at CVPR 2024. 2023.08: Our paper Layout Sequence Prediction From Noisy Mobile Modality has been accepted at ACM MM. 2022.09: I joined SMILE Lab at Northeastern University. |
Research (First-Author)My work connects efficient video understanding, multimodal reasoning, world models, and motion prediction and planning. |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Some Fun Projects
Sensors, embedded systems, signal processing → my early CV/AI journey
Click to expand
Several years ago, I delved into sensor modalities and signal processing, which sparked my interest in embedded platforms. That experience led me to explore further into AI and computer vision.
|
Research & Industry Experience
|
Google DeepMind |
|
|
Google | AI Edge Applied Research | Core ML |
|
|
Meta | Reality Labs Research, Redmond, WA |
|
|
LinkedIn | Video AI, Mountain View, CA |
|
|
Amazon | AWS AI Lab, Bellevue, WA |
|
|
Toyota InfoTech Lab, Mountain View, CA |
|
|
Tencent, Shanghai, China |
|
|
SMILE Lab, Northeastern University, Boston |
|
|
University of California, San Diego |
|
|
the Chinese University of Hong Kong, Shenzhen |
|
|
Zhejiang University, Hangzhou China |
|
Selected Honors & Awards
| NeurIPS Scholar Award | |
| ACM MM Travel Grant Award | ACM SIGMM |
| National Biomedical Engineering Innovative Design Competition | National First Prize |
| Challenge Cup Competition of Science Achievement in China | Provincial Grand Prize |
| Mobile Application Innovation Contest of North China | Provincial First Prize |
| 'Holtek cup' microcontroller application and design competition, Tianjin (6/453, < 1.3%) | Provincial First Prize |
| Tianjin IOT Innovation and Engineering Application Design Competition | Provincial First Prize |
| Tianjin Undergraduate Robotics Competition | Provincial First Prize |
| Tianjin International Student Internet Innovation and Entrepreneurship Competition | Provincial Second Prize |
| Northern China Robotics Competition | Provincial Second Prize |
Academic Service
• ES-Reasoning Workshop @ ICLR
• MMRAgI Workshop @ CVPR
• Voxel51 Best of CVPR Panel
• Conference: NeurIPS 2023–2026; IJCAI 2025; ACM MM 2024; ICLR 2024–2025; ECCV 2024, 2026; ICCV 2025; ICML 2025; AISTATS 2024; WACV 2025; ACL 2025
• Journal: IEEE TPAMI; IEEE TIP; Pattern Recognition; Multimedia Tools and Applications; ACM TKDD; IEEE TIV
