Haichao ZhangI am a fifth-year Ph.D. candidate in Computer Engineering at Northeastern University, where I am part of the SMILE Lab and the Physical AI Research Initiative (PAIR), fortunate to be advised by Professor Yun Raymond Fu (Member of the Academy of Europe; Fellow of ACM, IEEE, AAAI, and AAAS). My research focuses on generative and multimodal AI for computer vision, spanning Vision-Language Models (VLMs), video understanding and generation, world models, trajectory prediction, and embodied AI. I am particularly interested in connecting these areas to build token-efficient video intelligence, generative models for visual content creation, and intelligent systems capable of understanding, predicting, and interacting with the physical world. Prior to my Ph.D., I received my M.Sc. degree from Zhejiang University (ZJU). During my graduate studies, I also conducted research as a visiting student and remote research intern with The Chinese University of Hong Kong (CUHK) and the University of California, San Diego (UCSD). Beyond academia, I have worked on both fundamental and applied AI research across several industry research teams. Most recently, I interned with Google AI Edge Applied Research and Google DeepMind (Summer 2026), where I worked on efficiency planning for generative AI and collaborated on VLM evaluation for UI coding agents. Previously, I was a Research Scientist Intern at Meta Reality Labs Research (Fall 2025), working on world models and VLMs; a Ph.D. Research Intern with LinkedIn Video AI (Summer 2025), focusing on VLMs and recommendation; and an Applied Scientist Intern at Amazon AWS AI Lab (Summer 2024), working on video understanding and video large language models. Earlier, I was a research intern at Tencent (Spring 2021), where I worked on generative models for images and videos.
I am seeking full-time positions starting in January 2027. Feel free to reach out if you are interested in working together!
Email Me / Twitter / LinkedIn / GitHub / Hugging Face / GoogleScholar / CV |
|
News2026.03: ThinkJEPA was featured by Turing Post as one of 14 JEPA milestones. 2026.09: Two first-author papers, GeneralistJEPA and ThinkJEPA, were accepted to NeurIPS 2026! 2026.08: Two papers were accepted to EMNLP 2026, including one Main Conference paper and one Findings paper. 2026.06: I joined Google AI Edge Applied Research (Core ML) as a Research Intern, working on an efficiency agent for generative AI. I also contributed part-time to Google DeepMind on VLM evaluation for UI coding agents. 2026.06: DIVE-Bench was accepted to ECCV 2026! 2026.04: One paper was accepted to ICML 2026! 2026.03: ThinkJEPA is now available on arXiv: paper, code, and Hugging Face data cache. 2026.02: Featured in the Northeastern College of Engineering Spotlight (profile/interview): COE Spotlight feature. 2026.02: My first-author paper Out-of-Sight Embodied Agents (journal extension of OOSTraj) was accepted by IEEE TPAMI. 2026.02: My first-author paper LinkedOut (the first-ever MLLM-based video recommender) was accepted to the CVPR 2026 Findings Track. 2025.09: My first-author paper VQToken (Extreme Token Reduction) was accepted to NeurIPS 2025. 2025.08: I joined Meta Reality Labs Research as a Research Scientist Intern. 2025.05: I joined LinkedIn Video AI as a Research Intern in Video GenAI. 2024.05: I joined Amazon AWS AI Labs as an Applied Scientist Intern. 2024.02: My first-author paper Out-of-Sight Trajectory Prediction was accepted to CVPR 2024. 2023.08: My first-author paper Layout Sequence Prediction From Noisy Mobile Modality was accepted to ACM MM 2023. 2022.09: I joined SMILE Lab at Northeastern University. |
Research (First-Author)My work connects efficient video understanding, multimodal reasoning, world models, and motion prediction and planning. |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Some Fun Projects
Sensors, embedded systems, signal processing → my early CV/AI journey
Click to expand
Several years ago, I delved into sensor modalities and signal processing, which sparked my interest in embedded platforms. That experience led me to explore further into AI and computer vision.
|
Research & Industry Experience
|
Google DeepMind |
|
|
Google | AI Edge Applied Research | Core ML |
|
|
Meta | Reality Labs Research, Redmond, WA |
|
|
LinkedIn | Video AI, Mountain View, CA |
|
|
Amazon | AWS AI Lab, Bellevue, WA |
|
|
Toyota InfoTech Lab, Mountain View, CA |
|
|
Tencent, Shanghai, China |
|
|
SMILE Lab, Northeastern University, Boston |
|
|
University of California, San Diego |
|
|
the Chinese University of Hong Kong, Shenzhen |
|
|
Zhejiang University, Hangzhou China |
|
Selected Honors & Awards
| NeurIPS Scholar Award | |
| ACM MM Travel Grant Award | ACM SIGMM |
| National Biomedical Engineering Innovative Design Competition | National First Prize |
| Challenge Cup Competition of Science Achievement in China | Provincial Grand Prize |
| Mobile Application Innovation Contest of North China | Provincial First Prize |
| 'Holtek cup' microcontroller application and design competition, Tianjin (6/453, < 1.3%) | Provincial First Prize |
| Tianjin IOT Innovation and Engineering Application Design Competition | Provincial First Prize |
| Tianjin Undergraduate Robotics Competition | Provincial First Prize |
| Tianjin International Student Internet Innovation and Entrepreneurship Competition | Provincial Second Prize |
| Northern China Robotics Competition | Provincial Second Prize |
Academic Service
• ES-Reasoning Workshop @ ICLR
• MMRAgI Workshop @ CVPR
• Voxel51 Best of CVPR Panel
• Conference: NeurIPS 2023–2026; IJCAI 2025; ACM MM 2024; ICLR 2024–2025; ECCV 2024, 2026; ICCV 2025; ICML 2025; AISTATS 2024; WACV 2025; ACL 2025
• Journal: IEEE TPAMI; IEEE TIP; Pattern Recognition; Multimedia Tools and Applications; ACM TKDD; IEEE TIV
