Haichao Zhang

I am a fifth-year Ph.D. candidate in Computer Engineering at Northeastern University, where I am part of the SMILE Lab and the Physical AI Research Initiative (PAIR), fortunate to be advised by Professor Yun Raymond Fu (Member of the Academy of Europe; Fellow of ACM, IEEE, AAAI, and AAAS).

My research focuses on generative and multimodal AI for computer vision, spanning Vision-Language Models (VLMs), video understanding and generation, world models, trajectory prediction, and embodied AI. I am particularly interested in connecting these areas to build token-efficient video intelligence, generative models for visual content creation, and intelligent systems capable of understanding, predicting, and interacting with the physical world.

Prior to my Ph.D., I received my M.Sc. degree from Zhejiang University (ZJU). During my graduate studies, I also conducted research as a visiting student and remote research intern with The Chinese University of Hong Kong (CUHK) and the University of California, San Diego (UCSD).

Beyond academia, I have worked on both fundamental and applied AI research across several industry research teams. Most recently, I interned with Google AI Edge Applied Research and Google DeepMind (Summer 2026), where I worked on efficiency planning for generative AI and collaborated on VLM evaluation for UI coding agents. Previously, I was a Research Scientist Intern at Meta Reality Labs Research (Fall 2025), working on world models and VLMs; a Ph.D. Research Intern with LinkedIn Video AI (Summer 2025), focusing on VLMs and recommendation; and an Applied Scientist Intern at Amazon AWS AI Lab (Summer 2024), working on video understanding and video large language models. Earlier, I was a research intern at Tencent (Spring 2021), where I worked on generative models for images and videos.

I am currently seeking a full-time Research Scientist position starting in January 2027 and am also open to research collaborations. If you are interested in working together, please feel free to reach out!

Email Me   /  Twitter   /  LinkedIn   /  GitHub   /  Hugging Face   /  GoogleScholar   /  CV  

profile photo

News

2026.06: I joined internship at Google AI Edge Applied Research (Core ML), working on an efficiency agent for generative AI. I also collaborated part-time with Google DeepMind on VLM evaluation for UI coding agents.

2026.03: ThinkJEPA was featured by Turing Post as one of 14 JEPA milestones.

2026.03: ThinkJEPA is now available on arXiv: paper, code, and Hugging Face data cache.

2026.06: DIVE-Bench was accepted to ECCV 2026!

2026.02: Featured in Northeastern College of Engineering Spotlight (profile/interview): COE Spotlight feature.

2026.02: Out-of-Sight Embodied Agents (journal extension of OOSTraj) has been accepted by IEEE TPAMI.

2026.02: LinkedOut (the first-ever MLLM-based video recommender) has been accepted to CVPR 2026 Findings Track.

2025.09: Our paper VQToken (Extreme Token Reduction) has been accepted to NeurIPS 2025.

2025.08: I joined Meta Reality Labs Research as a Research Scientist Intern.

2025.05: I joined LinkedIn Video AI as a Research Intern in Video GenAI.

2024.05: I joined Amazon AWS AI Labs as an Applied Scientist Intern this summer.

2024.02: Our paper Out-of-Sight Trajectory Prediction has been accepted at CVPR 2024.

2023.08: Our paper Layout Sequence Prediction From Noisy Mobile Modality has been accepted at ACM MM.

2022.09: I joined SMILE Lab at Northeastern University.

Research (First-Author)

My work connects efficient video understanding, multimodal reasoning, world models, and motion prediction and planning.

ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model

ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
Haichao Zhang, Yijiang Li, Shwai He, Tushar Nagarajan, Mingfei Chen, Jianglin Lu, Ang Li, Yun Fu

Preprint, 2026.
ThinkJEPA marks a new milestone in JEPA-style world models by unifying a cortex-like VLM for high-level semantics with a cerebellum-like JEPA branch for low-level dynamics and physical consistency.
Featured in Turing Post’s “14 JEPA Milestones” .


arXiv GitHub Code Hugging Face Cache Turing Post Feature

DIVE-Bench: Codec-Style High-fps Video Understanding

DIVE-Bench: Codec-Style High-fps Video Understanding
Haichao Zhang, Wenhao Chai, Shwai He, Ang Li, Yun Fu

Accepted to ECCV 2026.
A benchmark for information-dense video understanding, spanning educational content and high-motion trajectories.
Gated Residual Tokenization reuses static visual content and preserves changes across frames. See the project page for the revised manuscript, code, and dataset access.


Earlier arXiv version: 2509.14199   Project Website   HuggingFace Dataset   GitHub Code

VQToken: token dynamics visualization

VQToken: Neural Discrete Token Representation Learning for Extreme Token Reduction in Video Large Language Models
Haichao Zhang, Yun Fu
Extreme Token Reduction for Video LLM.

the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

arXiv   Project Website   Hugging Face Model   GitHub Code

LinkedOut: World Knowledge Representation out of Video LLM for Recommendation

LinkedOut: Linking World Knowledge Representation Out of Video LLM for Next-Generation Video Recommendation
Haichao Zhang, Yao Lu, Lichen Wang, Yunzhe Li, Daiwei Chen, Yunpeng Xu, Yun Fu

Accepted to CVPR 2026 (Findings Track)
I led the development of LinkedOut, the first video LLM for video recommendation.
Internal evaluation on real-world LinkedIn data showed 12% higher accuracy than the model currently serving video recommendations in the LinkedIn app.


CVPR Findings 2026 arXiv

Out-of-Sight Trajectories: tracking fusion prediction

Out-of-Sight Embodied Agents: Multimodal Tracking, Sensor Fusion, and Trajectory Forecasting
Haichao Zhang, Yi Xu, Yun Fu
Out-of-sight embodied agents via multimodal tracking, sensor fusion, and trajectory forecasting.

IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

TPAMI

OOSTraj: Out-of-Sight Trajectory Prediction With Vision-Positioning Denoising
Haichao Zhang, Yi Xu, Hongsheng Lu, Takayuki Shimizu, Yun Fu
(First work on out-of-sight trajectory prediction.)

IEEE/CVF Conference on Computer Vision and Pattern Recognition 2024 (CVPR'24)
arXiv   PapersWithCode Benchmark   Project Website   GitHub Code

Layout Sequence Prediction From Noisy Mobile Modality
(See Beyond Vision: Denoising Diffusion Model for Layout Trajectory Prediction from Noisy Mobile Modality)
Haichao Zhang, Yi Xu, Hongsheng Lu, Takayuki Shimizu, Yun Fu

31st ACM International Conference on Multimedia (ACM MM'23)
arXiv   Project Website   Video

Camouflaged Image Synthesis Is All You Need to Boost Camouflaged Detection
Haichao Zhang, Can Qin, Yu Yin, Yun Fu
arXiv

Sketch Me A Video

Sketch Me A Video
Haichao Zhang, Gang Yu, Tao Chen, Guozhong Luo
arXiv

Fine-grained Identity Preserving Landmark Synthesis

Fine-grained Identity Preserving Landmark Synthesis for Face Reenactment
Haichao Zhang, Youcheng Ben, Weixi Zhang, Tao Chen, Gang Yu, Bin Fu
arXiv

Restore DeepFakes Video Frames

Restore DeepFakes Video Frames by Identifying Individual Motion Styles
Haichao Zhang, Zhe-Ming Lu, Hao Luo, Ya-Pei Feng
Electronics Letters   Journal Page

Selected Co-authored Publications

Demystifying When Pruning Works via Representation Hierarchies
Shwai He, Guoheng Sun, Haichao Zhang, Yun Fu, Ang Li.
ICML 2026 · Paper · Project · Code

Some Fun Projects
Sensors, embedded systems, signal processing → my early CV/AI journey
Click to expand

Several years ago, I delved into sensor modalities and signal processing, which sparked my interest in embedded platforms. That experience led me to explore further into AI and computer vision.

Wheelchair Control System via analysis eye-blinking EMG and EEG

Provincial Grand Prize at the Challenge Cup Competition of Science Achievement in China
project video
Mar. 2017

Proposed to detect eye blink EMG noise mixed in EEG signal, using intense eye blink signals to control wheelchair direction, while analyzing EEG to predict tension/relaxation degree to control speed.

An affordable solution for paralyzed patients to control their wheelchairs and move independently.

Low power abnormal ECG detection system based on MSP430

National First Prize at National Biomedical Engineering Innovative Design Competition
project video
Nov. 2016

Responsible for developing upper computer software to receive/filter signals in the spectral domain from the MSP430 PCB board, and developing an algorithm to detect abnormal ECG.

Sign language recognition system of wearable bending sensor gloves

First Prize at Mobile Application Innovation Contest of North China
Jul. 2016

Responsible for programming the embedded microprocessor to sample analog signals of bending sensors on gloves, predict sign language, and display results in the app.

Vision-based paper money and coin sorting machine

Summer 2015

Responsible for programming embedded microprocessors to control the mechanical structure and developing upper machine software to detect paper money types using traditional image processing, then sort them.

Multimedia Information Hiding Technology of Unstructured Data

Alibaba-ZJU Joint Research Institute of Frontier Technologies Research Project
Summer 2018

Responsible for developing C++ software "Shared Memory Based Code Hiding Platform."

Participated in video stream watermarking algorithm.

Research & Industry Experience

Google DeepMind
Part-time Student Researcher, Summer 2026 (informal)
Supervisor: Dr. Tony Nguyen
Focus: OS Coding Agents Evaluation Benchmark

Google | AI Edge Applied Research | Core ML
Research Intern, Jun. 2026 ~ Sep. 2026 · Sunnyvale, CA
Mentor: Dr. Daniel Pan
Focus: Efficiency Agent for Generative AI

Meta | Reality Labs Research, Redmond, WA
Research Scientist Intern, Sep. 2025 ~ Dec. 2025
Mentor: Prof. Stefan Scherer
Focus Areas: Vision-Language Models (VLM), World Model, Multimodal Learning

Meta Reality Labs Research

LinkedIn | Video AI, Mountain View, CA
Research Intern in Video GenAI, May. 2025 ~ Aug. 2025
Mentor: Dr. Yao Lu, Dr. Lichen Wang
Focus Areas: Vision-Language Models (VLM), Video Recommendation

LinkedIn Video AI

Amazon | AWS AI Lab, Bellevue, WA
Applied Scientist Intern, Jun. 2024 ~ Aug. 2024

Amazon AWS AI Lab

Toyota InfoTech Lab, Mountain View, CA
Part-time Research, Feb. 2023 ~ Nov. 2023

Toyota InfoTech Lab

Tencent, Shanghai, China
Research Scientist Intern, Jul. 2020 ~ May. 2021

Tencent

SMILE Lab, Northeastern University, Boston
Ph.D. Candidate, Sep. 2022 ~ Present
Expected graduation: January 2027
Supervisor: Prof. Yun Raymond Fu

SMILE Lab

University of California, San Diego
Summer Intern, May. 2021 ~ Sep. 2021
Supervisor: Prof. Xiaolong Wang

UCSD

the Chinese University of Hong Kong, Shenzhen
Visiting Student, Feb. 2020 ~ May. 2020
Supervisor: Prof. Xiaoguang Han

the Chinese University of Hong Kong, Shenzhen

Zhejiang University, Hangzhou China
Master Student, Sep. 2018 ~ Mar. 2021
Supervisor: Prof. Zhe-Ming Lu

Zhejiang University

Selected Honors & Awards

NeurIPS Scholar Award  
ACM MM Travel Grant Award  ACM SIGMM
National Biomedical Engineering Innovative Design Competition  National First Prize
Challenge Cup Competition of Science Achievement in China  Provincial Grand Prize
Mobile Application Innovation Contest of North China  Provincial First Prize
'Holtek cup' microcontroller application and design competition, Tianjin (6/453, < 1.3%)  Provincial First Prize
Tianjin IOT Innovation and Engineering Application Design Competition  Provincial First Prize
Tianjin Undergraduate Robotics Competition  Provincial First Prize
Tianjin International Student Internet Innovation and Entrepreneurship Competition  Provincial Second Prize
Northern China Robotics Competition  Provincial Second Prize

Academic Service

Area Chair
• ES-Reasoning Workshop @ ICLR • MMRAgI Workshop @ CVPR

Invited Talk
• Voxel51 Best of CVPR Panel

Reviewer
• Conference: NeurIPS 2023–2026; IJCAI 2025; ACM MM 2024; ICLR 2024–2025; ECCV 2024, 2026; ICCV 2025; ICML 2025; AISTATS 2024; WACV 2025; ACL 2025
• Journal: IEEE TPAMI; IEEE TIP; Pattern Recognition; Multimedia Tools and Applications; ACM TKDD; IEEE TIV