Generative Models · Multimodal LLMs · Computer Vision

Haojie Zhang | 张浩杰

I am currently a Senior Algorithm Researcher at Alibaba ATH, working in Token Foundry. I received my M.Sc. in Information and Communication Engineering from South China University of Technology (SCUT) in 2026, advised by Prof. Kui Jia and Dr. Xun Xu, and my B.Eng. in Information Engineering from SCUT in 2023. I previously visited the Department of Automation at Tsinghua University, where I worked with Prof. Jianhua Tao.

My research explores generative models, multimodal large language models, and foundation vision systems. My long-term goal is to bridge perception, understanding, and imagination in intelligent systems.

Hand-drawn portrait of Haojie Zhang

Hangzhou, China

News

  1. Aug. 2026I joined Alibaba ATH as part of Token Foundry.
  2. Jul. 2026The technical report for LingBot-World 2.0 was released.
  3. May 2026MuSS was accepted to ACM MM 2026 as an oral presentation.
  4. Apr. 2026LetsTalk was accepted for publication in IEEE Transactions on Multimedia.
  5. Jan. 2026PaDT was accepted to ICLR 2026.

Selected Publications

* Equal contribution. Selected work is ordered by recent publication milestones.

LingBot-World 2.0 teaser showing diverse interactive worlds
Technical Report · 2026

Infinite Worlds with Versatile Interactions

Zelin Gao, Qiuyu Wang, Jiapeng Zhu, Jingye Chen, Zichen Liu, et al., Haojie Zhang, et al.

An interactive video world model supporting real-time, persistent, and controllable virtual worlds.

DiffCap-Bench overview
Under Review · 2026

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning

Yuancheng Wei*, Haojie Zhang*, Linli Yao, Lei Li, Jiali Chen, et al.

A challenging benchmark and LLM-as-a-Judge protocol for evaluating image difference captioning.

WeSAM++ framework
TPAMI · Under Review

Improving the Generalization of Segmentation Foundation Models via Weakly-Supervised and Unsupervised Adaptation

Haojie Zhang, Yongyi Su, Nanqing Liu, Shijie Li, Xulei Yang, Xiangyu Yue, Kui Jia, Xun Xu

WeSAM++ extends foundation-model adaptation with patch-level contrastive learning and masked image modeling.

Education & Experience

Education

2024.04 – 2024.11

Tsinghua University

Visiting Student, Department of Automation
Advisor: Prof. Jianhua Tao

2023.09 – 2026.06

South China University of Technology

M.Sc. in Information and Communication Engineering
Advised by Prof. Kui Jia and Dr. Xun Xu

2019.09 – 2023.06

South China University of Technology

B.Eng. in Information Engineering (Innovation Class)
Graduated with distinction · GPA 3.64/4.0

Experience

2026.08 – Present

Alibaba ATH

Token Foundry · Hangzhou

2026.04 – 2026.07

Ant Research Institute

Lingbot-World Team · Hangzhou

2025.10 – 2026.04

Tencent TEG

Hunyuan Team · Shenzhen

2025.04 – 2025.10

Tencent WXG

WeChat Vision · Shenzhen

2024.12 – 2025.03

Tencent IEG

LIGHTSPEED · Shenzhen

Open-source Projects

Research code and project pages accompanying selected work.

Email copied