About Me
Yuhang Ma (马宇航) is an AI researcher at ByteDance, working on agent-based video understanding for short dramas, including automatic commentary/narration generation, video editing, and script understanding. She is also pursuing a PhD at King's College London, focusing on Multi-Agent Reinforcement Learning (MARL).
Previously, she worked at Fuxi AI Lab, NetEase Inc. (2022-2025), where she was responsible for pretraining the Danqing text-to-image generation model, fine-tuning the Danqing VLM and LLM models, and research on IP (character) consistency. She obtained her master's degree from University College London and her bachelor's degree from Hunan University, and was a visiting student at National University of Singapore in the 2019 winter semester.
Research Interests
- Multi-Agent Systems & Multi-Agent Reinforcement Learning
- Agent-based Video Understanding (narration, editing, script understanding)
- AIGC: Human Preference Modeling, Text-to-Image Generation, Consistent Character / Story Generation
- Multimodal Pretraining: LLMs and MLLMs
Experience
- [03/2025 - Present] ByteDance — AI Researcher
Agent-based video understanding for short dramas: commentary/narration generation, video editing, and script understanding. - [05/2022 - 03/2025] Fuxi AI Lab, NetEase Inc. — AI Researcher
Danqing text-to-image model pretraining, VLM/LLM fine-tuning, IP consistency research.
Education
- PhD, King's College London — Multi-Agent Reinforcement Learning
- MSc, University College London — Computer Science
- BEng, Hunan University — Mechanical Engineering
- Visiting Student, National University of Singapore — Computer Science
Contact
Personal Email: yuhang_ma0307 (at) 163.com / astronaut0307 (at) gmail.com
Work Email: mayuhang.26942 (at) bytedance.com
Publications
* indicates equal contribution, † indicates project leader, ✉ indicates advising
-
- HPSv3: Towards Wide-Spectrum Human Preference Score
- Yuhang Ma*, Yunhao Shui*, Xiaoshi Wu, Keqiang Sun✉, Hongsheng Li✉
- ICCV 2025
- [ProjectPage] / [Paper] / [Code] / [Model] / [Dataset]
-
- ComfyGPT: A Self-Optimizing Multi-Agent System for Comprehensive ComfyUI Workflow Generation
- Oucheng Huang*, Yuhang Ma*†, Zeng Zhao✉, Mingrui Wu, Jiayi Ji, Rongsheng Zhang, Zhipeng Hu, Xiaoshuai Sun, Rongrong Ji✉.
- arXiv 2025
- [ProjectPage] / [Paper] / [Code]
-
- Storynizor: Consistent Story Generation via Inter-Frame Synchronized and Shuffled ID Injection
- Yuhang Ma*†, Wenting Xu*, Chaoyi Zhao*, Keqiang Sun, Qinfeng Jin, Zeng Zhao✉, Changjie Fan, Zhipeng Hu.
- AAAI 2025
- [ProjectPage] / [Paper] / [Code]
-
- Character-Adapter: Prompt-Guided Region Control for High-Fidelity Character Customization
- Yuhang Ma*†, Wenting Xu*, Jiji Tang*, Qinfeng Jin, Rongsheng Zhang, Zeng Zhao✉, Changjie Fan, Zhipeng Hu.
- arXiv 2024
- [ProjectPage] / [Paper] / [Code]
-
- LLM4GEN: Leveraging Semantic Representation of LLMs for Text-to-Image Generation
- Mushui Liu*, Yuhang Ma*†, Xinfeng Zhang, Zhen Yang, Zeng Zhao✉, Bai Liu, Changjie Fan, Zhipeng Hu.
- AAAI 2025
- [ProjectPage] / [Paper] / [Code]
-
- You Can even Annotate Text with Voice: Transcription-only-Supervised Text Spotting
- Jingqun Tang*, Qiao Su*, Benlei Cui*, Yuhang Ma, Sheng Zhang, Dimitrios Kanoulas.
- ACM MM 2022
- [Paper]
Hobbies
Werewolf (狼人杀) fanatic, League of Legends, piano, vlogging, and a cat person with 4 cats.
