Muyao Niu (牛慕尧)

Hi, I'm Muyao Niu. I am a 2nd year Ph.D. student in the Department of Mechano-Informatics, the University of Tokyo (UTokyo). My supervisor is Prof. Yinqiang Zheng.

I received my master's degree from The University of Tokyo in 2024, and B.E. degree from Dalian University of Technology in 2022.

profile photo

My research interests include:

News 🔥

Selected Papers 📃

MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model — overview

MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model

Muyao Niu, Xiaodong Cun, Xintao Wang, Yong Zhang, Ying Shan, Yinqiang Zheng

ECCV, 2024

We introduce MOFA-Video to adapt motions from different domains to the frozen Video Diffusion Model. MOFA-Video can effectively animate a single image using various types of control signals, including trajectories, keypoint sequences, and their combinations.

WorldSculpt: Generating Compositional Worlds from Grounded Videos — overview

WorldSculpt: Generating Compositional Worlds from Grounded Videos

Muyao Niu, Jixuan He, Ruihan Yu, Lian Fu, Yonghao Yu, Zheng-Hui Huang, Yifan Zhan, Fengbo Lan, Yongtao Ge, Yinqiang Zheng, Kaipeng Zhang, Zhixiang Wang

Technical Report, 2026

WorldSculpt produces a compositional mesh representation for very complex scenes consisting of hundreds of individual objects given grounded RGB videos.

AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models — overview

AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models

Muyao Niu, Mingdeng Cao, Yifan Zhan, Qingtian Zhu, Weihang Ran, Yanhong Zeng, Xiao Sun, Zhihang Zhong, Yinqiang Zheng

ACM MM, 2026

We leverage "3DGS Avatar + Background Video" as hybrid guidance for the video diffusion model to insert and animate anyone into any scene following given motion sequence.

MAD-Avatar: Motion-Aware Animatable Gaussian Avatars Deblurring — overview

MAD-Avatar: Motion-Aware Animatable Gaussian Avatars Deblurring

Muyao Niu, Yifan Zhan, Qingtian Zhu, Zhuoxiao Li, Wei Wang, Zhihang Zhong, Xiao Sun, Yinqiang Zheng

CVPR, 2026

We introduce an innovative method for deriving sharp intrinsic 3D human Gaussian avatars from blurry videos via 3D human motion model and the physics-based blur formation model.

3DarkFusion: Toward Robust and 3D-Aware RGB-NIR Imaging in the Dark — overview

3DarkFusion: Toward Robust and 3D-Aware RGB-NIR Imaging in the Dark

Muyao Niu, Mingze Ma, Yifan Zhan, Qingtian Zhu, Zhihang Zhong, Wei Guo, Chang Wen Chen, Yinqiang Zheng

ACM MM, 2026

3DarkFusion is a novel RGB-NIR fusion framework that leverages the complementary information from RGB and NIR images to achieve robust and 3D-aware imaging under extremely low-light conditions.

RS-NeRF: Neural Radiance Fields from Rolling Shutter Images — overview

RS-NeRF: Neural Radiance Fields from Rolling Shutter Images

Muyao Niu, Tong Chen, Yifan Zhan, Zhuoxiao Li, Xiang Ji, Yinqiang Zheng

ECCV, 2024

We improve NeRF to consider the RS distortions with two technologies: camera trajectory smoothness regularization and multi-sampling strategy.

Visibility Constrained Wide-band Illumination Spectrum Design for Seeing-in-the-Dark — overview

Visibility Constrained Wide-band Illumination Spectrum Design for Seeing-in-the-Dark

Muyao Niu, Zhuoxiao Li, Zhihang Zhong, Yinqiang Zheng

CVPR, 2023

We designed an optimal illumination spectrum in the VIS-NIR range by considering human vision constraints, which significantly improves translation performance. A fully differentiable model was proposed, which includes the imaging process, human visual perception, and the enhancement network.

Physics-Based Adversarial Attack on Near-Infrared Human Detector for Nighttime Surveillance Camera Systems — overview

Physics-Based Adversarial Attack on Near-Infrared Human Detector for Nighttime Surveillance Camera Systems

Muyao Niu, Zhuoxiao Li Yifan Zhan, Huy H. Nguyen, Isao Echizen, Yinqiang Zheng

ACM MM, 2023

We introduced an innovative approach that passively manipulates the intensity distribution of NIR images and developed a 3D-aware, black-box attack algorithm to target deep learning-based NIR-powered human detection systems.

NIR-assisted Video Enhancement via Unpaired 24-hour Data — overview

NIR-assisted Video Enhancement via Unpaired 24-hour Data

Muyao Niu, Zhihang Zhong, Yinqiang Zheng

ICCV, 2023

We addressed the issue of collecting data for utilizing NIR images to improve low-light VIS videos. Physics-inspired algorithms are designed to simulate pseudo paired data of NIR and VIS images, simulating day-to-night situations. We then trained an enhancement network using the generated pseudo data.

Full Publication List 📚

WorldSculpt: Generating Compositional Worlds from Grounded Videos — overview

WorldSculpt: Generating Compositional Worlds from Grounded Videos

Muyao Niu, Jixuan He, Ruihan Yu, Lian Fu, Yonghao Yu, Zheng-Hui Huang, Yifan Zhan, Fengbo Lan, Yongtao Ge, Yinqiang Zheng, Kaipeng Zhang, Zhixiang Wang

Technical Report, 2026

AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models — overview

AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models

Muyao Niu, Mingdeng Cao, Yifan Zhan, Qingtian Zhu, Weihang Ran, Yanhong Zeng, Xiao Sun, Zhihang Zhong, Yinqiang Zheng

ACM MM, 2026

3DarkFusion: Toward Robust and 3D-Aware RGB-NIR Imaging in the Dark — overview

3DarkFusion: Toward Robust and 3D-Aware RGB-NIR Imaging in the Dark

Muyao Niu, Mingze Ma, Yifan Zhan, Qingtian Zhu, Zhihang Zhong, Wei Guo, Chang Wen Chen, Yinqiang Zheng

ACM MM, 2026

Surprise Forcing: What to Remember, When to Skip in Long Video Generation — overview

Surprise Forcing: What to Remember, When to Skip in Long Video Generation

Shuwei Shi, Zhen Li, Muyao Niu, Chuanhao Li, Bo Zheng, Kaipeng Zhang, Yinqiang Zheng

ECCV, 2026

Surprise Forcing is a training-free framework that uses "surprise" to keep the most informative frames discarded by the sliding-window KV cache (via a surprise-aware memory bank) and to adaptively cut denoising steps for easy chunks—preserving long-range scene consistency and narrative coherence in autoregressive diffusion long-video generation.

SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation — overview

SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation

Xiaolong Zhou, Yifei Liu, Ziyang Gong, Jiarui Li, Qiyue Zhao, Muyao Niu, Yuanyuan Gao, Le Ma, Xue Yang, Hongjie Zhang, Zhihang Zhong

arXiv, 2026

We introduce SpaceDG, the first large-scale dataset for degradation-aware spatial intelligence, and SpaceDG-Bench, a human-verified benchmark for evaluating MLLMs under visual degradations.

Composing Driving Worlds through Disentangled Control for Adversarial Scenario Generation — overview

Composing Driving Worlds through Disentangled Control for Adversarial Scenario Generation

Yifan Zhan*, Zhengqing Chen*, Qingjie Wang*, Zhuo He, Muyao Niu, Xiaoyang Guo, Wei Yin, Weiqiang Ren, Qian Zhang, Yinqiang Zheng

ECCV, 2026

CompoSIA is a powerful simulator for synthesizing rare driving scenes, enabling explicit control over structure, identity, and ego action.

MAD-Avatar: Motion-Aware Animatable Gaussian Avatars Deblurring — overview

MAD-Avatar: Motion-Aware Animatable Gaussian Avatars Deblurring

Muyao Niu, Yifan Zhan, Qingtian Zhu, Zhuoxiao Li, Wei Wang, Zhihang Zhong, Xiao Sun, Yinqiang Zheng

CVPR, 2026

Visionary: The World Model Carrier Built on WebGPU-Powered Gaussian Splatting Platform — overview

Visionary: The World Model Carrier Built on WebGPU-Powered Gaussian Splatting Platform

Yuning Gong, Yifei Liu, Yifan Zhan, Muyao Niu, Xueying Li, Yuanjun Liao, Jiaming Chen, Yuanyuan Gao, Jiaqi Chen, Minming Chen, Li Zhou, Yuning Zhang, Wei Wang, Xiaoqing Hou, Huaxi Huang, Shixiang Tang, Le Ma, Dingwen Zhang, Xue Yang, Junchi Yan, Yanchi Zhang, Yinqiang Zheng, Xiao Sun, Zhihang Zhong

arXiv, 2025

We build an open platform via WebGPU and ONNX Runtime, enabling real-time rendering of diverse Gaussian Splatting variants (3DGS, MLP-based 3DGS, 4DGS, Neural Avatars).

R3-avatar: Record and retrieve temporal codebook for reconstructing photorealistic human avatars — overview

R3-avatar: Record and retrieve temporal codebook for reconstructing photorealistic human avatarsPreprint

Yifan Zhan, Wangze Xu, Qingtian Zhu, Muyao Niu, Mingze Ma, Yifei Liu, Zhihang Zhong, Xiao Sun, Yinqiang Zheng

arXiv, 2025

We present R3-Avatar, incorporating a temporal codebook, to overcome the inability of human avatars to be both animatable and of high-fidelity rendering quality.

Tree-NeRV: A Tree-Structured Neural Representation for Efficient Non-Uniform Video Encoding — overview

Tree-NeRV: A Tree-Structured Neural Representation for Efficient Non-Uniform Video Encoding

Jiancheng Zhao, Yifan Zhan, Qingtian Zhu, Mingze Ma, Muyao Niu, Zunian Wan, Xiang Ji, Yinqiang Zheng

ICCV, 2025

We propose Tree-NeRV, a novel tree-structured feature representation for efficient and adaptive video encoding. Unlike conventional approaches, Tree-NeRV organizes feature representations within a Binary Search Tree (BST), enabling non-uniform sampling along the temporal axis.

ToMiE: Towards Explicit Exoskeleton for the Reconstruction of Complicated 3D Human Avatars — overview

ToMiE: Towards Explicit Exoskeleton for the Reconstruction of Complicated 3D Human Avatars

Yifan Zhan, Qingtian Zhu, Muyao Niu, Mingze Ma, Jiancheng Zhao, Zhihang Zhong, Xiao Sun, Yu Qiao, Yinqiang Zheng

ICCV, 2025

We highlight a critical yet often overlooked factor in most 3D human tasks, namely modeling complicated 3D human with hand-held objects or loose-fitting clothing by modeling the exoskeleton of the human body.

Adversarial Attacks on Event-Based Pedestrian Detectors: A Physical Approach — overview

Adversarial Attacks on Event-Based Pedestrian Detectors: A Physical Approach

Guixu Lin, Muyao Niu, Qingtian Zhu, Zhengwei Yin, Zhuoxiao Li, Shengfeng He, Yinqiang Zheng

AAAI, 2025

We developed an end-to-end adversarial framework for event-driven pedestrian detection, framing the design of adversarial clothing textures as a 2D texture optimization problem.

StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular Videos — overview

StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular Videos

Sijie Zhao*, Wenbo Hu*, Xiaodong Cun*, Yong Zhang#, Xiaoyu Li#, Zhe Kong, Xiangjun Gao, Muyao Niu, Ying Shan

Technical Report, 2024

We present a novel framework for converting 2D videos to immersive stereoscopic 3D, addressing the growing demand for 3D content in immersive experience. Leveraging foundation models as priors, our approach boosts the performance to ensure the high-fidelity generation required by the display devices.

MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model — overview

MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model

Muyao Niu, Xiaodong Cun, Xintao Wang, Yong Zhang, Ying Shan, Yinqiang Zheng

ECCV, 2024

CV-VAE: A Compatible Video VAE for Latent Generative Video Models — overview

CV-VAE: A Compatible Video VAE for Latent Generative Video Models

Sijie Zhao, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Muyao Niu, Xiaoyu Li, Wenbo Hu, Ying Shan

NeurIPS, 2024

We propose CV-VAE that is compatible with existing image and video models trained with SD image VAE. Our video VAE provides a truly spatio-temporally compressed latent space for latent generative video models, as opposed to uniform frame sampling.

RS-NeRF: Neural Radiance Fields from Rolling Shutter Images — overview

RS-NeRF: Neural Radiance Fields from Rolling Shutter Images

Muyao Niu, Tong Chen, Yifan Zhan, Zhuoxiao Li, Xiang Ji, Yinqiang Zheng

ECCV, 2024

KFD-NeRF: Rethinking Dynamic NeRF with Kalman Filter — overview

KFD-NeRF: Rethinking Dynamic NeRF with Kalman Filter

Yifan Zhan, Zhuoxiao Li, Muyao Niu, Zhihang Zhong, Shohei Nobuhara, Ko Nishino, Yinqiang Zheng

ECCV, 2024

We combine dynamic neural radiance field with a motion reconstruction framework based on Kalman filtering, enabling accurate deformation estimation from scene observations and predictions.

Physics-Based Adversarial Attack on Near-Infrared Human Detector for Nighttime Surveillance Camera Systems — overview

Physics-Based Adversarial Attack on Near-Infrared Human Detector for Nighttime Surveillance Camera Systems

Muyao Niu, Zhuoxiao Li, Yifan Zhan, Huy H. Nguyen, Isao Echizen, Yinqiang Zheng

ACM MM, 2023

NIR-assisted Video Enhancement via Unpaired 24-hour Data — overview

NIR-assisted Video Enhancement via Unpaired 24-hour Data

Muyao Niu, Zhihang Zhong, Yinqiang Zheng

ICCV, 2023

Visibility Constrained Wide-band Illumination Spectrum Design for Seeing-in-the-Dark — overview

Visibility Constrained Wide-band Illumination Spectrum Design for Seeing-in-the-Dark

Muyao Niu, Zhuoxiao Li, Zhihang Zhong, Yinqiang Zheng

CVPR, 2023

Region Assisted Sketch Colorization — overview

Region Assisted Sketch Colorization

Ning Wang*, Muyao Niu*, Zhihui Wang, Kun Hu, Bin Liu, Zhiyong Wang, Haojie Li

TIP, 2023

We proposed the Region-Assisted Sketch Colorization (RASC) method, which uses a 'Region Map' to better utilize regional information within the sketch, enhancing the perception of region-wise features.

Coloring anime line art videos with transformation region enhancement network — overview

Coloring anime line art videos with transformation region enhancement network

Ning Wang*, Muyao Niu*, Zhi Dou, Zhihui Wang, Zhiyong Wang, Zhaoyan Ming, Bin Liu, Haojie Li

Pattern Recognition, 2023

We propose a multi-scale Transformation Region Enhancement Network (TRE-Net) to enhance the learning on geometric transformation regions.

Awards & Service 🏆

CVPR 2025 Outstanding Reviewer
BOOST NAIS Special Research Scholarship (3,900,000 JPY / year), The University of Tokyo, 2024.10 - 2027.09
WING-CFS Special Research Assistant Scholarship (2,160,000 JPY / year), The University of Tokyo, 2023.04 - 2024.09

Reviewer: CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR, ACM MM, BMVC, TMLR.