AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models

ACM Multimedia 2026
The University of Tokyo Shanghai AI Lab Ant Group Shanghai Jiao Tong University

TL;DR: We control Video Diffusion Models with "3DGS Avatars + Background Videos" to insert and animate anyone into any scene, supporting character-scene occlusion, character-object interaction, and novel motion.

Showcase

Reference character 1
Reference character 2
Reference character 3

Animation with Dynamic Background

avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output

Animation with Scene Occlusion

avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output

Animation with Object Interaction

avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output

Animation with Extreme Body Shapes

avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output
avatar-background
output

Animation with Novel Motion

Original Video
Novel Motion Output
Original Video
Novel Motion Output

Animation with Interactive Occlusion Editing

Avatar-Background
Output
AniCrafter DiT editing pipeline

Comparison with Different Methods

Cross-Identity Reenactment

Reference
SMPL-X
HumanVid
MIMO
VACE
Ours
reference
reference

Self-Reenactment

Reference
GT
HumanVid
MIMO
VACE
Ours

Ablation Study

Relighting Augmentation

Avatar-Background
w/o Relighting
Ours

Avatar-Background Condition

Reference
SMPL-X
Avatar-Background
w/o Avatar
Ours
GT

Mask Strategy

Reference
SMPL-X
Avatar-Background
w/o Mask
Ours
GT

BibTeX


        TBA
  

Acknowledgement

This website is adapted from Nerfies and LLaVA, licensed under a Creative Commons Attribution-ShareAlike 4.0 International License. We thank the LLaMA team for giving us access to their models, and open-source projects, including Alpaca and Vicuna.