Interactive Real-Time Visual Avatar Generation: A Survey
A survey of 100 methods (2023-2026) for photorealistic, audio-driven digital avatars, spanning GAN-based synthesis, NeRF, 3D Gaussian Splatting, flow-matching, and hybrid parametric body representations.
Abstract
Photorealistic, audio-driven digital avatars have moved from offline research prototypes to real-time interactive production systems in the last few years, driven by the convergence of neural rendering, generative modelling, and parametric body representations. We review 100 methods published between 2023 and 2026, grouped into five families: GAN-based synthesis, Neural Radiance Fields (NeRF) and triplane representations, 3D Gaussian Splatting (3DGS), diffusion and flow-matching models, and hybrid parametric full-body representations. Methods are compared along four production axes rather than within a single representation family: rendering latency, anatomical coverage (head-only vs. upper-body vs. full-body), personalisation paradigm (generalizable vs. subject-specific), and input modality. A central focus is the head-body integration problem: how to combine the FLAME head model and the SMPLX body model despite their geometric and topological mismatch, and recent hybrid representations that claim to resolve it, in particular the Expressive Human Model (EHM). Three findings stand out. 3DGS coupled with 3DMM-driven expression control is now the dominant paradigm for real-time production deployment, reaching sub-50 ms frame latency with near-photorealistic quality. Flow matching is emerging as a preferred alternative to score-based diffusion for quality-critical, offline generation, with growing adoption among recent state-of-the-art systems. And generalizable one-shot methods are closing the quality gap with subject-specific fine-tuning. We close by investigating the open challenges that remain: temporal identity drift, co-speech gesture synthesis, and scalable personalisation from limited data, and by providing concrete research directions toward production-deployable interactive avatars.
Keywords: Talking head generation, 3D Gaussian splatting, audio-driven avatars, neural rendering, flow matching, parametric body models
Recommended citation:
@misc{dey2026avatarsurvey,
title={Interactive Real-Time Visual Avatar Generation: A Survey},
author={Dey, Arnab and Alcoverro, Marcel and Zabaleta, Itziar and Fern{\'a}ndez, Carla},
year={2026},
month={6},
note={Preprint},
doi={10.13140/RG.2.2.19626.17608}
}
Citation
Arnab Dey, Marcel Alcoverro, Itziar Zabaleta, and Carla Fernández. "Interactive Real-Time Visual Avatar Generation: A Survey." Preprint (2026). DOI: 10.13140/RG.2.2.19626.17608.
