Multimodal generative AI for human motion understanding and generation: A survey and way forward

Journal Publication ResearchOnline@JCU
Islam, Muhammad;Huang, Tao;Ahn, Euijoon;Naseem, Usman
Abstract

This paper presents an in-depth survey of the use of multimodal Generative Artificial Intelligence (GenAI) with autoregressive Large Language Models (LLMs) for human motion understanding and generation, offering insights into emerging methods and architectures and their potential to advance realistic and versatile motion synthesis. Focusing exclusively on text and motion modalities, this research investigates how textual descriptions can guide the generation of complex, human-like motion sequences. The paper explores various generative approaches, in-cluding multimodal autoregressive LLMs, multimodal diffusion, and multimodal transformers and their variants, and analyzes their strengths and limitations with respect to motion quality, computational efficiency, and adapt-ability. It highlights recent advances in text-conditioned motion generation, where textual inputs are used to control and refine motion outputs with greater precision. The use of LLMs further enhances these models by enabling semantic alignment between instructions and motion, improving coherence and contextual relevance. This systematic survey underscores the transformative potential of text-to-motion GenAI and LLM architectures in applications such as healthcare, humanoids, gaming, animation, and assistive technologies, while addressing ongoing challenges and research directions to guide future developments in human-centric GenAI.

Journal

Information Fusion

Publication Name

Information Fusion

Volume

135

ISBN/ISSN

1872-6305

Edition

N/A

Issue

N/A

Pages Count

27

Location

N/A

Publisher

Elsevier

Publisher Url

N/A

Publisher Location

N/A

Publish Date

N/A

Url

N/A

Date

N/A

EISSN

N/A

DOI

10.1016/j.inffus.2026.104435