Human Motion Instruction Tuning

November 25, 2024 · Declared Dead · 🏛 Computer Vision and Pattern Recognition

Authors Lei Li, Sen Jia, Jianhao Wang, Zhongyu Jiang, Feng Zhou, Ju Dai, Tianfang Zhang, Zongkai Wu, Jenq-Neng Hwang arXiv ID 2411.16805 Category cs.AI: Artificial Intelligence Cross-listed cs.CV Citations 14 Venue Computer Vision and Pattern Recognition Repository https://github.com/ILGLJ/LLaMo ⭐ 5 Last Checked 1 month ago

Abstract

This paper presents LLaMo (Large Language and Human Motion Assistant), a multimodal framework for human motion instruction tuning. In contrast to conventional instruction-tuning approaches that convert non-linguistic inputs, such as video or motion sequences, into language tokens, LLaMo retains motion in its native form for instruction tuning. This method preserves motion-specific details that are often diminished in tokenization, thereby improving the model's ability to interpret complex human behaviors. By processing both video and motion data alongside textual inputs, LLaMo enables a flexible, human-centric analysis. Experimental evaluations across high-complexity domains, including human behaviors and professional activities, indicate that LLaMo effectively captures domain-specific knowledge, enhancing comprehension and prediction in motion-intensive scenarios. We hope LLaMo offers a foundation for future multimodal AI systems with broad applications, from sports analytics to behavioral prediction. Our code and models are available on the project website: https://github.com/ILGLJ/LLaMo.

📄 View on arXiv 🌐 View on ar5iv 📑 PDF 💻 Repository 🎉 Report Code Found

Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

📜 Similar Papers

In the same crypt — Artificial Intelligence

R.I.P. 👻 Ghosted

A Unified Approach to Interpreting Model Predictions

Scott Lundberg, Su-In Lee

cs.AI 🏛 NeurIPS 📚 30.8K cites 8 years ago

R.I.P. 👻 Ghosted

Explainable Artificial Intelligence (XAI): Concepts, Taxonomies, Opportunities and Challenges toward Responsible AI

Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, ... (+10 more)

cs.AI 🏛 Inf. Fusion 📚 7.8K cites 6 years ago

R.I.P. 👻 Ghosted

Addressing Function Approximation Error in Actor-Critic Methods

Scott Fujimoto, Herke van Hoof, David Meger

cs.AI 🏛 ICML 📚 6.4K cites 8 years ago

R.I.P. 👻 Ghosted

Explanation in Artificial Intelligence: Insights from the Social Sciences

Tim Miller

cs.AI 🏛 AI 📚 4.9K cites 8 years ago

R.I.P. 👻 Ghosted

Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Peter Clark, Isaac Cowhey, ... (+5 more)

cs.AI 🏛 arXiv 📚 4.0K cites 8 years ago

R.I.P. 👻 Ghosted

Complex Embeddings for Simple Link Prediction

Théo Trouillon, Johannes Welbl, ... (+3 more)

cs.AI 🏛 ICML 📚 3.4K cites 9 years ago

Died the same way — ⚰️ The Empty Tomb

R.I.P. ⚰️ The Empty Tomb

DSFD: Dual Shot Face Detector

Jian Li, Yabiao Wang, ... (+7 more)

cs.CV 🏛 CVPR 📚 462 cites 7 years ago

R.I.P. ⚰️ The Empty Tomb

InstanceCut: from Edges to Instances with MultiCut

Alexander Kirillov, Evgeny Levinkov, ... (+3 more)

cs.CV 🏛 CVPR 📚 261 cites 9 years ago

R.I.P. ⚰️ The Empty Tomb

FLNet: Landmark Driven Fetching and Learning Network for Faithful Talking Facial Animation Synthesis

Kuangxiao Gu, Yuqian Zhou, Thomas Huang

cs.CV 🏛 AAAI 📚 62 cites 6 years ago

R.I.P. ⚰️ The Empty Tomb

Personalized Showcases: Generating Multi-Modal Explanations for Recommendations

An Yan, Zhankui He, ... (+3 more)

cs.IR 🏛 SIGIR 📚 58 cites 3 years ago