Real-Time Lip Sync for Live 2D Animation

October 19, 2019 · Entered Twilight · 🏛 arXiv.org

"Last commit was 6.0 years ago (≥5 year threshold)"

Evidence collected by the PWNC Scanner

Repo contents: 1_ChOnline_vs_Ours, 2_ChOffline_vs_Ours, 3_ToonBoom_vs_Ours, 4_OursNoAug_vs_Ours, 5_Ours4_vs_Ours, README.md, character_lipsync_video_summary.mp4, teaser.png

Authors Deepali Aneja, Wilmot Li arXiv ID 1910.08685 Category cs.GR: Graphics Cross-listed cs.CV, cs.HC, cs.LG Citations 16 Venue arXiv.org Repository https://github.com/deepalianeja/CharacterLipSync2D ⭐ 149 Last Checked 2 months ago

Abstract

The emergence of commercial tools for real-time performance-based 2D animation has enabled 2D characters to appear on live broadcasts and streaming platforms. A key requirement for live animation is fast and accurate lip sync that allows characters to respond naturally to other actors or the audience through the voice of a human performer. In this work, we present a deep learning based interactive system that automatically generates live lip sync for layered 2D characters using a Long Short Term Memory (LSTM) model. Our system takes streaming audio as input and produces viseme sequences with less than 200ms of latency (including processing time). Our contributions include specific design decisions for our feature definition and LSTM configuration that provide a small but useful amount of lookahead to produce accurate lip sync. We also describe a data augmentation procedure that allows us to achieve good results with a very small amount of hand-animated training data (13-20 minutes). Extensive human judgement experiments show that our results are preferred over several competing methods, including those that only support offline (non-live) processing. Video summary and supplementary results at GitHub link: https://github.com/deepalianeja/CharacterLipSync2D