Which phoneme-to-viseme maps best improve visual-only computer lip-reading?

October 03, 2017 Β· Declared Dead Β· πŸ› International Symposium on Visual Computing

πŸ‘» CAUSE OF DEATH: Ghosted
No code link whatsoever

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Helen L. Bear, Richard W. Harvey, Barry-John Theobald, Yuxuan Lan arXiv ID 1710.01093 Category cs.CV: Computer Vision Cross-listed cs.CL, eess.AS Citations 34 Venue International Symposium on Visual Computing Last Checked 6 months ago
Abstract
A critical assumption of all current visual speech recognition systems is that there are visual speech units called visemes which can be mapped to units of acoustic speech, the phonemes. Despite there being a number of published maps it is infrequent to see the effectiveness of these tested, particularly on visual-only lip-reading (many works use audio-visual speech). Here we examine 120 mappings and consider if any are stable across talkers. We show a method for devising maps based on phoneme confusions from an automated lip-reading system, and we present new mappings that show improvements for individual talkers.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

πŸ“œ Similar Papers

In the same crypt β€” Computer Vision

πŸŒ… πŸŒ… Old Age

Fast R-CNN

Ross Girshick

cs.CV πŸ› ICCV πŸ“š 27.7K cites 11 years ago

Died the same way β€” πŸ‘» Ghosted