A Dataset for the Recognition of Historical and Handwritten Music Scores in Western Notation

May 18, 2026 Β· Grace Period Β· + Add venue

⏳ Grace Period
This paper is less than 90 days old. We give authors time to release their code before passing judgment.
Authors Pau Torras, JiΕ™Γ­ Mayer, Carles Badal, Martina DvoΕ™Γ‘kovΓ‘, MarkΓ©ta HerzanovΓ‘ VlkovΓ‘, Gerard Asbert, VojtΔ›ch DvoΕ™Γ‘k, Samuel Ε omorjai, Jan Hajič, Alicia FornΓ©s arXiv ID 2605.18436 Category cs.CV: Computer Vision Citations 0
Abstract
A large amount of musical heritage has been digitised by memory institutions: libraries, museums, and archives. Nevertheless, the field of Optical Music Recognition (OMR) has struggled with making this music machine-readable, despite advances in deep learning, mostly because no datasets for training systems in realistic conditions were available. The MusiCorpus dataset aims to remedy this situation by providing 1,309 pages of historical sheet music, primarily handwritten, with MusicXML transcriptions and symbol annotations. It is the largest dataset of handwritten music to date and the first dataset containing a realistic and representative sample of musical document collections from memory institutions, suitable for training and evaluating both end-to-end and object detection-based OMR systems and comparing their performance.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

πŸ“œ Similar Papers

In the same crypt β€” Computer Vision

πŸŒ… πŸŒ… Old Age

Fast R-CNN

Ross Girshick

cs.CV πŸ› ICCV πŸ“š 27.7K cites 11 years ago