ImagiFilter: A resource to enable the semi-automatic mining of images at scale

August 20, 2020 · Entered Twilight · 🏛 arXiv.org

"Last commit was 5.0 years ago (≥5 year threshold)"

Evidence collected by the PWNC Scanner

Repo contents: .github, README.md, bash_automation_scripts, get_txt_split.py, license.md, load_data.py, models, requirements.txt, train.py, train_finegrained.py, visuals

Authors Houda Alberts, Iacer Calixto arXiv ID 2008.09152 Category cs.CV: Computer Vision Cross-listed cs.CL Citations 2 Venue arXiv.org Repository https://github.com/houda96/imagi-filter ⭐ 3 Last Checked 2 months ago

Abstract

Datasets (semi-)automatically collected from the web can easily scale to millions of entries, but a dataset's usefulness is directly related to how clean and high-quality its examples are. In this paper, we describe and publicly release an image dataset along with pretrained models designed to (semi-)automatically filter out undesirable images from very large image collections, possibly obtained from the web. Our dataset focusses on photographic and/or natural images, a very common use-case in computer vision research. We provide annotations for coarse prediction, i.e. photographic vs. non-photographic, and smaller fine-grained prediction tasks where we further break down the non-photographic class into five classes: maps, drawings, graphs, icons, and sketches. Results on held out validation data show that a model architecture with reduced memory footprint achieves over 96% accuracy on coarse-prediction. Our best model achieves 88% accuracy on the hardest fine-grained classification task available. Dataset and pretrained models are available at: https://github.com/houda96/imagi-filter.