Optimizing seed inputs in fuzzing with machine learning

February 07, 2019 · Declared Dead · 🏛 2019 IEEE/ACM 41st International Conference on Software Engineering: Companion Proceedings (ICSE-Companion)

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Liang Cheng, Yang Zhang, Yi Zhang, Chen Wu, Zhangtan Li, Yu Fu, Haisheng Li arXiv ID 1902.02538 Category cs.CR: Cryptography & Security Citations 30 Venue 2019 IEEE/ACM 41st International Conference on Software Engineering: Companion Proceedings (ICSE-Companion) Last Checked 3 months ago

Abstract

The success of a fuzzing campaign is heavily depending on the quality of seed inputs used for test generation. It is however challenging to compose a corpus of seed inputs that enable high code and behavior coverage of the target program, especially when the target program requires complex input formats such as PDF files. We present a machine learning based framework to improve the quality of seed inputs for fuzzing programs that take PDF files as input. Given an initial set of seed PDF files, our framework utilizes a set of neural networks to 1) discover the correlation between these PDF files and the execution in the target program, and 2) leverage such correlation to generate new seed files that more likely explore new paths in the target program. Our experiments on a set of widely used PDF viewers demonstrate that the improved seed inputs produced by our framework could significantly increase the code coverage of the target program and the likelihood of detecting program crashes.

📄 View on arXiv 🌐 View on ar5iv 📑 PDF 🎉 Report Code Found

Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

📜 Similar Papers

In the same crypt — Cryptography & Security

R.I.P. 👻 Ghosted

Towards Evaluating the Robustness of Neural Networks

Nicholas Carlini, David Wagner

cs.CR 🏛 IEEE S&P 📚 9.5K cites 9 years ago

R.I.P. 👻 Ghosted

Membership Inference Attacks against Machine Learning Models

Reza Shokri, Marco Stronati, ... (+2 more)

cs.CR 🏛 IEEE S&P 📚 4.9K cites 9 years ago

R.I.P. 👻 Ghosted

The Limitations of Deep Learning in Adversarial Settings

Nicolas Papernot, Patrick McDaniel, ... (+4 more)

cs.CR 🏛 IEEE S&P 📚 4.2K cites 10 years ago

R.I.P. 👻 Ghosted

Practical Black-Box Attacks against Machine Learning

Nicolas Papernot, Patrick McDaniel, ... (+4 more)

cs.CR 🏛 ASIACCS 📚 3.9K cites 10 years ago

R.I.P. 👻 Ghosted

Distillation as a Defense to Adversarial Perturbations against Deep Neural Networks

Nicolas Papernot, Patrick McDaniel, ... (+3 more)

cs.CR 🏛 IEEE S&P 📚 3.2K cites 10 years ago

R.I.P. 👻 Ghosted

Extracting Training Data from Large Language Models

Nicholas Carlini, Florian Tramer, ... (+10 more)

cs.CR 🏛 USENIX Sec 📚 2.6K cites 5 years ago

Died the same way — 👻 Ghosted

R.I.P. 👻 Ghosted

Language Models are Few-Shot Learners

Tom B. Brown, Benjamin Mann, ... (+29 more)

cs.CL 🏛 NeurIPS 📚 54.2K cites 6 years ago

R.I.P. 👻 Ghosted

PyTorch: An Imperative Style, High-Performance Deep Learning Library

Adam Paszke, Sam Gross, ... (+19 more)

cs.LG 🏛 NeurIPS 📚 49.7K cites 6 years ago

R.I.P. 👻 Ghosted

XGBoost: A Scalable Tree Boosting System

Tianqi Chen, Carlos Guestrin

cs.LG 🏛 KDD 📚 49.2K cites 10 years ago

R.I.P. 👻 Ghosted

Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift

Sergey Ioffe, Christian Szegedy

cs.LG 🏛 ICML 📚 46.0K cites 11 years ago