ExtremeBERT: A Toolkit for Accelerating Pretraining of Customized BERT
November 30, 2022 ยท Entered Twilight ยท ๐ arXiv.org
Repo contents: .gitignore, CODE_OF_CONDUCT.md, LICENSE, NOTICE, README.md, assets, collect_best_val.sh, conf, configs, dataset, docs, finetune_search.sh, install.sh, main.py, pretrain_search.sh, pretraining, requirements.txt, run_glue.py, run_glue.sh, run_pretraining.py, run_pretraining.sh, summarize_val.sh, translate_test_result.sh, vocab
Authors
Rui Pan, Shizhe Diao, Jianlin Chen, Tong Zhang
arXiv ID
2211.17201
Category
cs.CL: Computation & Language
Cross-listed
cs.LG,
math.OC
Citations
9
Venue
arXiv.org
Repository
https://github.com/extreme-bert/extreme-bert
โญ 269
Last Checked
6 months ago
Abstract
In this paper, we present ExtremeBERT, a toolkit for accelerating and customizing BERT pretraining. Our goal is to provide an easy-to-use BERT pretraining toolkit for the research community and industry. Thus, the pretraining of popular language models on customized datasets is affordable with limited resources. Experiments show that, to achieve the same or better GLUE scores, the time cost of our toolkit is over $6\times$ times less for BERT Base and $9\times$ times less for BERT Large when compared with the original BERT paper. The documentation and code are released at https://github.com/extreme-bert/extreme-bert under the Apache-2.0 license.
Community Contributions
Found the code? Know the venue? Think something is wrong? Let us know!
๐ Similar Papers
In the same crypt โ Computation & Language
๐
๐
Old Age
๐
๐
Old Age
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
๐
๐
Old Age
XLNet: Generalized Autoregressive Pretraining for Language Understanding
๐ฎ
๐ฎ
The Ethereal
Effective Approaches to Attention-based Neural Machine Translation
๐
๐
Old Age
A large annotated corpus for learning natural language inference
๐
๐
Old Age