TinyBERT: Distilling BERT for Natural Language Understanding
Xiaoqi Jiao low, Yichun Yin low, Lifeng Shang low, Xin Jiang, Xiao Dong Chen, Linlin Li, Fang Wang
Language model pre-training, such as BERT, has significantly improved the performances of many natural language processing tasks. However, pre-trained language models are usually computationally expensive, so it is difficult to efficiently execute them on resourcerestricted devices. To accelerate inference and reduce model size while maintaining accuracy, we first propose a novel Transformer distillation method that is specially designed for knowledge distillation (KD) of the Transformer-based models. By leveraging this new KD method, the plenty of knowledge encoded in a large "teacher" BERT can be effectively transferred to a small "student" Tiny-BERT. Then, we introduce a new two-stage learning framework for TinyBERT, which performs Transformer distillation at both the pretraining and task-specific learning stages. This framework ensures that TinyBERT can capture the general-domain as well as the task-specific knowledge in BERT.
What this paper cites, inside the corpus
What cites it, inside the corpus
Links
Topics
| Topic Modeling | Computer Science |
| Natural Language Processing Techniques | Computer Science |
| Multimodal Machine Learning Applications | Computer Science |
Is this record sound?
complete
Nothing in this record contradicts itself and no field we check is missing.
- supports7 author record(s) attached.
- supports46 reference(s) recorded.
- supportsThe DOI's year agrees with the publication year.
- supportsA title is present.
Provenance
sha256 db1645b78a57e29a…