Who Cited It

Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

2019 · arXiv (Cornell University) · 3,698 citations · 9 from inside this corpus

Colin Raffel, Noam Shazeer, Adam P. Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J. Liu low

The source holds an abstract for this work, but its best open-access copy is under no open licence, which does not permit us to republish the text. Read it at the source below.

Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (2019)Exploring the Limits of Trans…Dropout: a simple way to prevent neural networks from overfitting (2014)Dropout: a simple way to prev…Glove: Global Vectors for Word Representation (2014)Glove: Global Vectors for Wor…BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (2019)BERT: Pre-training of Deep Bi…BLEU (2001)BLEUDistributed Representations of Words and Phrases and their Compositionality (2013)Distributed Representations o…HISTORIAE, History of Socio-Cultural Transformation as Linguistic Data Science. A Humanit… (2019)HISTORIAE, History of Socio-C…Distilling the Knowledge in a Neural Network (2015)Distilling the Knowledge in a…Sequence to Sequence Learning with Neural Networks (2014)Sequence to Sequence Learning…Efficient Estimation of Word Representations in Vector Space (2013)Efficient Estimation of Word …ROUGE: A Package for Automatic Evaluation of Summaries (2004)ROUGE: A Package for Automati…[No title in the source record — Edinburgh Research Explorer (University of Edinburgh)][No title in the source recor…Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank (2013)Recursive Deep Models for Sem…Multitask Learning (1997)Multitask LearningSQuAD: 100,000+ Questions for Machine Comprehension of Text (2016)SQuAD: 100,000+ Questions for…Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Tr… (2016)Google's Neural Machine Trans…DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter (2019)DistilBERT, a distilled versi…A Learning Algorithm for Continually Running Fully Recurrent Neural Networks (1989)A Learning Algorithm for Cont…Get To The Point: Summarization with Pointer-Generator Networks (2017)Get To The Point: Summarizati…A Convolutional Neural Network for Modelling Sentences (2014)A Convolutional Neural Networ…Learning and Transferring Mid-level Image Representations Using Convolutional Neural Netw… (2014)Learning and Transferring Mid…SciBERT: A Pretrained Language Model for Scientific Text (2019)SciBERT: A Pretrained Languag…Federated Learning: Strategies for Improving Communication Efficiency (2016)Federated Learning: Strategie…An Overview of Multi-Task Learning in Deep Neural Networks (2017)An Overview of Multi-Task Lea…XLNet: Generalized Autoregressive Pretraining for Language Understanding (2019)XLNet: Generalized Autoregres…Deep Contextualized Word Representations (2018)Deep Contextualized Word Repr…Cross-lingual Language Model Pretraining (2019)Cross-lingual Language Model …Teaching Machines to Read and Comprehend (2015)Teaching Machines to Read and…On the Dangers of Stochastic Parrots (2021)On the Dangers of Stochastic …Survey of Hallucination in Natural Language Generation (2022)Survey of Hallucination in Na…Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Langu… (2022)Pre-train, Prompt, and Predic…Large language models encode clinical knowledge (2023)Large language models encode …HuggingFace's Transformers: State-of-the-art Natural Language Processing (2019)HuggingFace's Transformers: S…Affordance-Compiled Intelligence: Observable-Only Cognitive Impedance Matching for No-Met… (2026)Affordance-Compiled Intellige…FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence (2020)FixMatch: Simplifying Semi-Su…Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing (2021)Domain-Specific Language Mode…Pre-trained models for natural language processing: A survey (2020)
36 of 36 neighbouring works in this corpus. Blue is what this paper cites; orange is what cites it, and a dashed line is one neighbour citing another. Only the largest labels are drawn — every node carries its full title on hover.
this paper works it cites works citing it node size = global citations · hover for the full title

What this paper cites, inside the corpus

PaperYearCited
Dropout: a simple way to prevent neural networks from overfitting201434,236
Glove: Global Vectors for Word Representation201434,067
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding201933,416
BLEU200121,963
Distributed Representations of Words and Phrases and their Compositionality201318,054
HISTORIAE, History of Socio-Cultural Transformation as Linguistic Data Science. A Humanit…201917,489
Distilling the Knowledge in a Neural Network201514,099
Sequence to Sequence Learning with Neural Networks201413,351
Efficient Estimation of Word Representations in Vector Space201311,714
ROUGE: A Package for Automatic Evaluation of Summaries20048,302
[No title in the source record — Edinburgh Research Explorer (University of Edinburgh)]7,291
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank20136,849
Multitask Learning19976,478
SQuAD: 100,000+ Questions for Machine Comprehension of Text20166,435
Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Tr…20165,668
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter20194,600
A Learning Algorithm for Continually Running Fully Recurrent Neural Networks19894,490
Get To The Point: Summarization with Pointer-Generator Networks20173,935
A Convolutional Neural Network for Modelling Sentences20143,571
Learning and Transferring Mid-level Image Representations Using Convolutional Neural Netw…20143,197
SciBERT: A Pretrained Language Model for Scientific Text20193,109
Federated Learning: Strategies for Improving Communication Efficiency20163,050
An Overview of Multi-Task Learning in Deep Neural Networks20172,431
XLNet: Generalized Autoregressive Pretraining for Language Understanding20191,854
Deep Contextualized Word Representations20181,783
Cross-lingual Language Model Pretraining20191,624
Teaching Machines to Read and Comprehend20151,519

What cites it, inside the corpus

Topics

Topic ModelingComputer Science
Natural Language Processing TechniquesComputer Science
Multimodal Machine Learning ApplicationsComputer Science

Is this record sound?

partial

One field of this record is missing or disagrees with another. What is shown below is what the source publishes.

  • supports9 author record(s) attached.
  • supports125 reference(s) recorded.
  • weakensThe DOI names 1910 but the record dates this to 2,019. One of the two is about a different paper.
  • supportsA title is present.

Provenance

Everything above was read from one stored OpenAlex payload, fetched 2026-09-04T03:58:46+00:00.

sha256 bba2969b3567609a…