Who Cited It

ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning

2021 · IEEE Transactions on Pattern Analysis and Machine Intelligence · 2,446 citations · 0 from inside this corpus

Ahmed Elnaggar, Michael Heinzinger, Christian Dallago, Ghalia Rehawi, Yu Wang, Llion Jones, Tom Gibbs, T. Fehér, Christoph Angerer low, Martin Steinegger, Debsindhu Bhowmik, Burkhard Rost

Computational biology and bioinformatics provide vast data gold-mines from protein sequences, ideal for Language Models (LMs) taken from Natural Language Processing (NLP). These LMs reach for new prediction frontiers at low inference costs. Here, we trained two auto-regressive models (Transformer-XL, XLNet) and four auto-encoder models (BERT, Albert, Electra, T5) on data from UniRef and BFD containing up to 393 billion amino acids. The protein LMs (pLMs) were trained on the Summit supercomputer using 5616 GPUs and TPU Pod up-to 1024 cores. Dimensionality reduction revealed that the raw pLM-embeddings from unlabeled data captured some biophysical features of protein sequences. We validated the advantage of using the embeddings as exclusive input for several subsequent tasks: (1) a per-residue (per-token) prediction of protein secondary structure (3-state accuracy Q3=81%-87%); (2) per-protein (pooling) predictions of protein sub-cellular location (ten-state accuracy: Q10=81%) and membrane versus water-soluble (2-state accuracy Q2=91%). For secondary structure, the most informative embeddings (ProtT5) for the first time outperformed the state-of-the-art without multiple sequence alignments (MSAs) or evolutionary information thereby bypassing expensive database searches. Taken together, the results implied that pLMs learned some of the grammar of the language of life. All our models are available through https://github.com/agemagician/ProtTrans.

8 of 8 neighbouring works in this corpus. Blue is what this paper cites; orange is what cites it, and a dashed line is one neighbour citing another. Only the largest labels are drawn — every node carries its full title on hover.
this paper works it cites works citing it node size = global citations · hover for the full title

What this paper cites, inside the corpus

Topics

Topic ModelingComputer Science
Natural Language Processing TechniquesComputer Science
Innovative Teaching and Learning MethodsPsychology

Is this record sound?

complete

Nothing in this record contradicts itself and no field we check is missing.

  • supports12 author record(s) attached.
  • supports123 reference(s) recorded.
  • supportsThe DOI's year agrees with the publication year.
  • supportsA title is present.

Provenance

Everything above was read from one stored OpenAlex payload, fetched 2026-09-04T03:58:49+00:00.

sha256 a2172c1bbe00c44b…