LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard low, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix low, Baptiste Rozière, Naman Goyal, Eric Hambro low, Faisal Azhar low, Aurelien Rodriguez, Armand Joulin, Édouard Grave, Guillaume Lample
We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters. We train our models on trillions of tokens, and show that it is possible to train state-of-the-art models using publicly available datasets exclusively, without resorting to proprietary and inaccessible datasets. In particular, LLaMA-13B outperforms GPT-3 (175B) on most benchmarks, and LLaMA-65B is competitive with the best models, Chinchilla-70B and PaLM-540B. We release all our models to the research community.
What cites it, inside the corpus
Links
Topics
| Natural Language Processing Techniques | Computer Science |
| Topic Modeling | Computer Science |
| Speech Recognition and Synthesis | Computer Science |
Is this record sound?
partial
One field of this record is missing or disagrees with another. What is shown below is what the source publishes.
- supports14 author record(s) attached.
- weakensNo references are recorded despite 3,966 citations. A paper this heavily cited did not cite nothing, so the record is incomplete.
- neutralThe DOI carries no year to check against.
- supportsA title is present.
Provenance
sha256 bba2969b3567609a…