Large language model
- Sprouting
- c.
A large language model (LLM) is a form of neural network artificial intelligence (AI) designed to understand and process human text. The fundamental goal of a LLM is to statistically predict what will be the next word in a sequence.
LLMs are trained on massive datasets, analysing trillions of books, codebases, websites, and other text-based resources. Modern LLMs use transformer architecture (self-attention) to evaluate how each word relates to one another.
Concepts
- Tokens: subsets of text such as whole words, part of a word (“ing”), or punctation like a question mark.
- Parameters: adjustable, internal mathematical values fine-tuned during training that encode knowledge about how concepts and words relate to one another.
- Context window: the amount of text (measured in tokens) the model can process and retain in memory at a single time during a conversation.