Large language model

  • Sprouting
  • c.

A large language model (LLM) is a form of neural network artificial intelligence (AI) designed to understand and process human text. The fundamental goal of a LLM is to statistically predict what will be the next word in a sequence.

LLMs are trained on massive datasets, analysing trillions of books, codebases, websites, and other text-based resources. Modern LLMs use transformer architecture (self-attention) to evaluate how each word relates to one another.

Concepts

  • Tokens: subsets of text such as whole words, part of a word (“ing”), or punctation like a question mark.
  • Parameters: adjustable, internal mathematical values fine-tuned during training that encode knowledge about how concepts and words relate to one another.
  • Context window: the amount of text (measured in tokens) the model can process and retain in memory at a single time during a conversation.

Backlinks