Glossary
AGI and superintelligence, term by term.
Short definitions of the words that come up in every discussion of advanced AI.
- Agent
- An AI system that pursues a goal over many steps by planning, using tools and acting on the results, with little supervision between steps.
- AGI
- Artificial general intelligence. A system able to do most intellectual work at the level of a capable person, across domains and without task-specific redesign.
- Alignment
- The work of making an AI system pursue the goals its builders and users actually intend, including in situations they did not foresee.
- Benchmark
- A fixed set of tasks with known answers used to compare systems. Useful until scores reach the ceiling, when it is said to be saturated.
- Compute
- The processing power used to train and run models, usually counted in specialised chips and the energy they draw.
- Context window
- The amount of text a model can take into account at once, measured in tokens.
- Corrigibility
- The property of accepting correction or shutdown from authorised people without resisting or working around it.
- Emergent behaviour
- An ability that appears in a trained model without having been designed or targeted directly.
- Evals
- Evaluations. Structured tests of what a model can do and how it behaves, run before and after release.
- Fine-tuning
- Further training of an existing model on a smaller, specific dataset to adapt it to a task or style.
- Foundation model
- A large model trained on broad data that serves as the base for many different applications.
- Goodhart's law
- When a measure becomes a target, it stops being a good measure. The standard warning about optimising benchmarks.
- Hallucination
- Output that is fluent and confident but false or unsupported by any source.
- Inference
- Running a trained model to produce an output, as opposed to training it.
- Instrumental convergence
- The observation that many different goals share useful sub-goals, such as gaining resources and avoiding shutdown.
- Intelligence explosion
- I. J. Good's 1965 idea that a machine able to design better machines would trigger rapid, self-reinforcing progress.
- Interpretability
- Research that studies a model's internal computations to explain why it produces the outputs it does.
- Orthogonality thesis
- The claim that how intelligent a system is and what goals it has are independent of each other.
- Recursive self-improvement
- A loop in which an AI system improves the process that builds AI systems, so each generation speeds up the next.
- Reinforcement learning
- Training by trial and error, where a system learns from rewards for the outcomes of its actions.
- Singularity
- A point after which technological change is too fast for people to forecast. Popularised by Vernor Vinge in 1993.
- Superintelligence
- A system whose cognitive performance greatly exceeds the best humans in virtually every domain. Also written ASI.
- Takeoff
- The period between human-level AI and superintelligence. A slow takeoff lasts years; a fast one lasts months or less.
- Token
- The unit of text a language model reads and writes: a word, part of a word, or a punctuation mark.
- Transformer
- The neural network architecture, introduced in 2017, behind today's large language models. It relates every part of the input to every other part.
- Weights
- The numbers inside a neural network that are adjusted during training and that determine its behaviour.