AI Glossary

AI Alignment

AI alignment is the work of making an AI system pursue what its designers and users actually intend — including the goals they didn't think to spell out — rather than optimizing a literal objective in harmful or unintended ways. It spans training techniques, evaluation, and oversight.

Also known as: alignment, aligned AI

· Chain of Thought

AI alignment concerns whether a system’s behavior matches intended goals and constraints. For a language model, distinguish the objective used during training from what happens when the deployed model generates a response. A model does not necessarily keep optimizing its training objective through new weight updates while you chat with it.

The InstructGPT paper describes one approach: supervised demonstrations followed by reinforcement learning from human feedback. The authors report improvements in human preference and truthfulness on their evaluations, while noting that the resulting models still make simple mistakes. Those results support a particular training method, not a guarantee of aligned behavior in every setting.

For a deployed agent, make the intended behavior concrete. Specify which actions it may take, how to evaluate its responses, and when a person must review a decision. Guardrails and evaluation help test and constrain behavior; they do not by themselves solve alignment.

From the conversation