AI vocabulary in plain English: 14 terms
You do not need to be an engineer to take part in an AI decision. You do need a shared vocabulary for what a system does, what information it uses, and who stays in control.
This guide explains 14 terms business and industry leaders keep hearing. Start with generative AI, then work through models, company data, agents, and infrastructure. The business examples are illustrative.
Read straight through or jump to a group below. Each term includes a question or implication to take into your next conversation, with glossary definitions, deeper explainers, and selected guest discussions to follow.
Start with the umbrella: GenAI
GenAI: generative artificial intelligence
Generative AI produces content, such as text, images, audio, or code, from patterns learned during training. For example, an operations manager could use it to draft a supplier email from delivery notes, then check the dates and commitments before sending it.
Why it matters to you: Separate creating a useful draft from confirming that its contents are correct.
How the models work
LLMs: large language models
A large language model is a type of AI trained on large amounts of text that generates language by predicting the next token, or piece of text. It can help a finance team turn a long report into a short briefing, but a fluent summary can still contain mistakes.
Why it matters to you: Ask how important claims will be checked, even when the writing sounds confident.
Tokens
Tokens are the units of text a model processes: a whole word, part of a word, or punctuation. When a team asks an assistant to summarize a long incident report, both the supplied text and generated response use tokens; token counts affect request limits and, for token-priced services, cost.
Why it matters to you: Budget for the documents going in and the work coming out, not just the length of the answer you see.
Context
Context is the information available to a model for the task at hand, including instructions, conversation history, documents, and tool results. For an assistant answering a clinical-trial operations question, that might mean the relevant protocol section and an approved site procedure. The context window is the limited space available, not permanent memory of every document your organization holds.
Why it matters to you: Ask which records the assistant actually received, whether they are relevant, and whether they are up to date.
Prompts
A prompt is the input you give a model to guide its response, often combining a request, instructions, and examples. A procurement manager might ask: “Compare these three proposals on price, delivery dates, and exclusions; flag any missing information.” That is more specific than simply asking which proposal is best.
Why it matters to you: State the task, evidence, and output you need; wording alone cannot supply facts the model has not been given.
Hybrid reasoning
In model descriptions, hybrid reasoning can mean one model that supports both quick responses and extended reasoning before it answers. For example, a finance team could use a quick mode to rephrase an email and a reasoning mode to work through conflicting assumptions in a budget plan. Extra reasoning takes time and computing resources; it does not guarantee a correct result.
Why it matters to you: Ask what the vendor means by “hybrid,” then compare speed, cost, and accuracy on your own tasks.
Building on your own data
RAG: retrieval-augmented generation
RAG means finding relevant information and supplying it to a model when it writes an answer. A support assistant could retrieve the applicable warranty policy before explaining whether a repair is covered. This can bring company information into an answer without retraining the model, but the search or the answer can still be wrong.
Why it matters to you: Check both the retrieved evidence and whether the answer follows from it.
Vector databases
A vector database stores numerical representations of data, called embeddings, and searches for similar ones. That can help an operations team find maintenance notes about an overheating motor even when the question uses different wording. Similarity is a search signal, not proof that a result is the right document.
Why it matters to you: Ask whether search also respects access permissions, dates, and the equipment or customer the question concerns.
Agents and oversight
Agents
In language-model applications, an agent uses a model to choose actions toward a goal, use tools, and decide what to do next from the results. For example, an operations agent could investigate a late shipment by checking an order, reading a carrier update, and drafting a response. The application controls which tools and permissions it has.
Why it matters to you: Evaluate the steps and actions, as well as the final answer.
Autonomous
Autonomous describes how much work a system can do without a person directing or approving each step. An invoice assistant might collect records and prepare a reconciliation on its own while still needing approval to release a payment. Autonomy has a scope; it does not mean unlimited authority.
Why it matters to you: Define exactly which actions can happen without approval, and how a person can stop or correct them.
HITL: human in the loop
Human in the loop means a person takes part at a defined point in an AI process, such as reviewing an output or approving an action. For example, a trial coordinator could review an AI-drafted site communication before it is sent. A person who only monitors completed actions has a different role from someone whose approval is required first.
Why it matters to you: Specify who reviews what, when they can intervene, and what evidence they see.
MCP: Model Context Protocol
MCP is an open protocol: a shared way for AI applications to connect to tools and data. For example, an internal assistant could use an MCP connection to search an approved document library. The connection still needs compatible software and access controls; the protocol itself does not grant permission to read every file.
Why it matters to you: Ask which systems are connected and what each connection allows the assistant to do.
The hardware and the cloud
GPUs: graphics processing units
A GPU is a chip designed to perform many calculations in parallel, which helps with training and running AI models. For example, an operations team could rent GPU-backed computing to process a large batch of scanned invoices with an AI model.
Why it matters to you: Ask what the proposed hardware delivers on your workload and what each completed task costs.
Hyperscalers
In cloud discussions, hyperscalers are providers of computing, storage, and other services at very large scale, such as Amazon Web Services, Microsoft Azure, and Google Cloud. A manufacturer could rent capacity from one to run an internal maintenance assistant instead of building its own server facility.
Why it matters to you: Identify the provider behind the service, where your data goes, and which responsibilities remain with your team.
Go deeper
Bring one real task to your next AI conversation. Ask what information the system needs, what it may do, and how you will check the result. Use these resources to work through the details.
- The business leaders learning track — take this vocabulary into decisions about cost, value, oversight and agent autonomy.
- The AI glossary — look up a term as you encounter it.
- AI, decoded — work through a practical question in more depth.
- Start here — choose your first Chain of Thought episodes.
- The newsletter — follow the conversations in writing.