Graphics Processing Unit
A graphics processing unit (GPU) is a computer chip designed to perform many calculations in parallel. Originally developed for graphics, GPUs also accelerate the calculations used to train and run AI models.
Also known as: GPU, GPUs
A GPU divides suitable work across many calculations running at once. A central processing unit (CPU), the general-purpose processor in a computer, is designed to handle sequential work well. They often work together; adding a GPU does not make every part of a program faster.
Example: an operations team could use GPU-backed computing to run a model over a large batch of scanned invoices. The GPU accelerates the model’s calculations; the application still needs to load the files, check extracted values, and send exceptions for review.
For a buying decision, ask for measurements on your actual workload: how long a batch takes, how many requests it handles, and what the completed work costs. A chip’s specifications alone do not answer those questions. Renting capacity from a cloud provider and buying hardware are different ways to obtain the same kind of computing resource.
See AI vocabulary in plain English for the relationship between GPUs, models, and hyperscalers.
Sources
- NVIDIA CUDA Programming Guide: Introduction — Explains GPUs' graphics origins and parallel design, contrasts them with CPUs, and describes their use in AI. Checked October 6, 2026.
- NVIDIA: Why GPUs Are Great for AI — Describes parallel calculations in AI training and inference. Used for this mechanism, not its performance or market claims. Checked October 6, 2026.