Skip to content

AI Glossary

Convolutional neural network

A convolutional neural network (CNN) is a neural network that uses learned filters shared across positions to process data such as images. Each convolution produces feature maps that later layers use for tasks such as image classification.

Also known as: CNN

· Updated · Chain of Thought

Model Architecture

An image enters the network as a tensor. One 32-by-32 color image can have shape [1, 3, 32, 32]: batch, red-green-blue channels, height and width. A 5-by-5 convolution with six output channels learns six filters. Each filter combines values from a 5-by-5 neighborhood across the input channels and reuses its weights at different image positions, producing a feature map: an array of that filter’s responses.

The PyTorch tutorial supplies the layer configuration Conv2d(3, 6, 5), with default stride 1 and no padding. Applying it to our illustrative single-image batch produces [1, 6, 28, 28]: six feature maps, each 28-by-28. The tutorial then uses a nonlinear activation and pooling; pooling reduces the spatial dimensions while retaining six channels for the next convolution. Later layers predict the image’s class. These channel counts and sizes are one configuration.

The tensor entry explains how those dimensions relate to data types and memory. Transformer explains attention and shows why a text-generation decoder masks future tokens.

Sources

  • PyTorch: Conv2d — Defines shared learnable weights, input and output channels, and spatial dimensions for a two-dimensional convolution.
  • PyTorch: Training a Classifier — Demonstrates three input and six output channels, nonlinear activations, pooling and a classification head in a runnable color-image example.

From the conversation