AI Glossary

Vision Transformer

A vision transformer (ViT) processes an image as a sequence of patch representations using a transformer encoder. The original ViT embeds fixed-size patches and adds position information before attention mixes information across the sequence.

Also known as: ViT

· Updated · Chain of Thought

Split a 224×224 image into non-overlapping 16×16 patches and there are 14×14, or 196, patches. Each is projected into an embedding. The original classification design adds a class token, giving 197 tokens before the transformer encoder. Position embeddings preserve information about where patches came from.

An image classifier can use the class token’s output to predict a label. Attention lets representations incorporate information from other patches; this differs from a convolutional network’s local filtering structure. Hybrid designs can combine both approaches.

Patch size changes sequence length and therefore the attention workload. A smaller patch can retain finer detail while increasing the number of tokens. ViT’s original results also depend on pretraining scale. Choose using task accuracy, memory and latency under the actual image resolution; the architecture name alone does not promise an improvement.

Sources

Go deeper