Vision Transformer
A vision transformer (ViT) processes an image as a sequence of patch representations using a transformer encoder. The original ViT embeds fixed-size patches and adds position information before attention mixes information across the sequence.
Also known as: ViT
Split a 224×224 image into non-overlapping 16×16 patches and there are 14×14, or 196, patches. Each is projected into an embedding. The original classification design adds a class token, giving 197 tokens before the transformer encoder. Position embeddings preserve information about where patches came from.
An image classifier can use the class token’s output to predict a label. Attention lets representations incorporate information from other patches; this differs from a convolutional network’s local filtering structure. Hybrid designs can combine both approaches.
Patch size changes sequence length and therefore the attention workload. A smaller patch can retain finer detail while increasing the number of tokens. ViT’s original results also depend on pretraining scale. Choose using task accuracy, memory and latency under the actual image resolution; the architecture name alone does not promise an improvement.
Sources
- Dosovitskiy et al.: An Image is Worth 16×16 Words — Introduces patch embeddings, position embeddings, the class token and transformer image classification.
Go deeper
- Torchvision: VisionTransformer models docs
Compare available model builders, patch sizes and pretrained weights.