CANONICAL HISTORY
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy and coauthors submitted 'An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale' on October 22, 2020. The paper describes applying a pure Transformer directly to sequences of image patches and names the resulting model Vision Transformer (ViT).
Evidence / resource
This page preserves the public LINEAiGE record and its first-party source relationship.
Record identity
LINEAiGE IDvision-transformer-2020