In a nutshell, ViTs cut the image into patches and perform stacks of self-attention on all the patches, leading to the final feature map, also of lower resolution
https://lucasb.eyer.be/articles/vit_cnn_speed.html
https://blog.mdturp.ch/posts/2024-04-05-visual_guide_to_vision_transformer.html