What 22 Billion Parameters Could Change in Image AI

Google’s ViT-22B is a 22 billion-parameter vision model trained on four billion images. The paper reports strong benchmark results and a shape bias closer to human classification, while suggesting that larger vision models could develop new capabilities.

WTF Index TERMINATOR
◄ Terminator 2 Idiocracy 1 ►

ViT-22B’s scale and suggestion of emerging capabilities lean mildly toward more powerful AI, while the article reports no clear harm or social decline.

What 22 Billion Parameters Could Change in Image AI

Google’s ViT-22B expands the Vision Transformer into a 22 billion-parameter model trained on four billion images. The research examines how it performs across image and video tasks, whether it can teach smaller models, and how its object recognition compares with human perception.

From language transformers to image tasks

Google introduced the Vision Transformer, or ViT, in the fall of 2020. The architecture adapts Transformer models, influential in language processing, for image work such as recognizing objects. Instead of reading words, it processes small portions of an image.

Google’s first models came in three sizes: ViT-Base with 86 million parameters, ViT-Large with 307 million, and ViT-Huge with 632 million. They were trained on 300 million images. In June 2021, another model, ViT-G/14, broke the previous ImageNet benchmark record. It had just under two billion parameters and was trained on three billion images.

ViT-22B marks a much larger step. The new model has 22 billion parameters, ten times the size of ViT-G/14, and was trained using four billion images on 1,024 TPU-v4 chips. That scale required changes to the training approach: the team addressed stability issues, including by arranging transformer layers in parallel, which also made more efficient use of the hardware.

One model, several kinds of vision work

The team evaluated ViT-22B across image classification, semantic segmentation, depth estimation, and video classification. These tasks ask a model to do different things: identify what an image shows, distinguish regions within it, estimate depth, or classify video content.

In some benchmarks, the model achieved state-of-the-art results; in others, it ranked among the top performers. The report emphasizes that these results came without specializing the model for each task. Google also tested its classification ability on AI-generated images that were not included in its training data.

That breadth matters because a general model may be useful across more than one narrowly defined problem. Still, benchmark performance describes how a model does on particular evaluations. The source does not claim that strong scores alone settle how well it will work in every practical setting.

A large model can teach a smaller one

ViT-22B may also help improve smaller vision models. In a teacher-student setup, a ViT-Base model learned from ViT-22B and then scored 88.6 percent on the ImageNet benchmark. The paper describes this as a new high state-of-the-art result for a model of that size.

This finding points to a role for very large models beyond their own direct use: their learned representations can be passed on to smaller systems. The result does not mean that every smaller model will gain the same way, but it shows how a large teacher can transfer useful knowledge in at least this evaluated setup.

Recognizing shape, not only texture

The study also looks at a known weakness in image recognition: AI models can rely too heavily on texture when classifying objects, while people tend to focus on shape. According to tests cited in the paper, human attention in object classification is 96 percent shape and 4 percent texture.

Google reports that ViT-22B has an 87 percent shape bias and a 13 percent texture bias, a new state-of-the-art result on this measure. That makes its balance closer to the human figures reported, though it does not make the model identical to human perception. The paper also describes it as more robust and fairer.

Google argues that ViT-22B demonstrates the potential for image processing to benefit from scaling in a way comparable to large language models. The team suggests that further scaling could bring emergent capabilities that address some current limitations of Vision Transformers. That remains a possibility raised by the research, rather than a result already demonstrated by this model.