Google’s Gemini was expected to arrive in the fall as a new contender in the race to build capable AI systems. Reports described a multimodal model intended for Google’s own products as well as outside developers, with text, image and coding capabilities among its reported strengths.
A model designed for more than text
Gemini was described as a group of large AI models rather than a single model. That wording leaves room for different configurations: the models might be organized around specialized capabilities, or offered in different sizes. The report did not establish which interpretation Google would use.
The distinction matters because Gemini was expected to compete with OpenAI’s GPT-4. A family of models could give Google options for serving different tasks and users, although the details of how the group would work were not made public in the report.
Gemini was reportedly able to generate images as well as text, and was said to have significantly improved coding capabilities. Its training included YouTube video transcripts, prompting speculation that it might also generate simple videos. That possibility was described as potential, rather than a confirmed feature.
Google planned a gradual rollout
Google planned to integrate Gemini into its own services over time, including the Bard chatbot and Google Docs or Slides. Later in the year, it was also expected to become available to external developers through Google Cloud.
This rollout would put the model in two settings: consumer-facing Google products and tools developers could use to build their own AI applications. The report did not specify an exact launch date or explain which features would arrive first.
Making the model available through several products could also make its capabilities more visible to users. For developers, access through Google Cloud would create a route to experiment with Gemini beyond Google’s own apps. Both plans depended on the model being ready for those uses.
A large team and a major research effort
The Information reported that at least two dozen executives were involved in Gemini’s development and that the team included several hundred employees from Google Brain and Deepmind. Deepmind founder Demis Hassabis led the effort, with support from Oriol Vinyals, Koray Kavukcuoglu and former Google Brain chief Jeff Dean. Google founder Sergey Brin was also reportedly helping train and evaluate the model.
The report placed this work amid an organizational transition: Google Deepmind had recently been merged and was still working through questions such as remote work policies and the technology used to train models. It also said Deepmind had set aside a ChatGPT competitor codenamed “Goodall,” based on an unannounced model called “Chipmunk,” in favor of Gemini.
The article said earlier rumors put Gemini at at least a trillion parameters and described training with tens of thousands of Google’s TPU AI chips. These were reported claims, not specifications confirmed in the article as official details.
Training data and responsible development
Gemini’s training materials were reportedly under close review by Google’s legal department. The development team had to remove training data from copyrighted books. The Information’s source also said Gemini had inadvertently been trained on “offensive” content, which likely led to a partial retraining.
These details show that building a model involves choices about the material used to train it, as well as work to address problems discovered along the way. The report did not describe the scope of the data removal or explain what retraining changed, so those points remain unclear.
Gemini had been officially unveiled in May. In late June, Demis Hassabis said it would combine strengths of AlphaGo-type systems with the language capabilities of large models, alongside new innovations. The fall launch plans and reported capabilities therefore pointed to an ambitious project, but the article left key questions about its final design and features open.