How ChatGPT broadened the role of generative AI

ChatGPT helped make large language models easier to direct through conversation, while later models can also process information across a wider range of tasks. Their flexibility comes with uncertainty, so people still need to review what they produce.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

The article describes broader AI capabilities and warns that outputs can be unreliable, but does not focus on harm or human skill erosion.

How ChatGPT broadened the role of generative AI

Generative AI was once closely tied to the task it had been trained to perform. ChatGPT helped change that picture: large language models became easier to direct, and could be applied to work their creators had not specifically trained them to do.

From specialist model to flexible tool

Earlier AI systems were often built for narrow purposes. Google’s AlphaFold, for example, was trained on protein structure data to predict protein folding. Applying AI in a different field could mean investing time and money in a model designed for that specific work.

A robotics startup’s chief technology officer expected the same approach would be necessary for robots. The team found that, in many cases, off-the-shelf ChatGPT could help control robots without having been specifically trained for robotics. Technologists in areas including health insurance and semiconductor design have reported similar possibilities.

ChatGPT itself made large language models more responsive to people interacting with them. Its successors, GPT3.5 and GPT4, can also serve as general-purpose information-processing tools. That does not mean they know everything or can do anything. The comparison is to a general-purpose CPU: a tool that can be applied to many kinds of work, rather than a specialized chip designed for one kind of processing.

Why language models can be flexible

Large language models such as GPT4 work by estimating which words and phrases are likely to fit an input, then generating a likely continuation. This resembles a sophisticated form of autocomplete. The model is not directly sorting every possible answer into “right” and “wrong”; it is working with what is more or less likely given the input.

That approach has trade-offs. A model can produce useful responses across varied tasks, but its output can also be unpredictable and inexact. The same flexibility that lets it handle unexpected requests means people cannot assume that a plausible-sounding answer is correct.

Training shapes those probabilities. With gradient descent, a model’s outputs are compared with training data, and its parameters are adjusted in a direction that makes the outputs more like that data. Repeating this process in small steps can turn a network that produces incoherent text into one that generates coherent sentences. Related techniques can be adapted to pictures, DNA sequences, and other data.

Fine-tuning changes what a model can do

Once a model has been trained, gradient descent can also be used to adapt it. Fine-tuning starts with an existing model and trains it on a curated set of data to make it better at a specific kind of material or task.

For instance, Stable Diffusion has been fine-tuned to make anime images or landscapes. Language models can likewise be adapted for work with advertising copy or legal documents. Fine-tuning can also affect how a model responds and produces output, not only which subject areas it handles well.

That ability to adjust an existing model helps explain how generative AI can be useful beyond the purpose for which it was first trained. In some cases, a team may be able to give a model new information or instructions for a task rather than create a new model from scratch.

Useful does not mean autonomous

ChatGPT’s conversational format made it easier for people to express what they wanted from an AI model. But using these systems as general-purpose tools can involve a different approach: programming the model’s use and providing new data, rather than simply chatting with it.

The practical case for these tools is that they can multiply human productivity across different kinds of work. Their probabilistic nature still calls for care. Just as code needs review and testing before it is put into production, AI output needs processes for checking it. The model’s ability to attempt many tasks does not remove the need for people to assess what it produces.