GPT-4 brought several changes to OpenAI’s conversational AI: it could take images as input, work with longer text and perform better than GPT-3.5 on a range of reported evaluations. OpenAI also described efforts to improve safety, while stressing that the system remained imperfect.
Images become part of the conversation
GPT-4’s visual input capability lets users ask questions about an image rather than describe it first. OpenAI’s examples included explaining a meme, examining a visual motif, breaking down an infographic and summarizing a scientific graph.
That changes the kinds of tasks a text-based assistant can take on. A user might provide a chart and ask for an explanation of its parts, or share an infographic and request a step-by-step account. OpenAI said GPT-4 outperformed existing text-image models on common benchmarks, while noting that it was still discovering new visual tasks the system could handle.
Longer inputs and stronger performance
OpenAI characterized GPT-4 as more creative and collaborative than previous AI systems, with a broader knowledge base and stronger problem-solving ability. It highlighted structured tasks such as producing step-by-step instructions in response to a question about cleaning an aquarium.
The company also pointed to a simulated bar exam, where GPT-4 was expected to score in the top ten percent, compared with GPT-3.5 in the bottom ten percent. On common machine learning benchmarks, OpenAI reported that GPT-4 outperformed its predecessor by up to 16 percent, and by 15 percent on multilingual tasks.
GPT-4 could handle more than 25,000 words in some configurations, which OpenAI said could support longer documents and analyses. The available context depended on the version: one was limited to about 8,000 tokens, or roughly 4,000 to 6,000 words, while another could process up to 32,000 tokens, or about 25,000 words, with limited access at the time.
OpenAI said GPT-4’s training information ended in September 2021 and that the model did not learn from its own experience. These limits matter when interpreting its answers: more capacity for lengthy material does not mean the system has current information or continually updates itself.
Safety gains, with familiar shortcomings
OpenAI said GPT-4 was built using lessons from adversarial testing and feedback on ChatGPT. In the company’s internal evaluations, it performed on average 40 percent higher than GPT-3.5 on factuality, with average accuracy scores between 70 and 80 percent. OpenAI also said it was 82 percent less likely to answer critical requests that violated its content policies and 29 percent more likely to provide policy-compliant responses to sensitive questions, including medical topics.
Those results came with important caveats. OpenAI said GPT-4 was “far from perfect”; it could still produce hallucinations, and it continued to create or reinforce biases. The company described its safety work as ongoing and said it was developing ways to predict model performance in some domains, using models trained with only one-thousandth the computational effort of GPT-4.
OpenAI also began using GPT-4 to help people evaluate AI outputs, describing that work as the next phase of its alignment strategy. For developers using the API, system messages offered a way to shape the character of responses, such as making them more like a Hollywood actor or more Socratic.
Access, costs and what OpenAI left undisclosed
At launch, OpenAI made GPT-4 available to paying ChatGPT Plus customers. The service cost $20 per month and was available internationally. Developers could seek API access through a waitlist.
API pricing was higher than the prices stated for ChatGPT and GPT 3.5. For the 8,000-token version, OpenAI listed $0.03 per 1k prompt token and $0.06 per 1k completion token. For the 32,000-token version, the listed prices were $0.06 per 1k prompt token and $0.12 per 1k completion token. The article described gpt-3.5-turbo as costing about $0.002 per 1000 tokens.
OpenAI did not provide details about GPT-4’s architecture, model size, hardware, training computation or dataset construction, citing competition in the market. The absence of a parameter count also left readers without one familiar way of comparing model sizes; the article noted that parameter count alone does not determine quality.
OpenAI identified early customers including Duolingo, Be My Eyes and Morgan Stanley Wealth Management, as well as the Icelandic government, which was using GPT-4 to preserve its language. The varied examples show the range of applications being explored, while the reported limitations remain part of the picture users and developers must consider.