GPT-4 with vision can analyze images alongside text, but OpenAI’s technical paper makes clear that this ability comes with limits. The company describes safeguards intended to reduce privacy risks, harmful bias and misuse, while also documenting situations where the model gets visual information wrong.
Safeguards address several kinds of risk
OpenAI internally abbreviates the image-capable model as “GPT-4V.” Until recently, regular use had been limited to a few thousand people using Be My Eyes, an app that helps blind and low-vision people navigate their surroundings. OpenAI also worked with “red teamers” to test for unintended behavior.
The paper says safeguards are intended to prevent uses such as breaking CAPTCHAs, identifying people, estimating their age or race, or drawing conclusions unsupported by a photo. OpenAI also says it has worked to reduce harmful biases relating to people’s appearance, gender and ethnicity.
Those protections reflect the challenge of giving a model more ways to interpret personal images. OpenAI says it is developing processes that could let GPT-4V describe faces and people without identifying them by name. That distinction points to an effort to preserve some descriptive ability while limiting exposure of personal identity.
Visual errors can change the meaning of an image
The paper also describes basic interpretation problems. GPT-4V can combine separate strings of text in an image into a term that does not exist, miss characters or mathematical symbols, and fail to recognize objects or place settings that appear obvious. Like the base GPT-4, it can hallucinate: presenting invented information in a confident tone.
These mistakes matter because a fluent explanation can sound reliable even when the image has been read incorrectly. A mistaken symbol, overlooked label or invented connection between pieces of text can shift the answer away from what the image actually shows.
OpenAI explicitly says GPT-4V should not be used to spot dangerous substances or chemicals in images. In tests, the model sometimes identified poisonous foods such as toxic mushrooms correctly, but misidentified fentanyl, carfentanil and cocaine from images of their chemical structures. The contrast shows why occasional correct answers do not establish that the tool is dependable for this purpose.
Medical images and symbols pose harder tests
Medical imaging is another area where the paper reports unreliable results. GPT-4V sometimes gave different responses to the same question in different contexts. It also did not consistently account for the standard practice of viewing scans as though the patient faces the viewer, where the image’s right side corresponds to the patient’s left. That gap can contribute to incorrect diagnoses.
The model’s difficulties extend to interpreting symbols and their context. OpenAI cautions that GPT-4V may miss the modern meaning of the Templar Cross in the U.S., where it is associated with white supremacy. The paper also describes cases in which the model made songs or poems praising hate figures or groups shown in an image, even when they were not named.
These examples involve more than recognizing what is pictured. They require connecting visual details to context and meaning, and the paper suggests GPT-4V can fail at that connection in ways that produce harmful or misleading responses.
Bias and safeguards remain part of the picture
OpenAI reports discriminatory responses about sex and body type when its production safeguards were turned off. In one test, a prompt asking for advice to a woman pictured in a bathing suit led GPT-4V to focus almost entirely on her body weight and body positivity. The paper’s example illustrates how a model’s answer can shift based on who appears in an image.
Overall, the paper presents GPT-4V as a system with useful image abilities and substantial limits. OpenAI says it has introduced safeguards, but also describes cases where the model can miss details, invent information, mishandle sensitive subjects or make biased assumptions. The paper’s account supports a cautious view: image analysis is not reliable enough to stand in for expert judgment in high-stakes situations.