How DetGPT Connects Everyday Requests to Objects in Images

DetGPT pairs image recognition with a language model so people can ask for help related to objects in a photo. Its ability to interpret open-ended requests could also make visual surveillance more flexible, raising questions about how systems apply social norms.

WTF Index TERMINATOR
◄ Terminator 2 Idiocracy 1 ►

DetGPT’s flexible image understanding could enable more adaptable visual surveillance, though the article mainly describes a helpful research system.

How DetGPT Connects Everyday Requests to Objects in Images

A photo can contain many objects, but a useful AI assistant needs to work out which one matters to a person’s request. DetGPT is a research model designed to connect natural-language instructions with specific objects in an image, from finding a refrigerator when someone wants a cold drink to identifying an alarm clock when they want to wake up.

From describing an image to acting on a request

At the GPT-4 launch, OpenAI demonstrated multimodal capabilities such as turning a photographed, hand-sketched web design into code and answering questions about images. Image question answering was also available through the “Be My Eyes” application. Open-source models such as miniGPT-4 offered an early glimpse of similar possibilities, though these capabilities were not yet widely available.

Researchers at the Hong Kong University of Science and Technology and the University of Hong Kong developed DetGPT as a more capable alternative to miniGPT-4. The model combines the BLIP-2 visual encoder, a 13-billion-parameter Robin or Vicuna family language model, and a pre-trained detector called grounding-DINO.

That combination is intended to let a user take a photo and ask for a particular piece of information or action involving something in it. The model aims to locate the relevant object as well as understand the request, including when the instruction is complex.

How DetGPT interprets everyday needs

In one example, a user says, “I want a cold drink,” while showing a kitchen with no visible drink. DetGPT identifies the refrigerator as the best option. The request does not name an object; the model has to connect the person’s goal to something that appears in the scene.

Other demonstrations use the same idea in different settings. Asked, “I want to wake up in the morning,” DetGPT marks an alarm clock on a cluttered table. Asked, “Which fruits help with high blood pressure?” it marks fruits on a market stall that might help with high blood pressure.

These examples show the difference between listing what is visible and responding to what a person means to do. The system uses the image to find candidate objects and the language model to relate those objects to a request. That can make image interaction more practical: a person can describe a goal instead of knowing the exact name of the object they need.

For training, the team created a dataset of 30,000 examples from 5,000 images in the COCO dataset, along with examples of image instructions generated by ChatGPT. The code and dataset are available on GitHub, and a demo is available. The model itself is not yet available.

The same flexibility raises surveillance concerns

The ability to interpret broad instructions could have uses beyond personal assistance. DetGPT, miniGPT-4, multimodal GPT-4, and robotics-focused models such as PaLM-E illustrate the potential of adding more kinds of input to large AI models. How widely such systems will be used remains unclear, but they could affect many jobs and may make some forms of government surveillance easier.

A Reddit post presenting the research describes a surveillance scenario in which people tasked with monitoring public places might only specify a few examples of behavior, such as holding a knife or smoking. The post argues that a model could be asked more broadly to “detect behaviors that violate public order” and use its own knowledge to identify other behaviors it considers relevant.

That possibility is a meaningful shift. A system working from a fixed list looks for patterns people have already labeled. A model interpreting a broad instruction could also infer undesirable behavior from surrounding factors, extending detection beyond those initial examples.

Who defines what the system should detect?

Such a system would first need a language model fine-tuned to the norms of a particular society. That requirement points to a difficult question: whose understanding of acceptable behavior would shape the model, and how would those judgments be reflected in its detections?

When an AI system moves from finding an alarm clock to judging conduct in public, the consequences of interpretation change. A mistaken object suggestion may be inconvenient. A broad surveillance instruction could affect people who are classified according to norms embedded in a model. The article’s examples therefore show both the convenience of flexible image understanding and the need to examine how it is directed.

DetGPT offers a preview of how language and vision can work together: a person states a goal, and the model points to something in a scene that may help. The same capacity to generalize from open-ended requests is also why the technology’s uses and boundaries deserve attention.