Open multimodal AI expands image analysis, but leaves safety gaps

LLaVA-1.5 and Fuyu-8B bring image-and-text analysis to developers through open source models, with different strengths and limitations. Their capabilities raise practical questions about accuracy, bias, safeguards and how developers may use them.

WTF Index TERMINATOR
◄ Terminator 2 Idiocracy 1 ►

The story mildly leans toward Terminator because image analysis raises privacy and safety concerns, though it mainly reports an open-source model release.

Open multimodal AI expands image analysis, but leaves safety gaps

Open source multimodal AI models can interpret images alongside text, bringing some of the capabilities associated with GPT-4V to a wider range of developers. LLaVA-1.5 can identify objects and explain some visual context, while Adept’s Fuyu-8B is designed with screens, charts and other workplace information in mind. Their release also highlights unresolved questions about accuracy and safety.

What image-and-text models can do

A multimodal model processes more than one kind of information. In these examples, it can consider an image and a written question together, which lets it respond to visual details instead of relying on text alone.

That can make explanations more useful. A model might describe how to fix a bicycle, or suggest recipes based on ingredients visible in a refrigerator. The task goes beyond naming objects: the model must interpret how details in an image relate to the question.

OpenAI initially delayed releasing GPT-4V over concerns that people might use it to identify individuals in images without their consent or knowledge. The article also notes flaws OpenAI acknowledged in GPT-4V, including problems recognizing hate symbols and discrimination involving sex, demographics and body types.

LLaVA-1.5 makes image analysis easier to try

LLaVA-1.5 was released by researchers from the University of Wisconsin-Madison, Microsoft Research and Columbia University. It builds on an earlier LLaVA model by combining a visual encoder with Vicuna, an open source chatbot based on Meta’s Llama model.

The researchers expanded the training data, including by increasing image resolution and adding conversations from ShareGPT. The earlier LLaVA team had used text-only versions of ChatGPT and GPT-4 to generate conversations, questions, answers and reasoning problems from image descriptions and metadata.

One practical distinction is hardware access. The article describes LLaVA-1.5 as one of the first multimodal models that can run on consumer-level hardware with a GPU containing less than 8GB of VRAM. Its larger version has 13 billion parameters and can be trained in a day using eight Nvidia A100 GPUs, for a few hundred dollars in server costs.

Tests by Roboflow software engineers James Gallagher and Piotr Skalski found that LLaVA-1.5 could detect a dog in an image and identify where it appeared. It also recognized what was unusual in an image of a person ironing clothes on the back of a taxi, describing the scene as unconventional and potentially dangerous.

Those results had limits. The model struggled to distinguish multiple coins in a busy image and did not reliably read text in a webpage screenshot. The article contrasts this with GPT-4V, which handled that text-recognition task without the same problems.

Fuyu-8B focuses on workplace information

Adept released Fuyu-8B as an open source model aimed at sharing its work and gathering feedback from developers. Adept CEO David Luan described the company’s broader goal as building a copilot that knowledge workers could teach to perform computer tasks.

Fuyu-8B has 8 billion parameters. Adept says it performs well on standard image-understanding benchmarks, uses a simple architecture and training procedure, and answers questions quickly. Its intended subject matter includes websites, software interfaces, screens, charts, diagrams and ordinary photographs.

The model is meant to support tasks such as locating a specific screen element, extracting information from a software interface and answering questions about charts. But the article makes a distinction between the base model and those capabilities in Adept’s products: Adept fine-tuned larger, more sophisticated versions for document and software understanding.

Fuyu-8B is not licensed for commercial use, because some of its training data was provided to Adept under restrictive terms, according to Luan. The article likewise says LLaVA-1.5’s use of ChatGPT-generated training data creates a commercial-use concern under ChatGPT’s terms of use.

Capability and safeguards remain linked

Open access can help developers experiment, but it also means they may need to make important choices about model safeguards. The article reports that LLaVA-1.5 did not have the same toxicity filters as GPT-4V in one test: when asked for advice about a pictured larger woman, it recommended that she manage her weight and improve her physical health, while GPT-4V refused.

Luan said Fuyu-8B was released as a base model without moderation mechanisms or prompt-injection guardrails. He argued that protections should fit the intended use case, and said the model’s smaller size might reduce the chance of serious downstream risks. Adept had not tested it on CAPTCHA extraction, however.

The article also describes ways image-based instructions may be used to bypass safeguards, including attempts to get GPT-4V to solve CAPTCHAs. Strong text recognition could therefore bring both benefits and additional risks. The comparison between these models is not only about what they can recognize, but also about whether they respond accurately and how developers prepare them for use.

Both projects show how open source multimodal AI can make image understanding more accessible. Their reported limitations—from missed text to bias and absent guardrails—also show why model capability alone does not settle whether a system is ready for a particular application.