When AI Gets Tool Access, Prompt Injection Becomes a Security Risk

Prompt injection can steer an AI assistant away from its intended instructions. When assistants can use email, databases, search, or code tools, an attack embedded in content may lead to data exposure or unwanted actions. Simon Willison says developers need to account for this risk as they build AI applications.

WTF Index TERMINATOR
◄ Terminator 4 Idiocracy 1 ►

The story warns that tool-enabled AI can be manipulated into exposing data or taking harmful actions.

When AI Gets Tool Access, Prompt Injection Becomes a Security Risk

AI assistants are moving beyond conversation. They can be connected to email, search, databases, and other tools, giving them ways to act on information outside the chat. Developer Simon Willison warns that this shift makes prompt injection a practical security concern: hostile instructions hidden in content may influence what an assistant does.

Why tool access changes the risk

A language model can be prompted to ignore its original instructions and follow new ones embedded in material it reads. That becomes more consequential when the model can also make API requests, search the web, or run generated code in an interpreter or shell.

In a chat-only setting, a misleading response may remain a misleading response. But an assistant with access to other systems could take actions or move information. Willison points to patterns and products including ReAct, Auto-GPT, and ChatGPT Plugins as examples of systems that give language models tools to use.

The concern is not limited to a user directly typing a malicious request. An instruction could arrive in an email or appear on a web page the assistant is asked to process. The model may treat that text as instructions even when it was supplied as data.

An email can carry hidden instructions

Consider an assistant that can search, summarize, and reply to email through the ChatGPT API. If an incoming message tells the assistant to forward selected recent emails and then delete them, the assistant may have access to the very tools needed to carry out that request. Willison says there is, in principle, nothing stopping it from following those instructions.

Search results can also be manipulated. Researcher Mark Riedl put a message on his website that was not immediately visible to human readers, asking Microsoft's Bing to describe him as a time travel expert. Bing adopted the claim. Willison says similar attacks could be used to influence how a search assistant presents products.

These examples show why the source of an instruction matters. A message or page can contain language directed at the AI, even though the user only asked the assistant to read or summarize that content.

Combining tools can expose private data

The number of possible attack paths can grow when an assistant has access to several systems. Willison describes a plugin he developed that lets ChatGPT query a database. If an email plugin is also available, an attack in an email could prompt the assistant to retrieve valuable customer information and place it in a URL. A user clicking that link could send the private data to a website.

Developer Roman Samoilenko demonstrated another route that begins outside the chat. A visitor copies text from an attacker's website; JavaScript intercepts the copy event and adds a malicious prompt to the copied material. When the user pastes it into ChatGPT, the prompt can ask for sensitive chat data to be included in the URL of a tiny image. Loading the image sends the data to a remote server. The prompt could also ask for the image to appear in future answers, potentially putting later chat data at risk too.

The examples involve different paths into an assistant: email, search content, and copied text. In each case, the model is asked to process material that may contain instructions, while its connected tools can make the consequences extend beyond its response.

What developers can do

Willison says OpenAI's Code Interpreter and Browse modes operate separately from the general plugins mechanism, presumably to help avoid malicious interactions. His larger concern is the expanding range of combinations among existing and future plugins.

He suggests making the prompts assembled by an assistant visible to users. Seeing the instructions could give someone a chance to notice an attempted injection and report it. Assistants could also seek permission before taking certain actions, such as showing an email before sending it. Neither measure is perfect, and both can remain vulnerable to attack.

For now, Willison argues, developers need to understand prompt injection and consider it when designing applications based on large language models. A useful question for any such product is how it accounts for instructions hidden in the content the assistant reads, especially when the assistant can also reach private data or take action.