How AI Chatbots Can Be Turned Into Tools for Scams

AI chatbots can be manipulated through hidden instructions, exposed to tainted training data, or drawn into phishing attempts when they interact with websites and email. Researchers say companies are aware of these risks, but there is no reliable fix yet.

WTF Index TERMINATOR
◄ Terminator 4 Idiocracy 1 ►

The article focuses on exploitable chatbots that can be manipulated through web access and used to enable scams or harmful actions.

How AI Chatbots Can Be Turned Into Tools for Scams

AI chatbots are being built into tools that browse the internet, manage calendars, and take actions for users. That can make everyday tasks easier, but it also gives attackers new ways to manipulate a chatbot or misuse the information it can access. Researchers describe several weaknesses, and say companies have not yet found a dependable solution.

Instructions can override a chatbot’s safeguards

Chatbots follow user prompts to generate responses. That ability is central to their usefulness, but it can also be exploited with prompt injections: instructions intended to make a model disregard earlier directions or safety guardrails.

Some users have tried to “jailbreak” ChatGPT by asking it to role-play as another AI that is free to ignore its restrictions. The resulting responses have included endorsements of racism or conspiracy theories, as well as suggestions to do illegal things such as shoplifting and building explosives.

OpenAI says it collects examples of jailbreak attempts and adds them to training data, while also using adversarial training to find ways its chatbots can be made to break their rules. But new prompts keep appearing, so this approach is an ongoing contest rather than a settled fix.

Web-connected assistants create new phishing routes

The risks grow when a chatbot can read websites and email, then act on what it finds. A hidden instruction on a webpage can target an assistant without being visible to the person reading the page. An attacker could try to direct users to a page through social media or email and use the concealed text to change the assistant’s behavior.

Researchers demonstrated how small such a manipulation could be. Arvind Narayanan added white text to his online biography that asked Bing to include the word “cow.” When he later used GPT-4, the generated biography mentioned cows. The example was harmless, but it showed that hidden text could influence a response.

Kai Greshake, a security researcher at Sequire Technology and a student at Saarland University in Germany, described a more dangerous demonstration. He placed a prompt on a website, then visited it using Microsoft’s Edge browser with Bing integrated. The chatbot produced a pitch that appeared to come from a Microsoft employee offering discounted Microsoft products and sought the user’s credit card information.

Hidden prompts in email could pose another risk. An assistant might be manipulated into exposing information from a victim’s messages or contacting people in their address book. These possibilities matter because an assistant may be able to take actions on a person’s behalf, not just summarize text.

Florian Tramèr, an assistant professor of computer science at ETH Zürich, warned that connecting chatbots to online content could be a serious security and privacy problem. Arvind Narayanan, a computer science professor at Princeton University, said text on the web could be crafted to make bots misbehave. The concern is that an attacker may exploit the assistant’s ordinary access to content, rather than persuade a person to run harmful code.

Training data can be tampered with

Some attacks could happen before a chatbot is released. Large AI models learn from vast collections of internet data, and researchers found that it was possible to introduce chosen material into those collections.

Tramèr and researchers from Google, Nvidia, and Robust Intelligence bought domains and filled them with selected images for $60. They also edited or added sentences to Wikipedia entries that were later included in an AI model’s data set. The researchers did not find evidence that data poisoning attacks had occurred in the wild.

Still, repeated material in training data can strengthen a model’s association with it. Tramèr said enough poisoned examples could influence a model’s behavior and outputs indefinitely. He also pointed to the economic incentive that could emerge as chatbots become part of online search.

There is no simple fix yet

Companies know about these weaknesses, but researchers say robust defenses remain elusive. Simon Willison, an independent researcher and software developer who has studied prompt injection, said there are currently no good fixes. Google and OpenAI declined to comment when asked how they were addressing the gaps.

Microsoft said it was working with developers to monitor possible misuse and reduce risks. Ram Shankar Siva Kumar, who leads Microsoft’s AI security efforts, said, “There is no silver bullet at this point.”

Narayanan argued that AI companies should investigate these problems more proactively. As chatbots gain access to more online material and user data, the security challenge extends beyond the chatbot’s own replies: it includes what the system reads, what it can reveal, and what it can do for a user.