Why AI Coding Assistants Can Make Software Less Secure

A Stanford study found that developers using Codex were more likely to produce insecure answers to security-related programming problems—and more likely to believe those answers were safe. The researchers argue that AI coding tools can help with lower-risk work, but their suggestions need careful review.

WTF Index IDIOCRACY
◄ Terminator 2 Idiocracy 3 ►

The study suggests Codex can erode developers’ security judgment by encouraging insecure code and misplaced confidence.

Why AI Coding Assistants Can Make Software Less Secure

AI coding assistants can speed up programming, but their suggestions may also introduce security flaws. A Stanford study of developers using Codex found that access to the system was linked to more incorrect and insecure answers to security-related problems, along with greater confidence that those answers were secure.

What the Codex study examined

The researchers recruited 47 developers, from undergraduate students to industry professionals with decades of experience. Participants worked on security-related programming problems in Python, JavaScript and C, with some using Codex and others serving as a control group.

Codex, developed by OpenAI, powers Copilot. It was trained on billions of lines of public code and suggests code based on a developer’s description of a task and the context already in the project. That can make it useful for producing code quickly, but the study suggests that a plausible suggestion is not necessarily a secure one.

Compared with the control group, participants with access to Codex were more likely to write solutions the researchers considered incorrect and insecure. They were also more likely to judge their own insecure answers as secure. That second result matters because a developer who trusts flawed code may be less likely to examine it closely.

Speed does not replace security judgment

The study’s authors did not present their findings as a blanket rejection of code-generating systems. Megha Srivastava, a postgraduate student at Stanford and a co-author, pointed out that participants lacked security expertise that might have helped them recognize vulnerabilities. The results therefore describe how this group performed on these tasks, rather than establishing that every developer or use of Codex will have the same outcome.

Neil Perry, a PhD candidate at Stanford and the study’s lead co-author, said developers should be particularly careful when using these systems for work outside their expertise. Even when a task is familiar, he advised checking both the generated code and the context in which it will be used in the larger project.

The distinction is practical: an assistant can suggest a way to complete a task, but developers still need to assess whether that approach fits the software and its security needs. Familiar-looking output may save time, yet it does not by itself demonstrate that the code handles security risks correctly.

Where the researchers see value

Srivastava said code-generating tools can be reliably helpful for lower-risk work, such as exploratory research code. She also suggested that fine-tuning could improve their coding suggestions. Systems trained further on a company’s own source code might produce output that better matches its coding and security practices, she said.

That possibility does not remove the need for review. The study’s concern is partly about how people respond to generated answers: users may be inclined to trust a solution even when it contains a vulnerability. Better alignment with a team’s practices could help, but the article does not establish that it would prevent insecure code.

Possible safeguards and broader concerns

The co-authors suggested refining users’ prompts to encourage more secure code, in a way Perry compared to a supervisor revising a rough draft. They also recommended that developers of cryptography libraries make secure settings the defaults, because code-generating systems tend to use default values that are not always free of exploits.

Security is not the only concern raised in the article. It notes that at least some of the code used to train Codex was under restrictive licenses, and that Copilot had generated code from sources including Quake, personal codebases and books. Some legal experts have argued that companies and developers could face risk if copyrighted suggestions were unknowingly incorporated into production software.

GitHub introduced a filter for Copilot in June that compares a suggestion and about 150 characters of surrounding code against public GitHub code, hiding suggestions when there is a match or near match. The article describes the measure as imperfect: Texas A&M University computer science professor Tim Davis found that enabling it caused Copilot to emit large portions of his copyrighted code, including attribution and license text.

The researchers’ broader message is caution about using code-generation tools as a substitute for teaching strong coding practices, especially to beginning developers. These systems may assist with programming, but security still depends on people evaluating what they produce and how that code is used.