Skip to content

Nearly half of AI-written code fails security tests. Who is checking yours?

In Veracode's 2026 tests, 44% of AI coding tasks produced code with a known security flaw, and newer models haven't closed the gap. Where AI-written code goes wrong, why it gets through review, and the checks that catch it.

Two developers working on code at their monitors in an open office
AI now writes a large share of the code that ships. Someone still has to check it.

By Mayrian

· 3 min read

Share

AI coding assistants have gone from novelty to default in about two years. In Sonar's 2026 State of Code survey of more than 1,100 developers, AI accounted for 42% of all committed code, and developers expected that to reach 65% by 2027. The code compiles, the tests pass and the feature ships. The question is what else ships with it.

What the tests found

Veracode tests AI-generated code for security using 80 coding tasks in Java, JavaScript, C# and Python, and has run them against more than 150 models. Its 2026 GenAI Code Security Report found that 44% of tasks produced code with a known, risky vulnerability. The average security pass rate was 56%, barely changed from 55% in its first report a year earlier.

Bar chart of security pass rates for AI-generated code: insecure cryptography 87%, SQL injection 83%, cross-site scripting 15%, log injection 12%, average 56%
Share of AI coding tasks that produced secure code. Source: Veracode, 2026 GenAI Code Security Report.

The failures aren't evenly spread. Models mostly avoid SQL injection (83% pass) and weak cryptography (87%), but defend against cross-site scripting only 15% of the time and log injection 12%. Those two matter because they're everywhere: almost every web application displays user input and writes logs. In Veracode's spring update, Java code passed just 29% of the time. The best model, GPT-5.5, reached 68%, which still means failing about one task in three. Larger and coding-specialized models did no better on average: most models scored between 50% and 53%.

Why the flaws get through

Syntax is a solved problem: modern models produce code that compiles almost every time. Security depends on context the model often lacks, such as where a value came from and where it will be displayed, which is exactly what cross-site scripting and log injection turn on. Models also learn from public code, which includes plenty of insecure examples, and they reproduce common patterns rather than your application's rules. Code that looks clean and works in a demo is easy to approve.

Often it isn't checked at all. Sonar found that 96% of developers don't fully trust AI-generated code to be functionally correct, yet only 48% always check it before committing. Sonar credits Amazon CTO Werner Vogels with a name for the backlog that builds up: verification debt. Confidence makes it worse. A survey cited by the Cloud Security Alliance found that nearly 80% of developers believe AI tools write more secure code than humans do.

The consequences are showing up

The Cloud Security Alliance reports that a Georgia Tech project tracking published vulnerabilities traced to AI coding tools counted 6 in January 2026, 15 in February and 35 in March, and its researchers believe the real number is five to ten times higher. The same note cites research finding that about 20% of AI-generated code samples reference software packages that don't exist, an opening for attackers who publish malicious packages under those names. It also cites enterprise research in which developers using AI produced commits three to four times faster but introduced security findings at ten times the rate, and a scan of 1,400 apps built with AI app builders that found 2,038 highly critical vulnerabilities and more than 400 exposed secrets.

Who should be checking, and how

  • Automated checks on every change. Static analysis, dependency scanning and secret detection run before code merges, not at the end of a project.
  • Human review where it matters most. Logins, permissions, payments, cryptography and anything that handles user input get a named reviewer.
  • Limits on sensitive code. The Cloud Security Alliance recommends policies that restrict AI assistance on authentication, authorization, cryptography and input validation.
  • Security in the instructions. Project rules and prompts state your security requirements, so the assistant starts from them.
  • Verified dependencies. Every new package is checked against a real, maintained source before it's installed.
  • A record of what AI wrote. Knowing which code was generated tells you what to review and re-test.

Ask your team or vendor which tools write your code, which checks run before it merges, and who reviews security-sensitive changes. If a vendor builds your software, ask them to show you, not just tell you. AI makes building faster. It doesn't make checking optional.

Our custom software team uses AI assistants with these checks built into every project. If you'd like a second look at code you already have, we can help.

How we can help

Choose a service to see its capabilities. Point at one to see what it's used for and what you receive.

All services

Custom Software Development

Software Development

Web platforms and internal tools built around how your business works, released in small, tested increments.

Used for

  • Customer portals
  • Internal tools
  • SaaS products
  • Workflow and approvals
  • Marketplaces and booking platforms
  • Enterprise application extensions

What you receive

  • Source code in your repository, with a documented architecture
  • CI/CD pipeline and infrastructure as code
  • Automated test suite
  • Monitoring, logging and alerting
  • Security review against the OWASP Top 10:2025 and ASVS 5.0
  • Ownership and documentation handover
More on Software Development

Working on something like this?

Start with a free technical consultation: a plan covering the right tech stack, architecture, timeline and budget.

Start a project