AI-Generated Code Security Risks: 61% Correct, 10.5% Safe

AI-Generated Code Security Risks: 61% Correct, 10.5% Safe

Veracode's March 2026 update on AI-generated code security risks measured a security pass rate of roughly 55% across models, flat over the entire review period, while coding benchmarks like HumanEval showed consistent improvement.

The models got measurably better at writing code that runs and made no measurable progress at writing code that's safe. And those are not the same skill. Every other number in this piece sits underneath that one divergence.

The short version of whether to trust generated code: trust it to run, never trust it blindly where an attacker can reach it. The best-performing agent in a 2025 benchmark of 200 real-world tasks produced functionally correct code 61% of the time and secure code 10.5% of the time, per the SusVibes benchmark data. That's the whole problem in two numbers.

And the rest of this post is what those numbers mean for people shipping with a small team.

Working code and safe code are different skills

Start with the benchmark that isolates the gap. On those 200 real-world tasks, the top agent was functionally correct 61% of the time but secure only 10.5% of the time (SusVibes). An agent can ace "does it do the job" and flunk "does it resist attack" on the same task. My read is that nothing in the benchmark world rewards resisting attack. And nothing in the training loop does either.

Field data agrees. The Cloud Security Alliance's July 2025 research found 62% of AI-generated code solutions contained design flaws or known security vulnerabilities, even on the latest foundational models.

A February 2025 study in the ACM Transactions on Software Engineering and Methodology analyzed 733 real-world snippets from GitHub Copilot, CodeWhisperer.

And Codeium, and found weaknesses in 29.5% of the Python snippets and 24.2% of the JavaScript snippets, spanning 43 distinct CWE types (CSA summary).

Language moves the numbers hard in Veracode's testing: Java failed security checks 72% of the time, C# 45%, JavaScript 43%, Python 38%. If you're assembling a Java service with an assistant, you're sitting in the worst bucket on the board.

I'd treat that 72% as a planning input, not a trivia fact.

The official CVE count is a floor, and it's climbing fast

Georgia Tech's Systems Software and Security Lab launched the Vibe Security Radar in May 2025 to answer one question: how many publicly filed CVEs trace back to AI-generated code? The answer so far is a curve, not a count. The radar logged 6 such CVEs in January 2026, 15 in February. And 35 in March, a near-sixfold increase in two months. And that single March figure exceeded the total documented for all of 2025.

Here's the part that should bother you more.

The researchers estimate the true figure runs 5 to 10 times higher. Because most AI-generated code carries no metadata connecting it to the tool that wrote it. Attribution is the bottleneck. Code that can't be identified as AI-generated can't be counted, which means the public record undercounts everyone's actual exposure by design.

The tools themselves are attack surface too.

The December 2025 IDEsaster disclosure by researcher Ari Marzouk, which the CSA describes as the most comprehensive single-researcher audit of the category to date, identified more than 30 vulnerabilities across at least 10 products and resulted in 24 assigned CVEs. Your assistant writes risky code and is itself a risky program, and both directions deserve a policy from you.

Clean-looking code is exactly what gets past review

A December 2025 CodeRabbit study compared 320 AI-co-authored pull requests against 150 human-only pull requests from open-source repositories. The AI-authored PRs generated 1.7 times more issues overall and 2.74 times more security issues specifically. When you merge a generated PR, you're merging nearly triple the security issue load of a human one.

The uncomfortable mechanism comes from IEEE ISSRE 2025, where researchers analyzed more than 500,000 code samples in Python and Java from human developers and from models including ChatGPT, DeepSeek-Coder, and Qwen-Coder. AI code carried more high-risk vulnerabilities despite having simpler structural complexity than the human code (study summary).

That finding is a trap for reviewers, given that simpler code reads as trustworthy code.

Human reviewers are calibrated to slow down for tangled, weird, human-written mess and to wave through clean, idiomatic, confident-looking output. Generated code always looks like the latter.

The 2.74x issue rate survives review partly as review itself was built around human failure patterns.

What I'd do as a team of one

If you're a solo operator or a two-person shop, this lands harder on you than on an enterprise with a security org. The same person who wrote the prompt is the reviewer. And the data says the security review burden per PR nearly tripled. You can't hire your way out at your size, so process has to carry it.

My protocol is plain and boring on purpose:

- Treat generated code as untrusted input, the same policy you'd apply to data submitted by strangers. It enters your repo on probation, not on trust. - Put a scanner in CI and make it gate every AI-assisted commit. Your eye is not calibrated for these failure modes, and a scanner at least enumerates them. - Keep a written inventory of AI-assisted files. The attribution gap that keeps the public CVE count at a fraction of the real number will hide your own exposure too, unless you track it yourself. - Review line-by-line anything that touches authentication, payments, file handling, or database queries, no matter how clean it reads. Especially if it reads clean.

This costs time, which is exactly the time the assistant saved you.

That's the actual trade, and most people who got burned accepted it without ever pricing it.

The flat line is the story

The pass rate held near 55% while HumanEval climbed, and that tells you what the market optimizes for. Capability demos sell subscriptions, and no demo ever died on stage since of a missing security check. I don't expect the flat line to move until buyers start measuring it. And small teams don't have the option to wait for that.

So measure your own corner of it this week.

List the AI-assisted files in your active repo, run a scanner across them. And fix what it flags before you build the next feature on top. The numbers say AI-generated code security risks are compounding faster than anyone's review cycle. And an hour of cleanup beats explaining a breach that rode in on a pull request you approved without reading.