Blog

Claude's /security-review found 31 vulnerabilities, and 5 were real

A product with four years of security work behind it ran Claude's /security-review. It reported 31 vulnerabilities. After the engineers checked each one, 5 were worth fixing

Free, no call

Get every Fix AI Slop Code episode

The code it wrote, the code it should have written and a check, for every episode

Free, straight to your inbox. No call, no pitch

The short answer

Claude's /security-review found 31 vulnerabilities, and 5 were real

Claude's /security-review command is useful only if someone who understands security reads what it reports. On a product with its own security team, it reported 31 vulnerabilities. After the engineers ranked them, 5 were worth acting on, and the rest were low severity or already handled.

Is Claude's /security-review command worth running?

Yes, as a first pass. It is fast and it reads the whole codebase. Treat what it reports as a list of questions for a person who knows security, and do not treat it as a list of confirmed problems

Why does it report so many false positives?

It cannot see everything around the code: the network rules, how the app is deployed, or the fixes the team already made somewhere else. So it flags code that looks risky on its own but cannot be reached or is already covered

What if I do not have a security person?

Get one to look before you act on the list. Chasing every finding wastes days, and it can leave you thinking the app will never be secure when most of the list does not matter

The Blinkz team · Updated

Watch the video

Watch what happened

The same idea in a short video. It plays here, and nothing loads until you press play

What happened?

A product that has been in the market for four years ran the command against its codebase. The team already had a security group, site reliability engineers, and people whose job is to track the advisories and severity ratings for every package the product uses.

The command came back with 31 vulnerabilities. Engineers and managers sat down together and went through them one by one. Five were real issues, from medium to high risk. The other 26 were low severity or lower, or had already been taken care of.

Why is a list of 31 a problem?

A long list looks thorough. If you cannot rank it yourself, it does the opposite of what you want. Either you chase every line and lose a week, or you decide your app can never be secure and stop trusting your own code. A short list that is right is worth more than a long list that is mostly noise.

How should you read the findings?

  • Can it be reached? A risky function that no user input ever touches matters less than a small mistake in your login.
  • Is it handled somewhere else? A check in your middleware, a database rule or a network setting may already block it, and the model could not see any of those.
  • What would it cost you? Rank what is left by the damage: data that leaks, money that moves, or accounts that someone else can take over.
  • Write down the dismissals. For every finding you do not fix, write one line on why, so the next person does not have to redo the work.

Where do AI security checks fit?

Run the command. It is a cheap first pass, and it sometimes finds something real. Then put a person who does security for a living between that list and your next week of work.

On the apps we check, the problems that actually hurt people are rarely exotic. They are the ones in our Fix AI Slop Code series: rate limits that do not limit, database rules that are on but let everyone in, and retries that charge twice. If you want an engineer to read your app and tell you which of these it has, that is what an App Check is.

Free, no call

Get every Fix AI Slop Code episode

The code it wrote, the code it should have written and a check, for every episode

Free, straight to your inbox. No call, no pitch

Bring what you have. We’ll tell you where it stands

You’ll leave the call knowing what’s risky and what to do next

  • 30 minutes
  • Free
  • No slides, no sales team