Forty-five percent of AI-generated code contains security flaws. That is Veracode's 2025 number, and the follow-ups keep landing in the same grim neighborhood: CodeRabbit found AI-generated code 1.88x more likely to introduce vulnerabilities than human-written code, and an audit of 5,600 publicly available vibe-coded apps turned up more than 2,000 high-impact vulnerabilities plus 400 exposed secrets. Coding agents now ship code faster than human-led review cycles can read it — which is exactly why the most interesting security launch of May 2026 was not a better scanner, but a scanner that stopped being a dashboard.
On May 26, Swedish application-security platform Detectify launched the Detectify MCP Server, an integration layer that plugs its security testing engines into AI-driven coding workflows so agents can find, validate, and remediate exploitable vulnerabilities without leaving the loop. CEO Rickard Carlsson put the thesis in one line: "We are expanding from a dashboard humans check to a skill agents orchestrate." For a git-push PaaS, the takeaway is concrete: a scan phase between build and promote that your deploy agent can invoke, read, and gate on — every tenant deploy scanned before it ships, with triage and sign-off staying human. This post walks the Find & Fix loop step by step, sketches that pipeline, and names what still needs a person.
What shipped: the Find & Fix loop, step by step
The MCP Server's headline capability is "Find & Fix" automation, and it is worth stating as the literal loop an agent runs, because that loop is the whole product:
- Findings arrive as structured remediation tasks. Instead of a report a human triages, Detectify hands vulnerabilities to the agent as tasks it can act on — severity, domain, discovery date filterable, attack-surface data attached.
- The agent generates a patch. The coding agent does what it already does — writes the fix — but now against a finding with scanner-grade evidence rather than a linter hint.
- The agent triggers a validation scan. This is the step dashboards never had: the same agent that wrote the patch asks Detectify's engines to re-test and confirm the vulnerability is actually resolved.
- A verified fix goes to human review. The loop ends with a person, holding a patch plus a passing validation scan — not a raw alert queue.
Around that loop sits a conversational surface: query scan results, check asset status, pull high-severity findings for the month, all through natural-language prompts in Claude Code, Cursor, ChatGPT, or Claude Desktop. The server is remotely hosted — nothing to deploy or maintain — with per-request OAuth and a deliberately pre-selected tool surface: Detectify decides which tools and data agents may touch rather than exposing the whole platform API.
One honest footnote from Detectify's own launch post: public release went on brief hold while the company runs third-party security testing on the server, which it calls a best practice for any new MCP integration. A security vendor dogfooding "test your MCP server before agents trust it" is the right kind of slow, and it previews the governance theme this post returns to at the end.
What this means for your deploy pipeline
Here is the pipeline sketch the launch implies for any git-push platform — build, scan, promote, with the agent operating all three:
Build → Scan → Promote. The scan step sits between image build and production promotion, and it is agent-addressable: the deploy agent invokes the scan the way it invokes the test suite, reads findings as structured data rather than screenshots of a dashboard, and gates promotion on the result. "Every tenant deploy scanned before it ships" stops being a policy document and becomes a pipeline phase with a machine-readable verdict.
The gate policy is explicit and agent-readable. Severity thresholds (block on critical/high, warn on medium), validation-scan confirmation for any finding the agent patched, and a freshness rule (the scan must cover this exact build, not last week's). An agent can enforce all three without judgment calls — which is the point. Judgment calls are what humans keep.
Humans keep triage, false-positive calls, and fix sign-off. The Find & Fix loop ends at human review for a reason: scanners, deterministic or not, still flag things that are not exploitable in context, and only a person can accept that risk or send the patch back. The division of labor is clean — agents do invoke, read, and gate; humans do interpret, waive, and approve. Any pipeline that lets an agent waive its own findings has built a rubber stamp, not a gate.
This is the "table stakes" argument in concrete form: if your platform's pitch is agent-operated deploys, "the agent can run the scanner" is not a premium integration — it is the minimum that keeps the agent honest. A deploy agent that can push but cannot scan is a junior engineer with production access and no code review.
Why a tool beats a dashboard
Carlsson's framing is worth quoting in full because it names the mechanism, not just the mood: "We aren't competing with the AI's reasoning, we are providing the professional-grade tools that reasoning requires. By structuring our capabilities as modular, high-performance building blocks, we allow agents to call our scanner as naturally as they call a test runner."
Two ideas in there matter. The first is the test-runner analogy: a scanner the agent invokes mid-loop, gets a verdict from, and iterates against — not a separate system that renders a verdict hours later for a different human. The feedback latency collapses from "next sprint's security review" to "this deploy's pipeline run."
The second is determinism. Language models reason probabilistically — they do not produce the same answer to the same query every time. Detectify pitches its compiled scanning engines as the deterministic verification layer agents need before code reaches production: the probabilistic part proposes the patch, the deterministic part confirms the hole is closed. That division is load-bearing. An agent reviewing its own security work with only its own reasoning is a closed loop with no ground truth; the scanner is the ground truth.
The scale argument completes it: Detectify monitors millions of changing domains, and the MCP server is meant to bring that standing observation into agentic workflows so security operates at engineering velocity. The attack surface keeps exploding because agents keep shipping; the only response that scales at the same rate is tooling the agents themselves invoke.
The vendor wave: this is a pattern, not a stunt
Detectify is the fourth major appsec vendor to make its tooling agent-addressable over MCP, and the dates show the pattern compressing:
| Vendor | Move | Date |
|---|---|---|
| Legit Security | MCP server to secure AI-generated code | June 2025 |
| JFrog | MCP server for the software supply chain platform | July 2025 |
| TrojAI | Defend MCP to secure agentic AI workflows | November 2025 |
| JFrog | Cursor AI coding agent with remote MCP connection + automated security rules, aimed at 1M+ AI developers | March 2026 |
| Detectify | MCP Server with Find & Fix + conversational attack-surface queries | May 2026 |
JFrog has gone furthest on the governance side, adding an MCP registry for vetting which servers enterprise agents may reach — a sign the wave is already entering its "now govern it" phase. When four vendors in twelve months converge on the same interface for the same buyer (the agent, not the human), the interface has won. Scanner-as-MCP-tool is what CI badges were in 2015: the thing you assume is there.
Honest caveats before you wire it in
Three limits keep this recommendation credible rather than breathless.
A remotely hosted scanner is a trust boundary, not just a feature. Your deploy agent now sends code and queries to a vendor's MCP server over the network. That server's identity, its data handling, and its own update cadence are your supply-chain problem — the same week this post was written, census data showed thousands of MCP servers sitting outside any governance boundary at all. Vet the tool like infrastructure: pin what you can, audit what you cannot, and keep the scanner's access scoped to what the pipeline actually needs.
Validation scans cost deploy latency. A re-scan per patched finding is fast compared to a human review cycle and slow compared to a unit test. Budget the minutes in your pipeline SLOs, parallelize scan-against-build where the scanner allows it, and decide up front which severities block promotion versus file follow-up tasks. A gate that always blocks on everything becomes a gate everyone routes around.
Attack-surface scanning is not the whole of security. Detectify's engines test the running perimeter — domains, APIs, exposed services. That does not replace SAST on the source, secret scanning on the repo, dependency review on the lockfile, or image scanning on the artifact. The Find & Fix loop is one deterministic layer; a serious pipeline stacks several, and the agent should be able to invoke all of them.
Verdict: the agent that ships must be the agent that scans
Detectify's launch is small as a product announcement and large as a signal: the security vendors have accepted that the reader of their findings is increasingly not a human with a dashboard but an agent with a tool call. The Find & Fix loop — findings as tasks, agent patch, validation scan, human sign-off — is the shape every scanner integration will converge on, because it is the only shape that keeps humans where judgment lives while letting agents carry the velocity.
For a self-hosted git-push PaaS, the checklist is short: put a scan phase between build and promote, make its verdict machine-readable, let the deploy agent invoke it like a test runner, and keep triage, waivers, and sign-off with people. The platform that ships that is not buying a premium integration. It is meeting the new minimum for letting agents touch production.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Agent-operated deploys need deterministic gates: star the repo on GitHub or deploy your first app today.
Sources
- Detectify blog, "Introducing the Detectify MCP Server to connect security intelligence into your AI workflows" (May 26, 2026) — Find & Fix loop, conversational surface, remote-hosted OAuth, pre-selected tools, third-party-testing hold.
- SiliconANGLE, Duncan Riley, "Detectify debuts MCP server to let AI agents find and fix vulnerabilities in real time" (May 26, 2026) — Carlsson quotes, determinism framing, millions of domains, vendor-wave context (Legit, JFrog, TrojAI).
- SC Media, "Detectify launches MCP server to integrate security testing into AI coding workflows" (May 26, 2026).
- BetaNews, "New MCP server helps secure AI coding" (May 26, 2026) — "dashboard humans check to a skill agents orchestrate."
- Business Wire / ADVFN, "Detectify Launches MCP Server to Secure the Autonomous Coding Loop" (May 26, 2026) — press-release capability list.
- DEVOPSdigest, "Detectify Announces MCP Server" (May 27, 2026).
- Veracode GenAI Code Security Report (2025) via CIO Dive / SQ Magazine — 45% of AI-generated code contains security flaws.
- CodeRabbit report (December 2025) via Spiceworks / Medium — AI-generated code 1.88x more likely to introduce vulnerabilities; AI-co-authored PRs ~1.7x more issues.
- Escape, State of Security of Vibe-Coded Apps via CIO Dive — 5,600+ apps audited, 2,000+ high-impact vulnerabilities, 400+ exposed secrets.
- Business Wire, "JFrog Brings Enterprise-Grade Software Supply Chain Security to Over 1M AI Developers with New Cursor AI Coding Agent" (March 31, 2026).
- JFrog press release (July 17, 2025) via investors.jfrog.com; InfoWorld (July 2025) — JFrog MCP server launch.



