4,300 Unique Attacks in 30 Minutes: Superhuman Security Testing
One AI attack agent, one serverless login gate, 30 minutes: 4,300 unique attack probes and 105,000 credential guesses — certificate-transparency pivots, origin hunting, one accepted finding, and the finding the agent talked itself out of.
In 30 minutes, one PurpleSwarm attack agent sent 4,300 unique attack probes against a single login page — injection, request smuggling, token forgery, method tampering, cache poisoning, side-channel timing analysis — alongside 105,000 credential guesses and 2,200 infrastructure probes. No human pentester works at that rate. But volume alone is not the point: every probe was chosen for a reason, every response was read, and every next step depended on what came back. Machine-scale coverage with per-request judgment — that is what superhuman security testing looks like.
This post walks through that session, start to finish. The target: the pre-production environment of a B2B SaaS company, where the engagement's risk profile explicitly allowed credential attacks. Details are anonymized; the engagement was authorized, and every action below was logged.
The setup
The entire visible application was one page: a "Dev Environment" login. Username, password, submit.
One response header told the agent the whole story: X-Cache: LambdaGeneratedResponse from cloudfront. The site is not an app with a login page in front of it. It is an AWS Lambda@Edge viewer-request function attached to a CloudFront distribution — every path, every method, every request is intercepted at the edge and answered with the same 1,189-byte form. Behind the gate sits an origin that only speaks to CloudFront. The form is the only way in.
Rules of engagement: unauthenticated, black-box, non-destructive. The risk profile for this run explicitly allowed credential attacks — bounded wordlists, time budgets, full audit logging. More on what that means below.
A scanner's nightmare: the false-positive factory
Run a standard directory brute-force against this target and every word in your list comes back HTTP 200. /admin: 200. /.git/config: 200. /.env: 200. /wp-admin: 200. A status-code-based scanner "discovers" hundreds of sensitive paths on this site — every one of them a false positive — or worse, reports an exposed Git repository that is actually just the login page again.
The agent issued a handful of requests, noticed that every response was byte-identical, and reclassified the target: not an application, a gate. Then it proved the model before relying on it — 200 random paths, 40 auth-typical paths (/login, /oauth/token, /graphql, …), two dozen file extensions, every HTTP method including the WebDAV family, all returning the same form. Only then did it stop fuzzing paths and start attacking the gate itself.
That reclassification — from "website" to "edge function with a password" — is the move no scanner makes, because it requires understanding what the responses mean, not just what status code they carry.
What half an hour of agent coverage looks like
The most underrated way to read this session is as a coverage problem. A human pentester handed the same target would eventually try most of the same things — the offensive playbook is not secret. The constraint is time. Every technique costs minutes of setup, execution, and interpretation, so humans triage: they pick the handful of approaches most likely to work and defer the rest, hoping the report deadline allows a second pass.
The agent does not triage. In roughly half an hour, it worked through:
- Five credential-attack strategies — default lists, a 63,000-combination brand and environment mutation matrix, breach-corpus top passwords, seasonal patterns, and a person-targeted list built from the founders' OSINT footprint. ~105,000 attempts in total.
- Every injection class — SQLi, NoSQL operator injection (form and JSON), LDAP, SSTI, null bytes, type juggling — across 16 content-type and parser paths.
- Protocol-level attacks — more than two dozen HTTP verbs, request smuggling over raw sockets (CL.TE, TE.CL, eight Transfer-Encoding obfuscations, duplicate Content-Length), and HTTP/2 pseudo-header manipulation.
- Session and token forgery — dozens of cookie-name and token combinations,
alg=noneJWTs, CloudFront signed-cookie names, Basic auth. - Infrastructure intelligence — Certificate Transparency mining, passive DNS history, two subdomain brute-force lists (up to 1,900 names), origin-IP port sweeps, S3 bucket-name enumeration, full DNS record review.
- OSINT — GitHub organizations and members, three public code-search engines, four web archives, and the company's marketing site.
- Side channels — byte-level response differential analysis and 180 timed requests to check for constant-time credential comparison.
- Cache and intermediary attacks — cache-poisoning probes, unkeyed-header reflection, conditional-request quirks.
Add it up and the session produced roughly 4,300 unique, hand-crafted attack probes — each constructed for this specific target and its observed behavior, not replayed from a payload database — on top of the credential campaigns and infrastructure sweeps.
Every one of those is a technique a skilled human knows. Very few humans get past the first third of that list on a single login page before the clock forces them to move on — and no human enjoys the fortieth probe as much as the first. The agent's advantage is not superior technique; it is that the marginal cost of technique #35 equals the cost of technique #1. It never gets bored, never rounds "tested" up from "spot-checked," and when a hypothesis calls for a specialty — the gate sits behind a CDN, so request smuggling becomes relevant — it loads that methodology mid-session and applies it correctly.
A strong human pentester covers this breadth in days. The agent covers it before lunch — and writes down every negative.
Coverage scales with tokens, not headcount
There is a second axis to this session that matters for planning: the breadth above is not fixed. It is a function of budget.
This run had a ten-million-token budget. Watch what happened each time the agent was nudged to continue: it did not repeat itself — it opened new fronts. Round two added passive DNS history, a founder-targeted wordlist, content-type parser confusion, and cache-poisoning probes. Round three added GitHub organization reconnaissance, S3 bucket enumeration, DNS record analysis, web-archive searches, and timing-oracle statistics. Round four added WAF fingerprinting, HTTP/2 desynchronization probes, and CloudFront trusted-header fuzzing. Every additional slice of budget converted directly into additional, distinct attack coverage.
For a human pentester, more coverage means more days or more people — cost scales with time, and scheduling is the bottleneck. For an agent, more coverage means more tokens: a larger credential corpus (millions of guesses instead of 105,000), more exotic technique classes, deeper passes over every surface with fresh hypotheses. The budget is a dial you set per engagement, per target, per risk appetite.
And the dial has a built-in signal: diminishing returns. On this hardened target, each successive round produced fewer new hypotheses and no new findings — the later rounds mostly falsified increasingly exotic vectors. That plateau is itself evidence. When an agent with budget left to spend struggles to generate a new untested idea, "we have tested this thoroughly" stops being a figure of speech.
The pivots no scanner has
With the front door established as the only door, the agent widened the search in directions that never touch the target's HTTP responses:
- Certificate Transparency logs revealed two hostnames the application never references:
api-origin.dev.* andapi-origin.*— the origin servers behind the distribution, sitting on EC2. The agent port-scanned all five known origin IPs across a broad service-port sweep: everything filtered. Passive DNS history added a third, stale origin IP. Conclusion, recorded with evidence: no direct origin access, no bypassing the gate around the back. - The company's marketing site — a different property a scanner would never follow — yielded the founders' first names, the company's home city, and the community the founders met through. The agent built a person-targeted wordlist from exactly that: names, local themes, and year mutations.
- Public code and archive search: GitHub organizations (empty), Wayback Machine, urlscan, CommonCrawl, AlienVault OTX, three public code-search engines — zero references to the target, no leaked credentials, no historical captures.
- Infrastructure guessing: 13 plausible S3 bucket names (all nonexistent or fully private), DNS record review (clean), subdomain brute force (exactly one host resolves).
None of this is in any scanner's playbook, because none of it is a response to match a signature against. It is an agent forming hypotheses — where is the origin? who picks the passwords? has this credential leaked anywhere? — and then going wherever those hypotheses lead.
The finding: 62,920 proofs
The gate has no rate limiting. The agent did not assert that — it proved it, at scale, in three steps:
- 40 rapid sequential failed logins from one session: all HTTP 200, no slowdown.
- A 500-request parallel burst at ~243 requests/second: all HTTP 200, no 429, no CAPTCHA, no block.
- A full 62,920-attempt credential run sustained at ~297 requests/second for over three minutes: every response identical, no throttling at any point.
That finding was accepted as medium severity (CWE-307) with the complete reproduction evidence attached. The severity reasoning is the part a scanner can't do: this gate is the sole authentication barrier in front of the entire pre-production environment, and ~300 guesses per second is roughly 26 million guesses per day, forever, undetected.
The finding the agent talked itself out of
Midway through the session, the agent found that malformed percent-encoding in the POST body — a bare %, a truncated %C3%28 — crashes the Lambda function. CloudFront returns 503 with X-Cache: LambdaExecutionError. It filed a denial-of-service finding.
The validation step pushed back with the right question: a per-request crash fails the attacker's request — but does it affect anyone else?
So the agent designed an experiment. Eight threads flooded the gate with 297 malformed requests while a separate "victim" thread sent 37 legitimate logins. Result: 37 out of 37 served normally. Lambda@Edge is stateless per invocation; the crash is strictly per-request; there is no cross-user impact and no practical DoS.
The agent downgraded its own finding to a knowledge-base note: robustness bug, worth fixing, not a security finding.
A scanner reports "503 error, possible denial of service" and leaves a human to triage it. The agent measured the actual impact and refrained. Fewer findings, more truth.
Where credential attacks fit
This session sprayed roughly 105,000 credential combinations: default lists, brand and environment mutations, seasonal patterns, breach-corpus top passwords, and the founder-targeted list built from OSINT. Zero valid logins.
That only happened because this engagement's risk settings explicitly allow credential attacks. They are off by default. In runs where they are off, the agent does not guess a single password — it proves the preconditions for an attack instead (enumerable accounts, no throttling) and reports those as the finding. The risk profile decides which of the two proofs the agent is allowed to produce.
When credential attacks are enabled, they are bounded: curated wordlists rather than unbounded brute force, time budgets, a target scope the customer authorized, and every single attempt logged and auditable.
And notice what 105,000 open-throttle guesses bought here: nothing. That is not a failed attack — it is evidence. It is the strongest practical proof, short of source review, that this gate's credential is not derivable from any dictionary, brand theme, or founder biography. "We think the password is fine" became "105,000 targeted guesses failed." That is a quantifiable statement a customer can act on.
The negative space
The rest of the session was a systematic elimination of everything a login gate could theoretically be vulnerable to, each ruled out with evidence rather than silence: failure pages byte-identical for every input class (MD5-verified, so no oracle), no username reflection (no XSS surface at all), no CORS misconfiguration, no CRLF injection, no cache poisoning surface, and no denial-of-service via 60 MB request bodies.
Every one of those negatives was written back to the shared knowledge base, so the next agent — or the next run — starts from "already falsified" instead of repeating the work.
Scanner vs. agent
| Observation | DAST scanner | PurpleSwarm agent |
|---|---|---|
| Every path returns HTTP 200 | Hundreds of false-positive "found" paths | Recognizes a uniform gate in a few requests; changes strategy |
| 503 on malformed input | "Possible DoS" finding for a human to triage | Measures cross-user impact; downgrades to a robustness note |
| Login form | Maybe a rate-limit warning after a few probes | 62,920-attempt proof with sustained throughput numbers |
| Certificate Transparency logs | Not consulted | Origin hostnames discovered; direct-access bypass tested and ruled out |
| Marketing site | Out of scope | Mined for founder names, city, community → targeted wordlist |
| Failed login page | "Login form found" | Byte-identical oracle analysis; constant-time verification; XSS ruled out |
| Coverage in 30 minutes | One payload list, fired once | Dozens of technique classes, each executed to a conclusion |
| End of run | List of payloads sent | List of falsified hypotheses, persisted for the next run |
Takeaways
The headline of this session is not that an agent broke the gate — it didn't, and the gate deserved to hold. The headline is what "tested" means when testers are armed with AI.
A strong human tester does excellent work on the handful of techniques their schedule allows. A tester with agents gets the whole playbook — every injection class, every protocol attack, every OSINT pivot, every side channel — on every surface, on every run, with every negative documented. That changes the quality of the pentest itself: more techniques attempted means more chances to hit the one misconfiguration that matters; quantified evidence replaces "we think it's fine"; and falsified hypotheses make the clean parts of the report trustworthy instead of merely untested. And when more coverage is needed, the answer is no longer more days or more headcount — it is a larger token budget.
Scanners ask whether a response matches a known-bad pattern. On this target, the two most valuable responses were the ones that matched nothing — and the ones that matched everything.
The two most useful outputs of this session were a quantified absence (no rate limiting, proven at 297 requests/second) and a measured non-event (a crash that affects no one). Neither fits in a signature.
Final tally: 4,300 unique attack probes, ~105,000 credential attempts under an explicit risk profile, one accepted finding (medium), one crash bug responsibly downgraded after impact testing, and dozens of attack vectors falsified and documented — in about half an hour, against a single login page that a scanner would either drown in false positives or wave through as "form found."
That is the promise of AI in offensive security: not replacing the tester's judgment, but removing the ceiling on how much of the attack surface that judgment reaches. Testers test far more than they ever could before — and the pentest gets better.
What happens at a thousand agents?
Everything in this post was one agent, one target, one half-hour window.
Now send thousands of agents against hundreds of different API endpoints. Every endpoint gets the same treatment — 4,300 unique probes, the full technique playbook, the OSINT pivots, the validated findings — on every deploy, without scheduling, without fatigue. And because the agents share a knowledge base, coverage compounds: a vector one agent falsifies on Monday is never retried on Tuesday, and an oddity noticed on one endpoint becomes the hypothesis another agent tests on fifty more.
That is the swarm model of offensive security — not one superhuman tester, but a coordinated fleet with shared memory. The bottleneck is no longer human hours, and the question is no longer "can we afford to test this deeply?"
One agent, one gate, 30 minutes: 4,300 probes and one finding that mattered. Thousands of agents across your entire API estate — and for the first time, "we've tested everything" stops being a figure of speech.