Skip to main content
B-52 · Agentic Penetration Testing Platform

One platform carries out the assessment

B-52 maps the target, generates the test cases that target needs, works them through, proves what it finds, writes the report, and then puts that report through a quality gate that can send it back.

It can do that because of what it was built from. Security Brigade audits going back to 2006 sit inside it, and the harness runs the same practices our senior auditors run.

Coverage
Eleven classes, from web and mobile applications through to Active Directory, source code and AI applications.
Delivery
Fully autonomous, autonomous with expert verification, or human led. The models differ by where the human sits, never by what is tested.
Deployment
On-premise or inside your own VPC. Nothing is required inside your network for external and application testing; a deployment is needed only for internal testing.
Pricing
From $500. One scan is one application or one target, and the full ladder is published.

What it was built from

Every competitor can claim autonomy. None of them has our audit history.

An autonomous tester is only as good as the practice it learned from. Ours learned from Security Brigade’s.

The four practices the harness runs

These are not stages of a pipeline. They are the four things a Security Brigade auditor does on an application assessment, and each one produces something the next depends on.

The auditor practices the B-52 harness reproduces
Auditor practiceWhat it produces on a run
Mindmap creation A map of the target as it actually is: entry points, roles, the paths between them, and the places where one role can reach another role’s data.
Test-case generation The checks this particular application needs, derived from what the application does, rather than a generic list applied to whatever is in front of it.
Comprehensive JavaScript analysis The client-side routes, parameters and endpoints the interface itself never links to. Anything reachable but unlinked is where an assessment that only follows links stops.
Functional flow analysis The business logic: multi-step flows, the state they carry, and what happens when the steps are taken out of order or by the wrong role.

The auditor practices the B-52 harness reproduces

Mindmap creation

What it produces on a run
A map of the target as it actually is: entry points, roles, the paths between them, and the places where one role can reach another role’s data.

Test-case generation

What it produces on a run
The checks this particular application needs, derived from what the application does, rather than a generic list applied to whatever is in front of it.

Comprehensive JavaScript analysis

What it produces on a run
The client-side routes, parameters and endpoints the interface itself never links to. Anything reachable but unlinked is where an assessment that only follows links stops.

Functional flow analysis

What it produces on a run
The business logic: multi-step flows, the state they carry, and what happens when the steps are taken out of order or by the wrong role.

What happens on a run

Six phases, and the sixth is QA

Discovery, planning, scanning, exploitation, reporting, QA. Read the middle two carefully: they are separate phases. Scanning produces candidates and exploitation either proves one or drops it, so nothing that reaches the report arrived on a scanner’s say-so.

The six phases, with what enters and what leaves each one
PhaseWhat enters itWhat leaves it
Discovery The scope you authorised, and nothing outside it. The surface as it stands today: hosts, applications, endpoints and the roles that reach them.
Planning That surface. A mindmap of the target and the generated test cases this application needs.
Scanning The planned test cases. Candidates. Nothing that leaves this phase is a finding yet.
Exploitation Those candidates. Proof, or the candidate is dropped. What survives carries the request, the response and the steps that reproduce it.
Reporting The proven findings. The report, and the same findings in your dashboard with their evidence attached.
QA The finished report. A report that stands, or findings sent back to the phase that produced them.
QA is the sixth phase and it sits after reporting. A quality gate placed mid-pipeline can only stop work that is still in progress; one placed after reporting can return a report that has already been written.

The six phases, with what enters and what leaves each one

Discovery

What enters it
The scope you authorised, and nothing outside it.
What leaves it
The surface as it stands today: hosts, applications, endpoints and the roles that reach them.

Planning

What enters it
That surface.
What leaves it
A mindmap of the target and the generated test cases this application needs.

Scanning

What enters it
The planned test cases.
What leaves it
Candidates. Nothing that leaves this phase is a finding yet.

Exploitation

What enters it
Those candidates.
What leaves it
Proof, or the candidate is dropped. What survives carries the request, the response and the steps that reproduce it.

Reporting

What enters it
The proven findings.
What leaves it
The report, and the same findings in your dashboard with their evidence attached.

QA

What enters it
The finished report.
What leaves it
A report that stands, or findings sent back to the phase that produced them.

QA is the sixth phase and it sits after reporting. A quality gate placed mid-pipeline can only stop work that is still in progress; one placed after reporting can return a report that has already been written.

Putting QA last is the unusual choice and it is the deliberate one. A gate that runs mid-pipeline can only hold work back; a gate that runs after reporting can pull the finished report apart and send the parts back to the phase that produced them.

How each phase changes per coverage class

Where the human is

Scope sign-off, then nothing — or an auditor, if that is what you have to file

Three delivery models. They differ only by where the human sits. All three test the same eleven classes, so the choice is about the assurance you need and who has to sign it, and never about what gets looked at.

The three delivery models and their regulatory consequence
Delivery modelWhere the human isWhat you can file it as
Fully autonomous A person authorises the scope and the targets. After that, nobody acts. B-52 tests, chains, validates and delivers the report. Not signable under Security Brigade’s CERT-In empanelment.
Autonomous, expert verified A senior auditor verifies every finding before anything reaches you. Signable under Security Brigade’s CERT-In empanelment.
Human led A senior auditor runs the engagement with B-52 underneath. Signable under Security Brigade’s CERT-In empanelment.
Key
  • An empanelled auditor is in the engagement, so the output is signable for a regulated filing
  • No empanelled auditor in the engagement

If the assessment is going to a regulator, take the expert-verified or human-led model. The empanelled auditor’s involvement is what makes the output signable, and CERT-In empanelment is a condition of the testing itself, not only of the vendor who sells it. That is worth settling before you scope.

Authorisation

Three actions need your written approval before B-52 takes them

Production impact
Destructive or state-changing actions against a production system.
Beyond the entry host
Persistence, implants and lateral movement past the host B-52 first landed on.
Live data
Anything that touches live credentials or real customer data.

Persistence is the platform’s own work

Once persistence is approved for a scope, B-52 establishes it. There is no point in the run where a human operator takes over to do the part that needs a person, because that part is in the platform too.

Everything above the three gates runs without asking. Everything at them stops until you have said yes in writing.

The autonomy boundary in full

What comes out

Every finding arrives with the exploit that produced it

The request, the response, and the steps that reproduce it. In all three delivery models, without exception — a finding your engineers cannot reproduce still costs them a week to chase.

In the fully autonomous model, findings additionally pass an independent automated cross-check before they are reported. The exploit is what makes a finding verifiable; the cross-check is a second gate sitting on top of it.

We measured it against our own auditors

The benchmark that mattered to us was not a public leaderboard. It was running B-52 in parallel with the Security Brigade assessors it was built from, on the same targets, and counting what each of them found.

Comparable coverage, different blind spots. That is also the honest case for the expert-verified model: the platform and the auditor together find more than either does alone.

How the benchmark was measured

How the parallel benchmark was run
  • B-52 ran in parallel with Security Brigade’s own expert assessment team, on the same targets.
  • The measure is the combined findings set: everything either party found, counted once.
  • B-52 reached 90–95% of that combined set.
  • It also surfaced issues the human team did not find, which is why the combined set is larger than either party’s own.

Where you read it

Four screens, and you have a login for all of them

Your team works in the same place the run reports into. There is no separate portal, and no PDF to wait for.

01 Findings
Every finding with its severity and the evidence behind it — the request that triggered it, what came back, and how to run it again.
02 Attack chain
The findings in the order they were chained, so a low-severity item that only matters because of what follows it is read in that position rather than in a severity sort.
03 Test runs
Which runs are under way, what has been covered so far, and what is still outstanding on the current scope.
04 Remediation
Each finding tracked from open through to a retest and a verified fix, rather than closed on an assertion that it was fixed.

Look inside the dashboard

What it can be pointed at

Eleven coverage classes, each with its own page

A list of eleven names proves nothing. Each of these opens onto what B-52 does in that class, which methodology it works to, and what it needs from you to start.

Physical security, hardware and wireless testing are out of scope. That is the only exclusion, and it is stated here so you can price the gap rather than discover it.

One scan is one application or one target

The ladder starts at $500 and every tier on it is published. You can sign up and run the first scan with a card, or bring us a scope and we will size it with you.