Skip to main content
B-52 · Agentic Penetration Testing Platform

Security validation, measured against our own assessors

The benchmark that decided whether B-52 was ready was not a public scoreboard. Security Brigade put it on live engagements alongside the assessment team it was built from, ran both in parallel, and counted what each of them found.

Neither party found everything. That is the result, and it is what the expert-verified delivery model is built on.

Why this measurement

Why we benchmarked against our own assessors

A fixed public target set comes with an answer key, so a score against it reports how much of a known list was recovered. Your application has no answer key. The question that decides a purchase is whether the platform gets to what a senior assessor gets to on a target nobody has solved in advance, and what it gets to that the assessor does not.

That measurement needs a senior assessment team willing to be measured against, working the same engagements at the same time. It is expensive, it is slow, and it is the version of the question a buyer actually has.

Why the comparison was available to us at all

Security Brigade has been CERT-In empanelled since 2008, and every assessment the firm has run since it started in 2006 was worked inside Lemon, our own assessment platform, rather than leaving with the auditor who wrote each finding. 6,700+ assessments of test cases, vulnerabilities and threat models are what our models learned from.

The B-52 harness was then written to work the way those auditors work — mindmap creation, test-case generation, comprehensive JavaScript analysis and functional flow analysis. So the assessment team is not an arbitrary yardstick. It is the practice the platform came from, which is what makes running the two side by side worth the cost.

What the platform was built from

What the parallel run answers that a scored target set cannot
What a buyer is trying to settleWhat the parallel run says
Does it reach what a senior assessor reaches? Both worked the same engagements over the same period. The comparison is finding against finding on live applications, rather than a score against a fixed target set whose answers are already published.
Does it reach anything an assessor does not? Yes. The combined set is larger than what the assessment team produced on its own, because B-52 surfaced issues they did not.
What does it miss? The part of the combined set it did not reach. That part came from the assessors, and it is why a senior auditor sits inside two of the three delivery models.

What the parallel run answers that a scored target set cannot

Does it reach what a senior assessor reaches?

What the parallel run says
Both worked the same engagements over the same period. The comparison is finding against finding on live applications, rather than a score against a fixed target set whose answers are already published.

Does it reach anything an assessor does not?

What the parallel run says
Yes. The combined set is larger than what the assessment team produced on its own, because B-52 surfaced issues they did not.

What does it miss?

What the parallel run says
The part of the combined set it did not reach. That part came from the assessors, and it is why a senior auditor sits inside two of the three delivery models.

What the parallel run returned

B-52 reached 90–95% of the combined findings set

The measure is the combined set — everything either party found, each item counted once. B-52 reached 90–95% of it. Some of what it reached was work the assessment team did not produce on that engagement, which is why the combined set is larger than either party’s own output. The rest of the set came from the assessors and not from the platform.

The combined findings set, by which party reached each part
Part of the setWho reached itWhat it tells you
Reached by both The overlap — issues the platform and the assessment team each arrived at independently, working the same target. The same finding, confirmed from two directions.
Reached by B-52 only Issues the assessment team did not surface on that engagement. Counted inside the 90–95%, and the reason the combined set is larger than the team’s own output.
Reached by the assessors only The remainder of the combined set. B-52 did not get to it. The honest argument for putting a senior auditor in the engagement.
Key
  • Reached independently by the platform and by the assessment team
  • Reached by B-52 and not by the assessment team
  • Reached by the assessment team and not by B-52
How the parallel run was set up
  • B-52 and Security Brigade’s expert assessment team worked the same engagements, in parallel, over the same period.
  • The denominator is the combined findings set: everything either party found, each item counted once.
  • B-52 reached 90–95% of that set.
  • Part of what B-52 reached was absent from the assessment team’s own output, so the combined set is larger than what either party produced alone.
  • The remainder of the set came from the assessors. Neither party reached all of it.

What follows from it

The platform and the assessors have blind spots in different places

Read plainly, that is a statement about blind spots — the platform’s and the assessors’ are in different places. It is also why the three delivery models differ by assurance and not by coverage. The choice you are making is how many of those blind spots you want a second party looking for.

What a finding passes through before it reaches you, by delivery model
Delivery modelWhat stands between the finding and your report
Fully autonomous The exploit artefact, then an independent automated cross-check that runs before the finding is reported.
Autonomous, expert verified The exploit artefact, then a senior auditor who verifies every finding before anything reaches you.
Human led The exploit artefact, produced inside an engagement a senior auditor is running with B-52 underneath.

What a finding passes through before it reaches you, by delivery model

Fully autonomous

What stands between the finding and your report
The exploit artefact, then an independent automated cross-check that runs before the finding is reported.

Autonomous, expert verified

What stands between the finding and your report
The exploit artefact, then a senior auditor who verifies every finding before anything reaches you.

Human led

What stands between the finding and your report
The exploit artefact, produced inside an engagement a senior auditor is running with B-52 underneath.

Every model tests the same coverage classes. What moves between them is where the person is, so the choice is about the assurance you need and never about what gets looked at.

The proof unit

A finding arrives as something your engineer can run again

Three things travel with every finding, in all three delivery models: the request that produced the behaviour, the response that came back, and the steps that reproduce it. A result nobody can reproduce is a ticket for your team to argue about.

Evidence

What travels with a finding

The request
The call that produced the behaviour, as B-52 sent it.
The response
What the target sent back. This is the part that makes the behaviour observed rather than asserted.
The steps that reproduce it
The sequence one of your own engineers can run to see the same thing, without us on the call and without taking our word for the first two.

The cross-check sits on top of the artefact

In the fully autonomous model, a finding additionally passes an independent automated cross-check before it is reported. The artefact is what makes a finding checkable; the cross-check is a second gate above it, for the model in which no person reads the run before you do.

The artefact does not thin out in the classes that are not web-shaped. Mobile is the sharpest case of that — B-52 decompiles and analyses the release binary itself, obfuscated and hardened builds included, and a finding there arrives on the same terms as a finding against a login form.

Mobile application testing in full

Where this sits in the platform

The benchmark is the evidence. The run that produces it — six phases ending in a quality gate that sits after reporting, three delivery models, and the boundary the platform works inside — is on the platform page.

The platform in full