Skip to main content
B-52 · Agentic Penetration Testing Platform

Autonomous penetration testing, and exactly where it stops

Autonomy is a boundary, and a boundary is worth nothing unless you can say where it runs. On a fully autonomous engagement the last human action is scope sign-off. Three things after that point stop and wait for you in writing. Everything else, B-52 does on its own.

What follows is that boundary line by line: the three delivery models and where the human sits in each, the three actions held for your approval, the persistence B-52 establishes itself once it has that approval, and what stands behind a finding in each model.

Last human action
Scope sign-off, in the fully autonomous model. Nobody acts again between that signature and the report.
Held for written approval
Three actions: production impact, movement beyond the entry host, and anything touching live credentials or real customer data.
Performed by the platform
Persistence, once it is approved for a scope. B-52 establishes it itself; no human operator is handed the keyboard for that part.
Proof on every finding
A reproducible exploit artefact in all three delivery models. The fully autonomous model adds an automated cross-check before reporting.

Where the human sits

Three delivery models, and one of them has nobody in it after sign-off

Where the human sits is the only axis these three differ on. All eleven coverage classes are tested in every one of them, so the choice you are making is about assurance and about who has to sign the output, and never about what gets looked at.

The three delivery models, where the human sits, and what the output can be filed as
Delivery modelWhat the human doesWhat you can file it as
Fully autonomous Somebody authorises the scope and the targets. After that nobody acts: B-52 maps, tests, chains, proves and delivers the report. No empanelled auditor is in the engagement, so the output is not signable for a regulated filing.
Autonomous, expert verified A senior Security Brigade auditor verifies every finding before any of it reaches you. An empanelled auditor is in the engagement, so the output is signable for a regulated filing.
Human led A senior auditor runs the engagement, with B-52 working underneath them. The auditor is running it, so the output is signable for a regulated filing.
Key
  • A Security Brigade auditor is inside the engagement
  • No auditor is inside the engagement after sign-off

If the assessment is going to an Indian regulator, take the expert-verified or human-led model. An empanelled auditor’s involvement is what makes an assessment signable, and CERT-In empanelment attaches to Security Brigade the firm rather than to any platform, this one included.

Coverage parity across the three is easiest to check on the class where a machine has the most to prove. On a fully autonomous run B-52 decompiles and analyses the APK or IPA itself, which is the work a mobile tester would otherwise do by hand before the assessment proper starts.

The fully autonomous model

Scope sign-off, then nothing

A run with nobody in it is only as good as the practice it learned from. The B-52 harness reproduces how Security Brigade’s own senior auditors work — mindmap creation, test-case generation, comprehensive JavaScript analysis and functional flow analysis — and it was built on 6,700+ assessments of test cases, vulnerabilities and threat models held in Lemon, our own assessment platform. Security Brigade has been CERT-In empanelled since 2008, and that work, going back to the firm’s start in 2006, is what sits underneath a run nobody is watching.

What sign-off actually commits you to

In the other two models a scope document is the start of a conversation that continues through the engagement. Here it is the whole conversation. Everything the platform is permitted to do is in it, and everything it is not permitted to do is either outside it or behind one of the three gates below.

That puts real weight on the document, and it is the honest cost of this model. It is also why the gates exist as a separate mechanism: the two decisions a scope cannot sensibly carry — impact on production and reach past the first host — are taken out of it and asked for on their own.

Turnaround on the fully autonomous model
  • Median one to three business days from scope sign-off to report delivery.
  • That is an observed median across engagements, and it is published as one.
  • It is not a service level, and nothing in the pricing ladder converts it into one.

What is held

Three actions need your written approval before B-52 takes them

These three hold in every delivery model, including the one with nobody in it. They are the same three whether an auditor is running the engagement or not, because what they protect is your estate rather than our process.

The three actions held for the customer’s written approval
Held actionWhat it coversWhy it is held at run time rather than settled in the scope
Production impact Destructive or state-changing actions against a production system. A scope can name a production host without anybody having decided that data inside it may be altered. Those are two different permissions, given by two different people more often than not, so they are asked for separately.
Beyond the entry host Persistence, implants and lateral movement past the host B-52 first landed on. The entry host is the one you expected to be reached. What sits behind it is usually discovered during the run, so it could not have been authorised before the run started — and an approval written in advance for something nobody had seen yet is not an approval.
Live data Anything that touches live credentials or real customer data. A test account is your decision to make in advance. A real one belongs to somebody who is not in the conversation, and reaching it changes who is exposed by the test rather than only what the test finds.

The three actions held for the customer’s written approval

Production impact

What it covers
Destructive or state-changing actions against a production system.
Why it is held at run time rather than settled in the scope
A scope can name a production host without anybody having decided that data inside it may be altered. Those are two different permissions, given by two different people more often than not, so they are asked for separately.

Beyond the entry host

What it covers
Persistence, implants and lateral movement past the host B-52 first landed on.
Why it is held at run time rather than settled in the scope
The entry host is the one you expected to be reached. What sits behind it is usually discovered during the run, so it could not have been authorised before the run started — and an approval written in advance for something nobody had seen yet is not an approval.

Live data

What it covers
Anything that touches live credentials or real customer data.
Why it is held at run time rather than settled in the scope
A test account is your decision to make in advance. A real one belongs to somebody who is not in the conversation, and reaching it changes who is exposed by the test rather than only what the test finds.
After approval

Persistence is the platform’s own work

What you approve
Persistence, implants and movement beyond the entry host, for a named scope, in writing. The approval is for a boundary, not for a single action inside it.
Who performs it
B-52. With the approval in place it establishes persistence itself, inside the authorised scope, as part of the same run that found the way in.
Where a person would normally be
Nowhere in this step. On a conventional engagement the platform reaches a foothold and a human operator takes the keyboard for what follows; here the platform carries on, which is why the approval is written before the run rather than requested during it.

Why that is the harder half of the claim

Reaching a foothold without a person is one problem. Deciding what to do with it — what to leave behind, which host to move to, when the chain has proved enough and can stop — is a different one, and it is the half most autonomous testing hands back to an operator.

B-52 keeps it. That is precisely why this action is gated rather than scoped: an authorisation that a machine will act on unsupervised has to be given deliberately, for a named boundary, in writing, before the run starts.

What stands behind a finding

Every finding carries a reproducible exploit, in all three models

The request, the response, and the steps that reproduce it. That holds for the human-led model, for the expert-verified model, and for the fully autonomous one, with no exception and no reduced form on the model that has no person in it. A finding nobody on your side can reproduce is an argument, and it is the wrong kind of argument to hand an engineering team.

In the fully autonomous model a finding then passes an independent automated cross-check before it is reported. Both words in that phrase are load-bearing and both are narrower than they look. Independent means it is not the component that produced the finding. Automated means no person is involved in it. What makes the finding reproducible is the exploit artefact, and the artefact is there in every model — so the cross-check is a second confirmation sitting on top of proof that already exists, rather than a stand-in for one.

The model in which a senior auditor reads every finding before you do is the expert-verified one. It is one of the three, it is chosen at scoping, and it is what an empanelled auditor signs.

What sits between a finding and your inbox, per model
  • Human led: a senior auditor runs the engagement, and the finding carries its reproducible exploit artefact.
  • Autonomous, expert verified: the same artefact, and a senior auditor verifies every finding before any of it reaches you.
  • Fully autonomous: the same artefact, and an independent automated cross-check before the finding is reported.
  • The artefact is the same object in all three — the request, the response, and the steps to reproduce it.

Deliberately excluded

  • A person reading the finding, in the fully autonomous model. The cross-check runs without one, and the model that puts an auditor on every finding is the expert-verified model.

Measured against the people it learned from

B-52 was run in parallel with Security Brigade’s own expert assessment team, on the same targets, and what each of them found was counted. B-52 reached 90–95% of the combined findings set — everything either party found, counted once — and it surfaced issues the human team did not.

Comparable coverage, different blind spots. Read plainly that is the case for the fully autonomous model and the case for the expert-verified one at the same time: the platform and the auditor together find more than either of them finds alone.

The benchmark in full

What the boundary leaves behind

Four written records of what B-52 was permitted to do

At most four written records define what B-52 was permitted to do on an engagement. Three of them exist only if the corresponding gate was released, so what was refused is as legible afterwards as what was allowed.

  1. 01 The scope authorisation

    What was in scope, what was outside it, and who signed it off. It is the document every action is checked against, and in the fully autonomous model it is the only human input to the engagement.

  2. 02 The production-impact approval

    Present only if a destructive or state-changing action against a production system was permitted. Absent otherwise, and the absence is readable.

  3. 03 The movement approval

    Present only if persistence, implants or movement past the entry host was permitted, and scoped to the hosts it names.

  4. 04 The live-data approval

    Present only if the run was allowed to touch live credentials or real customer data.

Each finding carries its own record of what was sent and what came back, which is the same artefact your engineers reproduce from. What the platform did to produce a given finding is therefore read off the finding rather than reconstructed from a log afterwards.

The same boundary in all three delivery models

All three run against the same scope authorisation, the same three gates and the same reproducible exploit behind every finding. What changes between them is who reads that finding before you do. Choose it at scoping, and change it between engagements when what you need the output for changes.