Skip to main content
Coverage class · People

Social engineering penetration testing aims at a foothold, not a score

This class is an assessment, not a training exercise. It is built from reconnaissance of your own organisation rather than from a template library, it runs against a population you authorised in writing, and it is judged on one thing: whether it obtained a way in. What comes back is the route, the evidence, and what your own controls did about it.

The boundary

The authorisation is not a gate — it is the condition

Every coverage class has three actions that wait for written approval. This one has something before them: a document that must exist before anything is sent to anybody.

The authorisation is not a gate — it is the condition
StateWhat it meansWhat follows
Sending anything to a real person Nothing reaches anybody until a written authorisation naming the population, the window and the permitted pretexts exists. It cannot be released mid-run, because it is what makes the engagement an engagement. Required before the work begins.
Using what a captured credential opens A credential that was entered is the evidence. Signing in with it is the next act, and it is the third approval gate under a different name. Stops and waits, in writing.
Persistence and movement past the entry host Where a pretext yields access to something, going further into the estate is an internal network question rather than a social engineering one. Stops and waits, in writing.
Everything else inside the authorised scope Terminal Open-source reconnaissance, target enumeration, campaign construction, delivery inside the agreed window, and reporting. Runs without asking.
Key
  • A precondition of the engagement, not something released during it
  • Requires your written approval before B-52 proceeds
  • Authorised by the scope you signed off
  • TerminalNo state follows this one

What the authorisation names

Four things, and none of them can be settled afterwards

The population, by group rather than by whoever reconnaissance happens to surface. The window, so your own team can tell this apart from a real campaign afterwards — and so a genuine incident during it is not filed as ours. The pretexts that are permitted and the ones that are not, which is your decision and one most organisations turn out to have firm views about the moment they are asked. And where the engagement stops: at the credential being entered, or at what the credential opens. In the fully autonomous model that document is the last human action of the engagement, which is precisely why it has to be this specific.

Class and boundary

Three forms, and the one that runs first

Reconnaissance is not preparation for the other two. It is the form that most often produces the finding.

Form one

Reconnaissance and target enumeration

Who is there, what is already published about them, which of their credentials are in circulation somewhere nobody put them deliberately, and which of that makes a campaign credible. It runs first and it decides the other two.

Form two

Targeted phishing simulation

A campaign written for the organisation it is aimed at, out of what reconnaissance actually found — not a template with a different logo dropped into it.

Form three

Pretexting and credential harvesting

A scenario with a plausible reason to exist, run to obtain a credential or an action, and treated as one link in a chain rather than as the result.

Adjacent

What the foothold leads to

Where a pretext or a recovered credential yields access, how far that access then travels is the internal network class. This one ends at the way in; that one measures what the way in was worth.

Not in this class

Voice, physical and hardware

Telephone pretexting sits outside this class. Physical, hardware and wireless testing are out of scope for the platform entirely, in every class.

The other thing this could be

A phishing simulation for training and one for testing are different exercises

A security awareness programme runs phishing simulation to measure and improve how a workforce responds: at a cadence, repeated, with everybody eventually included. That is a real discipline and Security Brigade runs one. This class is doing a different job — it is trying to obtain a foothold, against a named population, once, as part of an assessment, and it succeeds or fails on whether it got in rather than on what proportion of people clicked. The two produce different reports for different readers. Buying either one does not give you the other, and the failure mode of confusing them is a programme that never tests what an attacker would actually do, and an assessment that quietly turns into a staff appraisal.

Who this page is for

Two readers, and both have already spent on the technical half

Both are answered here. Where they should start is not the same.

01 The whole route

The team testing whether the perimeter is the only way in

What brought them
An external assessment that came back clean, and an uncomfortable awareness that the clean result only covers one of the two routes.
What they need
A campaign that ends where a real one would — at what the credential opened. Start at the worked example and at the boundary above it.
What they check first
Whether the engagement stops at the click or continues into what the click yielded.
02 Assurance

The team whose technical controls are already good

What brought them
Multi-factor authentication rolled out, a mail gateway doing its job, and no evidence about what happens when somebody is simply persuaded.
What they need
An answer about the controls rather than about the people — what stood between a captured credential and an account.
What they check first
Whether the report names individuals. It does not, unless you asked for that in the authorisation.

The run

Six phases, worked as a campaign

The phase names are the platform’s. On this class the first one is where most of the value is, and the fourth is the only one anybody outside the engagement ever sees.

On the people in it

The report describes populations, not individuals

An engagement like this generates a list of people who did something, and what happens to that list decides whether running it was worth anything. The report describes the population, the pretext and the outcome; it does not hand a manager a league table of names. Where you need individual detail for a specific reason, that is agreed in the authorisation with everything else and it is your decision rather than a default. The findings worth having are structural in any case: a pretext that worked says something about what your staff have been trained to expect, and a captured credential that opened something says rather more about what a credential is permitted to open.

What the assessment returns

What a campaign produces, beyond who clicked

The first two rows are the ones that survive into a remediation plan. The rest are what explain them.

ClassWhat it looks like in a real engagement
Credentials already in circulation Accounts belonging to the organisation that appear in corpora assembled elsewhere — from unrelated breaches, from software running on somebody’s personal machine, from a repository that was made public by accident. Found before anything is sent, and they cost the campaign nothing.
What a credential opens Whether a recovered or captured credential reaches anything at all, and what stands between it and the account. This is the finding. The credential is only the evidence for it.
The published surface of the organisation’s people Who is identifiable, in which role, with which address format, and what has been published about the processes they take part in. The material any campaign is built from, and all of it already public.
What the pretext relied on Which assumption made the scenario credible — an expected process, a known supplier, a system people already receive mail from. This is the part that turns into a change somebody can actually make.
What the mail controls did Whether the campaign arrived, was quarantined, or was stopped outright. Reported either way, because a campaign that does not arrive is evidence about a control you paid for.
What happened after the interaction Whether the destination was reachable, whether anything warned the person, and whether what they did could be undone.
Whether anybody said something Whether the activity was reported internally at all. That is the half of the exercise that describes the organisation rather than the campaign.

Credentials already in circulation

What it looks like in a real engagement
Accounts belonging to the organisation that appear in corpora assembled elsewhere — from unrelated breaches, from software running on somebody’s personal machine, from a repository that was made public by accident. Found before anything is sent, and they cost the campaign nothing.

What a credential opens

What it looks like in a real engagement
Whether a recovered or captured credential reaches anything at all, and what stands between it and the account. This is the finding. The credential is only the evidence for it.

The published surface of the organisation’s people

What it looks like in a real engagement
Who is identifiable, in which role, with which address format, and what has been published about the processes they take part in. The material any campaign is built from, and all of it already public.

What the pretext relied on

What it looks like in a real engagement
Which assumption made the scenario credible — an expected process, a known supplier, a system people already receive mail from. This is the part that turns into a change somebody can actually make.

What the mail controls did

What it looks like in a real engagement
Whether the campaign arrived, was quarantined, or was stopped outright. Reported either way, because a campaign that does not arrive is evidence about a control you paid for.

What happened after the interaction

What it looks like in a real engagement
Whether the destination was reachable, whether anything warned the person, and whether what they did could be undone.

Whether anybody said something

What it looks like in a real engagement
Whether the activity was reported internally at all. That is the half of the exercise that describes the organisation rather than the campaign.

Three of them, worked

The same engagement, as a headline and as a finding

Each of these reads one way in a summary and another way in a remediation plan.

01 Reconnaissance

A credential nobody in the organisation exposed

As a headline
Staff passwords found online.
As a finding
An account still accepting a password that has been circulating outside the organisation, with nothing else asked for at sign-in.
What changes because of it
A control decision about second factors and about credential reuse — not a training slide.
02 Delivery

A pretext that worked

As a headline
People clicked a link.
As a finding
A scenario that was credible because it matched a process the organisation genuinely runs, arriving from infrastructure nothing flagged.
What changes because of it
Either the process gains a way to be verified, or the mail controls gain something they were missing. Usually the first.
03 Outcome

A campaign that got nowhere

As a headline
Nothing to report.
As a finding
Where it stopped, and what stopped it — the gateway, the second factor, or somebody who raised it.
What changes because of it
Nothing needs to. But it is evidence that a control works, which is worth more written down than assumed.

Methodology

The standards a social engineering engagement is worked against

The last row is applied to the technical findings a campaign produces rather than to the campaign itself, and the row says so rather than pretending a scoring vector fits everything.

StandardVersionWhat it carries here
MITRE ATT&CK v19.2, released 6 August 2026. Read 2026-09-13 The technique vocabulary for reconnaissance, initial access and credential access — so a finding can be handed straight to the team that owns the detection for it.
NIST SP 800-115 Final, September 2008. Read 2026-09-13 The structure of a technical assessment, and the handling rules for what the attack phase turns up.
Still the current edition: it has been neither withdrawn nor superseded.
PTES Current The engagement structure — intelligence gathering, authorisation in writing, and the shape of the report.
CVSS v4.0, November 2023. Read 2026-09-13 The severity vector on the technical findings a campaign produces — what a captured credential reached, and what did not stand in its way.
Not applied to the campaign itself. A pretext is not a vulnerability with a vector, and scoring one as though it were would be inventing precision.

MITRE ATT&CK

Version
v19.2, released 6 August 2026. Read 2026-09-13
What it carries here
The technique vocabulary for reconnaissance, initial access and credential access — so a finding can be handed straight to the team that owns the detection for it.

NIST SP 800-115

Version
Final, September 2008. Read 2026-09-13
What it carries here
The structure of a technical assessment, and the handling rules for what the attack phase turns up.

Still the current edition: it has been neither withdrawn nor superseded.

PTES

Version
Current
What it carries here
The engagement structure — intelligence gathering, authorisation in writing, and the shape of the report.

CVSS

Version
v4.0, November 2023. Read 2026-09-13
What it carries here
The severity vector on the technical findings a campaign produces — what a captured credential reached, and what did not stand in its way.

Not applied to the campaign itself. A pretext is not a vulnerability with a vector, and scoring one as though it were would be inventing precision.

Before a run starts

What is fixed in writing, and what is never in scope

The authorisation for a social engineering engagement As of 2026-09-13
  • The target population, named by group. Reconnaissance will surface people outside it, and those people stay outside it.
  • The window, so that your own team can separate this from a real campaign — and so that a genuine incident during it is not filed as ours.
  • The permitted pretexts, and the ones that are not permitted. That decision belongs to you, and it is worth taking properly.
  • Where the engagement stops: at the credential being entered, or at what the credential opens.
  • Who inside your organisation knows it is happening, which decides whether the exercise also measures your own response.

Deliberately excluded

  • Telephone pretexting, which sits outside this class.
  • Physical, hardware and wireless testing, which are out of scope for the platform in every class.
  • Anybody outside the authorised population, however reachable reconnaissance shows them to be.
  • Anything delivered outside the agreed window.

Per finding

What arrives with every finding

Always

What was sent, and from where

The message as delivered and the infrastructure it came from, so your team can find it in their own logs rather than taking our word for what arrived.

Always

What came back, by population

The interaction and its timestamp, described by group and role. Individual detail appears only where the authorisation asked for it.

Where you authorised it

What the credential reached

What a captured credential authenticated to, and what stood in the way. This is the finding the engagement exists to produce, and it is also the one behind an approval gate.

Always

What your controls did

Whether the campaign arrived at all, what happened to the parts that did not, and whether anybody inside the organisation raised it.

Fully autonomous only

An independent automated cross-check

In the model with no auditor in it, findings pass a second automated gate before they are reported. It is a check on top of the evidence, not a substitute for it.

Where this sits

Against the other things aimed at the same people

No vendor is named here — these are categories of work, and one of them is a thing Security Brigade also sells.

An awareness programmeMail gateway and filteringA scheduled red teamB-52
What counts as success A measurable change in how people respond, over time. A campaign that never arrives. A foothold. A foothold, or a written account of what stopped there being one.
Who it is aimed at Everybody, repeatedly, by design. Everybody, continuously. A chosen few. The population your authorisation names, once.
Built from reconnaissance of your organisation Usually from a template library, which is appropriate for the job it does. Not applicable. Yes. Yes — and the reconnaissance is reported whether or not the campaign ever ran.
Continues past the credential No, and it should not. Not applicable. Yes, within scope. Only where you authorised it, and the report says which findings went that far.
Cadence Continuous, by design. Continuous. When it is scheduled and staffed. The interval you set.
Whose signature it carries Not that kind of output. None. The firm that ran it. Security Brigade’s, in the two models with an empanelled auditor in them.

What counts as success

An awareness programme
A measurable change in how people respond, over time.
Mail gateway and filtering
A campaign that never arrives.
A scheduled red team
A foothold.
B-52
A foothold, or a written account of what stopped there being one.

Who it is aimed at

An awareness programme
Everybody, repeatedly, by design.
Mail gateway and filtering
Everybody, continuously.
A scheduled red team
A chosen few.
B-52
The population your authorisation names, once.

Built from reconnaissance of your organisation

An awareness programme
Usually from a template library, which is appropriate for the job it does.
Mail gateway and filtering
Not applicable.
A scheduled red team
Yes.
B-52
Yes — and the reconnaissance is reported whether or not the campaign ever ran.

Continues past the credential

An awareness programme
No, and it should not.
Mail gateway and filtering
Not applicable.
A scheduled red team
Yes, within scope.
B-52
Only where you authorised it, and the report says which findings went that far.

Cadence

An awareness programme
Continuous, by design.
Mail gateway and filtering
Continuous.
A scheduled red team
When it is scheduled and staffed.
B-52
The interval you set.

Whose signature it carries

An awareness programme
Not that kind of output.
Mail gateway and filtering
None.
A scheduled red team
The firm that ran it.
B-52
Security Brigade’s, in the two models with an empanelled auditor in them.

On the awareness programme

Keep it — it is doing the job this one is not

Running phishing simulation as a training programme should include everybody, should repeat, and should be measured over months. None of that is true of this class, which aims at an authorised population once and is judged on whether it got in. Both are worth having and Security Brigade runs both, on separate terms with separate reports. What neither of them should be is the other: a training programme asked to prove whether an attacker could get in will not, and an assessment asked to produce a departmental score has stopped being an assessment.

The record

What makes this lawful, and what it is evidence for

Coverage is identical across the three delivery models. What differs is whose signature the report carries.

Authorisation

The document that separates this from what it imitates

A social engineering engagement is testing performed on people. The written authorisation naming the population, the window and the permitted pretexts is what makes it an engagement, and it is retained as part of the record rather than treated as paperwork.

Evidence

One record, not three

Where an obligation asks for testing that covers the route into an organisation through its people, this class produces it — the campaign, the outcomes and the authorisation that permitted them, together.

Remediation

The state a finding closes in

A finding here closes on a control change rather than on a patch: a second factor that now stands in the way, a process that can now be verified, a credential that no longer works.

Delivery

Choosing a model for a filing

Where the report goes to a regulator or an assessor, start with the expert-verified model.

Worked example

A chain where the campaign was never needed

Four steps, and the first of them was found in reconnaissance before anything was sent to anybody.

A chain from a social engineering engagement, read step by step and then in sequence
LinkAloneIn sequence
1 A credential already in circulation An address and a password, in a corpus assembled somewhere else.It belongs to somebody inside the authorised population.
2 An account asking for one thing A sign-in page.The credential works, and nothing further is asked for.
3 A mailbox the organisation trusts An inbox.A message sent from it does not look like a campaign, because in every technical sense it is not one.
4 The action somebody took A routine request, actioned.The route in never touched the perimeter, so nothing watching the perimeter saw it.
A chain from a social engineering engagement, read step by step and then in sequence Reconstructed from a real engagement. Sector, organisation and every identifier are generalised; the structure is what carries across. Every step after the first was taken under written approval.

Measured

Benchmarked against our own assessors

B-52 was run beside Security Brigade’s expert assessment team on the same targets at the same time, and everything either of them produced was pooled into a single set with each item counted once. B-52 reached 90–95% of that pooled set. Part of what it reached the team had not, which is why the pool is larger than either side alone — and why the expert-verified model is a genuine option on this class in particular, where judgement about what to send is worth having a person on.

A campaign is scoped around an authorisation, not a card

One scan is one application or target and the entry tier is $500, but this class is scoped around a population, a window and a written authorisation. That is a conversation, and it is a short one.