Pentest Scoping

Black box, gray box, or white box? How to scope a penetration test

More than one client has asked for black box testing believing it's the cheap, no-strings option. It's usually the opposite. Here's how the three testing models actually differ, what each one costs for the coverage you get, and the specific situations where black box is genuinely the right call.

The three models

How much the tester knows going in

The difference between the three models is simple: how much information about the target the tester is given before testing starts. Everything else, including cost and coverage, follows from that one decision.

White box

Full access: source code, schematics, architecture and design documentation, keys and credentials, and often debug access. Almost all engagement time goes to deep analysis instead of discovery.

Black box

No inside information. The tester starts exactly where an external, uninformed attacker would: probing an interface with nothing but what's observable from outside.

Side by side

What each model actually buys you

"Coverage" here means the depth and completeness of what gets tested, not just hours spent. Two engagements can cost the same and deliver very different coverage depending on how they're scoped.

Comparison of white box, gray box, and black box penetration testing
Dimension White box Gray box Black box
Time spent on recon vs. research Minimal recon; nearly all time is research Some targeted recon; most time is research Large share on recon and rediscovery
Relative cost for equivalent coverage Lowest Moderate Highest
Fits a standard length engagement Yes, with full-depth results Yes, with solid depth Only with reduced depth, or a longer window
Best-fit scenario Pre-production design validation; source-available components Most production hardware, firmware, and network engagements Third-party components; narrow perimeter scope; detection & response validation
The cost myth

Why black box costs more, not less

An engagement's budget is fixed hours. Every hour a black box tester spends re-discovering your architecture, mapping your protocols, or reverse engineering your firmware from scratch is an hour that a gray or white box tester spent instead on actual vulnerability research, because they already had that information on day one.

For the same fee, black box buys you shallower coverage. For the same coverage, black box costs more. That tradeoff is the same regardless of how it's marketed, and it's worth being explicit about it before scoping, not after the report comes back thinner than expected.

Vulnerabilities don't care

They exist whether they're found or not

A vulnerability in your system is there regardless of whether a time-boxed engagement happened to uncover it. Handing a tester more information up front doesn't make the system less secure; it means the fixed budget you're paying for goes toward research a skilled person can do that a scanner or a script can't, instead of toward manual reconnaissance that a script largely could have done anyway.

The goal of a pentest is to spend the limited time on the questions only a researcher can answer. Black box scoping routinely spends that time on questions a document handoff could have answered for free.

A common request

"We want you to attack us black box, like a real attacker would"

This comes up often, and the instinct behind it is reasonable. But a penetration test can't actually replicate a real attacker's position, for reasons that have nothing to do with how much information you hand over.

Time

A pentest is scoped to a few weeks. A real attacker going after a sufficiently valuable target can spend months, and on high-value targets, years. No amount of black box penetration testing skill and sophistication can close that gap.

Rules of engagement

A real attacker isn't bound by a signed scope document. They can spearphish your engineers and program managers for documentation, firmware, and source, or target a trusted manufacturing or supply chain partner for the same material. A tester, by design, can't.

Destructive, slow techniques

Attacks like side-channel and fault-injection analysis take real attackers as long as they need, and the setup phase is very likely to destroy hardware along the way. That requires more test units than a typical engagement has budget or lead time for, especially on expensive hardware or before production has ramped.

What black box actually simulates

Black box testing recreates an attacker's starting information, not their time budget, their willingness to go around the target entirely, or their disregard for rules of engagement. It's a narrower simulation than it's usually sold as.

A hybrid approach

Adjust scope as the engagement goes, not just at the start

Scoping doesn't have to be a single, fixed decision made before day one. One effective approach is to start an engagement black box or dark-gray box, then progressively hand over specific items, such as protocol documentation, debug access, or even source, once the tester hits a wall that a sufficiently motivated real attacker would eventually get through anyway, whether by brute-force effort, reverse engineering, or simply attacking a softer target like an engineer's inbox or a supply chain partner.

  • It answers the question clients actually want answered. "What walls would an attacker really run into" is answered more honestly by a wall a real attacker could eventually get past than by an artificial one imposed by a two-week calendar.
  • It keeps most of gray and white box's efficiency. Access is unlocked only when it's needed, not handed over by default, so most of the engagement's budget still goes toward research instead of rediscovery.
  • The wall itself becomes a finding. How long a barrier held up, and what it would have taken a determined attacker to get past it, is useful information on its own, even after the tester is let through it to keep testing what's behind it.
A better mental model

Borrow "assume compromise" from network security

Assumed-breach testing has been standard practice in network and cloud penetration testing for years: instead of spending the whole engagement trying to achieve an initial foothold, you start from the assumption that a foothold already exists, and spend the budget on what matters more, what happens next. Product and embedded security testing should adopt the same posture.

  • "Our firmware is compromised. What can an attacker reach with it?" Gray and white box testing answer this directly, without spending the engagement's budget actually extracting the firmware first.
  • "An attacker recovers keys from our HSM. What's the blast radius?" Providing the keys lets a tester map the actual downstream impact instead of burning the engagement on a key-extraction attempt that a well-funded attacker, unlike the tester, has the time and rules of engagement to eventually pull off anyway.
  • It's honest about the tester's constraints. A tester's rules of engagement rule out some feasible attacks a real adversary would try. Assuming compromise upfront acknowledges that gap instead of pretending a time-boxed black box test closes it.
To be clear

There are times and places for black box testing

None of this means black box testing is never the right call. It's the right call in a few specific situations, and the honest version of each one still accounts for its cost.

Third-party components

Evaluating a vendor or third-party module you don't own and can't get internal documentation for. The external perspective isn't a choice, it's the only information anyone has.

Narrow, perimeter-only scope

When the real question is "can an outside attacker get past this specific interface at all," not full internal coverage of everything behind it, black box is a proportionate answer to a narrow question.

Detection & response validation

Closer to a red-team exercise than a coverage-driven pentest: the goal is testing whether monitoring and your SOC notice an intrusion in progress, not finding every flaw in the target.

The cost tradeoff doesn't go away. Even in these cases, black box still delivers less genuine research per dollar than gray or white box would for the same target. Scoping it for the right reason, with that tradeoff explicit, is what makes it a good decision rather than a default one.

White box doesn't mean skip the hardware checks

Handing over firmware for review is convenient, but it doesn't confirm that a device with physical access is actually blocked from pulling that same firmware off the chip. A pentest should still verify that flash and debug readout protections, chip-level read-out protection, locked or disabled debug interfaces, are enabled and effective on production hardware, regardless of whether firmware was supplied for the engagement. Providing the source doesn't substitute for testing whether it's supposed to be reachable in the first place.

Questions clients actually ask

FAQ

Is black box penetration testing more expensive than white box or gray box testing?

For the same depth of coverage, yes. A black box tester spends a large share of a fixed engagement budget rediscovering information the client already has internally, such as architecture, protocols, and interfaces, before any genuine vulnerability research can start. White and gray box testing skip that rediscovery, so more of the same budget goes toward actual research. Black box testing isn't cheap because there's less to give the tester; it's more expensive per unit of real coverage because more of the clock is spent on reconnaissance instead of research.

Doesn't black box testing better simulate what a real attacker would do?

Not as well as it sounds like it should. A penetration test is constrained to a fixed window, typically a few weeks, while a real attacker targeting a valuable system can spend months or years on it. Real attackers also aren't bound by rules of engagement: they can spearphish engineering and program-management staff for documentation, firmware, or source code, target a trusted manufacturing or supply-chain partner for the same material, and take the time to run destructive or slow techniques like side-channel analysis. A black box test recreates the attacker's starting position, but not the attacker's time budget or their willingness to go around the target entirely.

When should I actually choose black box testing?

Three cases are the clearest fit: testing a third-party or vendor component you don't own and can't get internal information about; a narrow, perimeter-only scope where the real question is whether an outsider can get past a specific interface at all, not full internal coverage; and validating detection and response, where the goal is to see whether your monitoring and SOC notice an intrusion in progress rather than to find every possible flaw. Even in these cases, black box still costs more per hour of genuine research than gray or white box would, so it should be scoped with that tradeoff explicit.

What's the difference between gray box and white box penetration testing?

White box testing gives the tester full access: source code, schematics, architecture and design documentation, keys or credentials, and often debug access. Gray box testing gives partial information, commonly firmware images and interface or protocol documentation without full source, roughly matching what a moderately resourced real attacker could plausibly obtain on their own. Gray box is the default for most production hardware and firmware engagements because it mirrors realistic attacker knowledge without spending the engagement's budget on rediscovery.

If I give the tester my firmware, do I still need hardware-level testing?

Yes. Providing firmware for review doesn't confirm that an attacker with physical access to a production unit is actually blocked from extracting it. A penetration test should still verify that flash and debug readout protections, such as chip-level read-out protection and disabled or locked debug interfaces, are enabled and effective on production hardware, independent of whether firmware was supplied for the engagement.

What does "assume compromise" mean for product security testing?

It's the same posture network penetration testing adopted years ago with assumed-breach engagements: instead of spending the whole budget trying to achieve an initial compromise, you start from the assumption that some component is already compromised and test what happens next. Applied to embedded and cyber-physical products, gray and white box testing let you directly answer questions like "our firmware leaks, what can an attacker reach with it" or "an attacker recovers our HSM keys, what's the blast radius," without spending the engagement's time and budget actually achieving that initial compromise first.

Let's talk

Not sure which model fits your system?

Scoping is a conversation, not a form. Tell me what you're building and what you need to find out, and I'll recommend the model that actually gets you there, not just the one that sounds the most dramatic.