SOC 2 Type II: How to Decide What Technical Testing Your Controls Actually Need

SOC 2 Type II: How to Decide What Technical Testing Your Controls Actually Need
You are a few months into a SOC 2 Type II window, the budget is finite, and someone has told you that you need a penetration test. The reasonable questions follow immediately. Does a web and API test satisfy this? Do you also need the network? Is a scan enough? And who decides?
The answer starts somewhere other than the price list. It starts with what your system description claims and which controls you are asserting, because those are what the testing has to exercise.
What the criteria actually say
SOC 2 runs on the AICPA's Trust Services Criteria, currently the 2017 edition with revised points of focus from 2022. The relevant criterion is less prescriptive than most people expect.
CC4.1 says the entity "selects, develops, and performs ongoing and/or separate evaluations to ascertain whether the components of internal control are present and functioning."
Evaluations, ongoing or separate, that establish whether your controls are there and working. The criterion does not name a method.
Penetration testing does appear in the document, exactly once, in a point of focus under that same criterion, listed as one type of evaluation management may use. Points of focus are illustrative rather than mandatory, and the criteria say so directly: "use of the trust services criteria does not require an assessment of whether each point of focus is addressed", and "some points of focus may not be suitable or relevant to the entity or to the engagement to be performed."
So the honest position is that SOC 2 does not require a penetration test. It requires that you evaluate whether your controls work, and it leaves the method to you and your auditor.
Why your auditor may still ask for one
None of the above means an auditor asking for a penetration test is overreaching.
An auditor has to form an opinion on whether the controls in your system description operated effectively over the period. That requires evidence, and the evidence has to be credible to a reader who was not there. A penetration test is one well-understood form of independent technical evidence, and it speaks directly to whether technical controls hold up under pressure. CC4.1 requires that you evaluate your controls. It does not require that you evaluate them this way. An auditor asking for a penetration test is asking for a form of evidence they can read with confidence, which is a different thing from the criteria demanding it.
Customer contracts are the other driver. Plenty of enterprise agreements name annual penetration testing regardless of what the criteria say, and that obligation is just as real as the audit.
Three parties shape what you end up buying. The criteria set the objective, your auditor judges whether your evidence meets it, and your environment determines what evidence would be convincing in the first place. Ask your auditor what they expect to see before you scope anything. They will tell you, and it costs less than a rescope.
Where testing fits against readiness work
First-cycle teams often run SOC 2 readiness work and testing in parallel and then discover the two were scoped against different assumptions.
Readiness work, sometimes sold as a gap assessment, answers whether you have the controls and the documentation to support them. It helps define and document what management intends to include, since the control set and the description of the system are management's to own. Testing evidences that the technical controls hold up.
The ordering matters more than it sounds. The control list is the input to scoping a test, so a test scoped before readiness finishes is scoped against a list that may still change. Running testing first is not a mistake, and sometimes the calendar forces it, but budget for the possibility that the scope moves once management finalizes the control list.
If you are early enough that the description of the system is still in draft, that is the cheapest moment to ask what testing it will commit you to.

Start with the environment and the controls
Testing should exercise the controls you are claiming, over the system you described. Working outward from those two documents rather than from a package is the step that scoping conversations tend to skip.
Some rough correspondences, offered as starting points rather than a formula:
If your description covers a multi-tenant SaaS product, web application and API testing is usually where the scope sits, with tenant isolation named explicitly in it
If you assert controls over cloud configuration, identity or key management, a Cloud Configuration Review evidences those in a way an app test does not
If your description brings a corporate network, VPN or on-premise systems into scope, external and internal network penetration testing answer different questions. ENPT covers what is reachable from outside, INPT covers what someone already inside can reach, and which one your controls call for depends on where you have asserted them
If your program includes workforce security awareness or access request controls, social engineering is one way to exercise them. Awareness controls are commonly evidenced through training records and policy instead, so treat it as an option to discuss rather than something the control automatically requires
If almost nothing you run is internet-facing and the risk sits internally, an external-only engagement can look thorough while leaving the part that matters untested
Against any proposed scope, take your list of asserted controls and ask which ones the testing would exercise. Controls with no corresponding evidence are the gap to close first. Testing that does not map to an asserted control is not automatically waste, since risk, contractual commitments and architecture changes all justify testing the audit never asks about. It just should not crowd out coverage of a control you have asserted.
Where teams overbuy and underbuy
Overbuying usually looks like a full external, internal and social engineering package bought in the first cycle by a company whose system description covers one SaaS application and its supporting cloud account. The extra coverage is not wasted in a security sense, but it is not evidencing anything the audit asked about, and it consumes budget in a cycle where audit-relevant coverage is what the calendar is paying for.
Underbuying tends to cost more. It usually looks like testing the application while the system description also claims controls over cloud configuration and internal access. The report comes back clean, the auditor asks how you evaluated the rest, and there is no answer inside the window.
Both come from scoping against a catalogue rather than against the description of the system.
Type I and Type II ask different things of the same test
The two report types put different weight on the same engagement, and it changes what you should buy.
A Type I report addresses whether controls are suitably designed at a point in time. A test dated near that point evidences that a control exists and works as designed.
A Type II report addresses operating effectiveness across a period, often three to twelve months. The auditor forms a view about the whole window rather than a single day inside it.
That does not mean the penetration test has to be repeated through the window. A test is one evidence source among many, and the ongoing part of CC4.1 is generally carried by the controls and monitoring you operate continuously, with the test sitting alongside them rather than standing in for them. What tends to matter is that findings were addressed and that you can show it. How your particular auditor weighs any of this is theirs to say, so ask rather than assume.
What should come back
Whatever you scope, the deliverable has to work as audit evidence, which is a higher bar than working as a security document. Look for:
Findings mapped to the relevant Trust Services Criteria, when requested or included in scope, so your auditor is not doing the translation
A written scope that names what was tested and, just as importantly, what was not
The methodology, stated clearly enough that a reader can tell manual testing from automated output
Reproduction detail and evidence for each finding, with severity and the reasoning behind it
Remediation status, and whether remediation testing is included and inside what window
Who performed the work, since independence from the team that built the system is often part of what makes a report persuasive
Third-party independence can strengthen the assurance a report carries. SOC 2 does not universally require an independent penetration test, so treat this as something to raise with your auditor rather than a rule to design around.
The document your customers will ask for
Many customer reviewers ask for a letter, and scopes tend to leave it out.
Your full technical report is rarely the right thing to hand a prospect's security reviewer. It contains findings, sometimes unremediated ones, and it is longer than anyone outside your team wants to read. Many reviewers will accept a short attestation or summary letter naming who performed the testing, when, against what scope, using what methodology, and where remediation stands.
That letter is also what tends to unblock a deal while your SOC 2 report is still months away. Ask whether it is included before you sign, because retrofitting one after an engagement closes is slower than producing it as part of the work.
It does not substitute for the SOC 2 report, and your auditor may request whatever supporting evidence they need, up to and including the full report. The letter exists so you have something shareable in the meantime.
Platform-based and human-led testing
Buyers in a first SOC 2 cycle usually meet both models and are told the other one is inadequate. Neither framing is much use, so here are the trade-offs as they fall.
Continuous platforms are good at cadence and breadth. They run often, they cover a lot of surface cheaply, and they produce the sort of ongoing evidence that CC4.1's "ongoing evaluations" language sits comfortably with. What they tend to be weaker at is anything that depends on knowing what your application is supposed to do: authorization logic, tenant boundaries, multi-step business flows.
Human-led testing goes deeper on exactly those, and produces the reproduction detail and reasoning that reads well as audit evidence. It costs more per engagement and happens less often.
Some programs run both, and the criteria's own wording about ongoing and separate evaluations accommodates that combination comfortably. If a provider tells you their model is the only acceptable one for SOC 2, that claim is not coming from the criteria.
Scope, stated plainly
A penetration test is one piece of SOC 2 evidence. It is not the audit, and it does not make you compliant on its own.
Red Sentry performs the testing and produces the evidence. We are not your auditor, we do not issue SOC 2 reports, and we do not determine whether you pass. The firm doing your examination does that, and they should be the ones telling you what they expect to see.
Our SOC 2 security testing page sets out what comes back in the report and how scope gets locked before testing starts. If you are still deciding between Type I and Type II, we have a separate piece on choosing between them.
Book a scoping call and bring your system description and your control list, even in draft. We will tell you what the testing would cover, what it would not, and how that scope lines up with the controls you are asserting.