# 10. Pen Testing, Red Team, Offensive Services
Onward to a topic where vendor quality varies enormously and the buyer often can't tell the difference between elite and mediocre work, which is exactly how mediocre vendors stay in business.
Context: This is one of the most opaque categories in security purchasing. You're buying a service where:
- The deliverable is a report you may not be technically equipped to evaluate
- The "quality" of testing is invisible to most buyers
- The market spans from elite boutique firms charging $50K/week to compliance-mill shops charging $5K for a "pen test" that's actually just a Nessus scan with a logo
- Vendors hide behind NDAs and "methodology" claims
The category itself fragments into:
- Vulnerability Assessment (VA), automated scanning + light validation. Often mislabeled as "pen testing."
- Penetration Testing (Pen Test), manual exploitation against a defined scope, usually time-boxed (1–3 weeks). The compliance staple.
- Red Team Engagement, adversary simulation against a specific objective (e.g., "steal the crown jewels"), often weeks to months, includes physical, social, and cyber.
- Purple Team, collaborative exercise where the offensive team works with the defensive team to improve detections.
- Continuous Pen Testing / PTaaS (Pen Test as a Service), subscription model with platforms like Cobalt, HackerOne, Synack, Bugcrowd, NetSPI.
- Adversary Emulation / BAS (Breach & Attack Simulation), automated tools (AttackIQ, SafeBreach, Picus) that continuously test defenses against known TTPs.
- Bug Bounty Programs, crowdsourced testing via platforms (HackerOne, Bugcrowd, Intigriti).
The market reality: most "pen tests" sold to mid-market companies are actually compliance-driven vulnerability assessments dressed up as pen tests, and most buyers can't tell, because they're buying for the audit checkbox, not for security improvement.
The right question isn't "did you get a pen test?", it's "did the test simulate real threats relevant to my business, expose meaningful risk, and result in measurable improvement?"
14 Questions to Ask Pen Test / Red Team / Offensive Services Vendors
1. "Walk me through the difference between what you're proposing and a vulnerability assessment. What manual exploitation will actually happen?"
Why: Many "pen tests" are 80% scanner output + 20% light validation. A real pen test involves significant manual exploitation, chaining of vulnerabilities, and creative attack paths a scanner can't find.
Good answer: Clear delineation of automated vs. manual work, hours of manual effort committed, examples of manual findings. Red flag: "Our methodology combines automated and manual testing" with no specifics.
2. "Who specifically will be doing the testing, names, certifications, years of experience, and prior similar engagements? Can I interview them before signing?"
Why: The single biggest determinant of pen test quality is who's doing the work. Senior testers find different things than junior ones. Some firms quote senior names in sales calls, then staff with juniors. Get names in writing.
Good answer: Named individuals, willingness to put them on a call, contractual commitment to specific personnel. Red flag: "We have a team of 50+ testers and assign based on availability."
3. "Show me a sanitized example of your previous report, not just the executive summary, but the actual technical findings."
Why: The report is the deliverable. Some vendors produce excellent technical narrative; others produce reformatted scanner output. The quality difference is enormous and visible in 5 minutes of reading.
Good answer: Real (sanitized) reports with detailed reproduction steps, business impact analysis, and remediation guidance. Red flag: "We can't share reports due to NDAs" (then provide a sanitized template).
4. "What's the threat model you're testing against? Are you simulating a specific adversary, or running a generic methodology?"
Why: Generic methodology testing finds generic issues. Threat-informed testing finds what's relevant to your business and your likely adversaries. A healthcare company should be tested against ransomware groups; a defense contractor against APTs. Your industry, threats, and assets matter.
Good answer: Tailored threat model based on your industry and crown jewels, MITRE ATT&CK alignment, specific adversary profiles. Red flag: "Our methodology covers OWASP/PTES/etc." (necessary but not sufficient).
5. "How do you handle scope, and how do you push back when scope is too narrow to find meaningful risk?"
Why: Many pen tests are scoped so tightly (10 IPs, no social engineering, no exploitation chaining, no production systems) that meaningful findings are impossible. Mature vendors will tell you when scope makes the test useless. Immature ones just take the money.
Good answer: Honest pushback on scope limitations, willingness to recommend broader engagement, refusal to do useless work. Red flag: "We'll work within whatever scope you provide."
6. "What's your stance on production testing, social engineering, and physical access? When are these warranted, and what are your safeguards?"
Why: Real attackers don't respect scope. Tests that exclude production, humans, and physical access are missing the most common attack vectors. Mature vendors discuss the trade-offs honestly; immature ones avoid the conversation.
Good answer: Risk-based recommendations, robust safeguards, communication protocols, escalation paths. Red flag: "We only test what you ask us to."
7. "Walk me through your most impressive engagement of the last year. What did you find that nobody expected?"
Why: Top-tier offensive teams have stories. They've found creative paths, chained obscure vulnerabilities, demonstrated impact that surprised the customer. Forces a real-world demonstration of capability beyond methodology slides.
Good answer: A specific (sanitized) story showing creativity, persistence, and technical depth. Red flag: Generic answers about "finding misconfigurations."
8. "How do you handle remediation guidance, high-level recommendations or detailed, actionable, environment-specific advice?"
Why: Findings without actionable remediation are wallpaper. The best pen test reports give engineers exactly what they need to fix the issue. The worst say "implement defense in depth" and leave you guessing.
Good answer: Step-by-step remediation, code samples where relevant, prioritization, post-engagement Q&A access. Red flag: "We provide industry-standard recommendations."
9. "What's your policy on retesting, and is it included or extra?"
Why: Findings need verification of fixes. Some vendors include free retesting within X days; others charge full price. This affects your total cost and your audit story significantly.
Good answer: Defined retesting window included, clear scope of retest, audit-ready before/after evidence. Red flag: "Retesting is a separate engagement."
10. "What's the difference between this pen test and a red team engagement, and which do I actually need?"
Why: Pen tests find vulnerabilities; red team engagements test detection and response. Most companies need both over time, but in different orders. A vendor that doesn't help you understand the difference is selling, not consulting.
Good answer: Honest framework, pen test for technical risk discovery, red team for measuring SOC/program maturity, recommendations based on your maturity. Red flag: Selling whatever they happen to offer.
11. "How do you work with my blue team, adversarial only, or purple team collaborative?"
Why: Pure adversarial testing has its place but often results in defensive teams learning nothing. Purple teaming (offensive working with defenders) generates more actionable improvement. Mature programs use both at different times.
Good answer: Both modes offered, with guidance on when each is appropriate, willingness to debrief defenders deeply. Red flag: "We test silently to be realistic." (Sometimes correct, but not always; depends on goal.)
12. "How do you handle continuous testing vs. point-in-time engagements? What's your view on PTaaS, BAS, and bug bounty?"
Why: The annual pen test is becoming an outdated model. Continuous testing approaches (PTaaS, BAS, bug bounty) provide ongoing validation. A vendor's stance reveals whether they're stuck in the old model or evolving.
Good answer: A nuanced view, point-in-time still has value, but layered with continuous approaches; recommendations based on maturity. Red flag: Defending the annual pen test as sufficient (it isn't, alone).
13. "What credentials, certifications, and accreditations does your firm hold, and which ones actually matter?"
Why: The cert landscape is crowded (OSCP, OSCE, OSEP, GPEN, GXPN, CREST, CESG, etc.). Some are genuine signals of skill (OSCP, CRT, OSCE, GXPN); some are paper certs. CREST/CHECK accreditation matters in UK; PTES alignment is a methodology baseline. A mature vendor explains which matter and why.
Good answer: Honest discussion of which certs reflect skill vs. checkbox, named accreditations relevant to your jurisdiction. Red flag: "We have all the major certifications" without nuance.
14. "How do you protect my data during and after testing, credentials harvested, exploits developed, findings documented?"
Why: During a pen test, the vendor briefly holds the keys to your kingdom. After the engagement, they hold a roadmap of your weaknesses. Their security posture matters as much as their offensive skill. A breach of your pen test firm = a guided tour for attackers.
Good answer: Encrypted handling, defined data retention, secure communication, NDA + insurance, transparency about their own security. Red flag: "We follow industry best practices."
🎯 The Meta-Test
- Q1, Q2, and Q3 are the quality test. A vendor weak on these is selling a compliance product, not a security product.
- Q4 and Q5 are the maturity test. Threat-informed, scope-honest vendors are the ones worth hiring.
- Q14 is the trust test. They're going to know your weaknesses better than you do, make sure they protect that information.
💡 Honest Observations
For "regular" companies, the pen testing market has some uncomfortable truths:
- Most pen tests are bought for compliance, not security. PCI DSS, SOC 2, ISO 27001, HITRUST, these all require pen testing. So companies buy the cheapest one that satisfies the auditor and treat it as a checkbox. The result: useless tests, no real security improvement, and a false sense of security. If you're going to spend the money, spend it on a test that actually finds things.
- The price range tells you everything. A real pen test of a meaningful environment costs $20K–$100K+. If someone's quoting you $5K for a "comprehensive pen test," they're running scanners. There's no shame in buying a vulnerability assessment if that's what you need, but call it what it is.
- The "we'll do everything" pen test firm is usually mediocre at all of it. Elite offensive work is specialized, application security, cloud, OT/ICS, mobile, hardware, social engineering. Boutique firms with deep specialty often outperform full-service generalists for that specific domain.
- Annual pen tests are outdated as a sole strategy. The world has changed; your environment changes weekly. A point-in-time test is a snapshot, not a posture. Modern programs combine: annual pen test (compliance + deep dive) + continuous PTaaS or bug bounty (ongoing coverage) + BAS (control validation) + red team (every 1–2 years for maturity testing). Smaller companies can't afford all of this, but at minimum, pen test annually + bug bounty or PTaaS for ongoing visibility is the modern baseline.
- The single most underused offensive service is purple teaming. A purple team exercise often produces more actionable improvement than three pen tests. Your defenders learn what attacks look like in your environment, and your detection content improves immediately. If you've never done one, do one.
- The remediation matters more than the finding. A pen test report that sits in a drawer is worse than no test (false sense of security + audit trail of known issues you didn't fix = legal liability if breached). Budget time for remediation before you budget the test. If you can't remediate, don't test.
- Beware the "we'll keep finding new issues every year" model. If your annual pen test keeps finding the same categories of issues (or worse, the same specific issues), your security program isn't maturing, your testing is. Either change vendors, change scope, or invest in fixing root causes.
For most "regular" mid-market companies, the practical sequence is:
- Build a baseline first, vulnerability management, EDR, MFA, basic hygiene. Pen testing an immature program just generates a long list of obvious findings.
- Get an annual pen test focused on internet-facing assets and critical applications, from a reputable firm with named senior testers.
- Layer in PTaaS or bug bounty for continuous coverage of your changing attack surface.
- Add purple teaming once you have detection capability worth testing.
- Graduate to red team engagements once your program is mature enough to learn from them. Doing red team on an immature program is like sending someone with no fitness baseline to run a marathon, you'll just get hurt.