Red Team Assessment: Scope, Process, Deliverables

Red Team Assessment: Scope, Process, Deliverables

15 min read October 8, 2026
Nazar Tymoshyk Nazar Tymoshyk Author

TL;DR

  • Red teaming measures whether your people, processes and tools catch a realistic attacker. In CISA’s 2026 test of two organizations, the team took both domains, and only one SOC noticed.
  • Scope starts with objectives, such as a customer database or domain administrator rights, then fixes the starting position and the hard limits in written rules of engagement.
  • Under the EU’s Digital Operational Resilience Act (DORA), regulated red teaming needs at least 12 weeks of active testing, and the red team report is due within 4 weeks of that phase ending.
  • A useful report shows the attack paths that worked and failed, what your defenders saw, root causes and ranked fixes, followed by a replay with your blue team.

In August 2026, the US Cybersecurity and Infrastructure Security Agency (CISA) published two red team assessments it ran at the same time, with similar tradecraft, against two critical infrastructure organizations. In both, the team reached full domain compromise and sensitive business systems. Only one of the two security operations centers (SOCs) noticed.

According to CISA’s 2026 advisory, Organization B’s analysts isolated the three phished workstations within 10, 2 and 20 minutes, forcing the red team to continue from a foothold the organization handed it. At Organization A, the same activity raised medium- and low-severity EDR alerts that nobody acted on, buried under thousands of false positives from normal business activity.

That gap, between owning detection tools and acting on what they show, is what a red team assessment measures.

Ready to book a pentest? Talk to a team that has run incident response on real breaches and knows where attackers get in

What Is a Red Team Assessment?

A red team assessment is an authorized simulation of a real attack, run by a team whose job is to reach agreed objectives without being caught. NIST’s glossary, quoting the Committee on National Security Systems (CNSSI 4009-2022), defines a red team as a group “authorized and organized to emulate a potential adversary’s attack or exploitation capabilities against an enterprise’s security posture.”

The same definition gives the purpose: to improve security “by demonstrating the impacts of successful attacks and by demonstrating what works for the defenders.” So a red team answers a different question from a vulnerability scan: would we notice, and what would we do about it?

That makes your defenders, usually called the blue team, as much the subject of the test as your systems. The European Central Bank’s TIBER-EU framework puts it plainly: the outcome “is not a pass or fail,” and the point is to learn where resilience holds and where it breaks.

Reaching an objective and being noticed are separate questions, so every objective can end four ways.

Red team assessment outcomes as a 2x2 grid: objective reached or stopped, noticed or not

Each objective is usually called a flag, meaning a system, a dataset or a level of access the team has to reach and prove. A flag reached quietly tells you something very different from a flag reached while your SOC was chasing the intruder, which is why both outcomes get recorded.

Red Team Assessment vs Penetration Test

The red team assessment vs penetration test question usually comes up first in scoping, because the two share tools and often the same people. A penetration test tries to find as many exploitable weaknesses as it can inside a defined scope. A red team picks a goal and heads for it the way a real attacker would, quietly and by whatever route works.

The differences show up in almost every part of the engagement:

Penetration test Red team assessment
Main question What can be exploited in this scope? Would we detect and stop an attacker going after this objective?
Scope Named systems, applications or IP ranges Objectives, with many possible routes in
Who knows it’s happening Usually the IT and security teams A small control team only
Stealth Not required; testers can be noisy Central; the team works to avoid detection
People and physical access Usually out of scope Often in scope: phishing, phone pretexts, site entry
What gets measured Vulnerabilities and their severity Detection, response and the paths that reached each flag
Typical output A findings list with severity ratings and fixes An attack narrative, flags reached and missed, a blue team scorecard
Duration Usually a few weeks Weeks to months; at least 12 weeks of active testing under DORA

UnderDefense, which runs both, puts its average penetration test at about 3 working weeks. A regulated red team’s active phase alone runs at least four times that.

The two work best in sequence. Start with penetration testing if the basics aren’t fixed yet, because a red team that walks straight through unpatched systems tells you what a cheaper pentest would have.

Penetration tests come in several forms, and the types of penetration test differ mainly in how much the testers are told before they start. A red team sits at the far end of that scale, told almost nothing and expected to find its own way.

Signs You’re Ready for Red Teaming

A red team engagement is expensive feedback, so it pays off only when you can act on what it finds. Before you commission one, check how many of these already hold:

  • Critical and high findings from your last penetration test are fixed and retested.
  • Someone acts on alerts around the clock, in-house or through a provider, with the authority to isolate a host.
  • Endpoint, identity, cloud and network logs are collected centrally and kept long enough to reconstruct an attack weeks later.
  • You have a written incident response plan and have exercised it at least once.
  • You know which systems and data would hurt most if an attacker reached them.
  • An executive will sign the authorization letter and back the remediation budget.
  • Your team has time in the following quarter to fix what the test finds.

If several of those are missing, a penetration test or a tabletop exercise will teach you more per dollar, since a red team run against a program without central logging mostly proves that nobody could have seen the attack.

What prompts one depends on who you answer to: the DORA testing cycle for an EU financial entity, a large customer’s security review for a SaaS vendor, or a mid-market board asking whether anyone would notice an intruder a clean pentest missed.

Scoping a Red Team Security Assessment

Scoping a red team security assessment comes down to deciding what the attacker wants, where it starts and what it must never touch.

Start With Objectives an Attacker Would Pay For

Write each objective as an outcome: read the customer database, change a payment instruction, push code into your CI/CD pipeline or hold domain administrator rights. Each one becomes a flag the team has to reach and prove, usually with a screenshot or a marker file your control team can verify.

One or two flags is plenty for a first engagement. Every extra objective adds scenarios, testing time and reporting, and a scope with eight flags tends to end with shallow attempts at all of them.

Choose the Starting Position

A full-chain scenario starts outside with only public information, so it tests your perimeter, your email filtering and your people on the way in. An assumed-breach scenario hands the team a foothold, such as a standard employee laptop or a VPN account, and spends the budget on what happens after initial access.

Laid over one attack chain, the two starting positions differ only in where the testing begins.

Full-chain versus assumed-breach red team assessment, bracketed over one five-step attack chain

Assumed breach is often the better buy, since phishing tends to work eventually and weeks spent proving it add little. Regulated tests accept this kind of help openly: the EU’s rules require the report to list every “leg-up” granted, defined as assistance the control team gives testers who can’t advance an attack path on their own.

Decide What’s Out of Scope

List what’s off-limits before anyone starts: databases that can’t tolerate a crash, medical or industrial devices, executives who are traveling and any tactic your legal team hasn’t cleared.

Agree every social engineering pretext in writing before it’s used, because a fake HR email about bonuses upsets staff in a way a fake IT password reset doesn’t, even when both are legal.

Rules of Engagement and Deconfliction

Rules of engagement turn the scope into a document both sides sign. It protects your business from a test that goes too far, and it protects the testers when they do exactly what you hired them to do.

A usable rules-of-engagement document covers at least these points:

  • In-scope domains, IP ranges, cloud accounts, offices and named scenarios
  • Systems, people and actions that are explicitly out of scope
  • Blackout windows, such as payroll runs, trading hours or a product launch
  • Stop conditions, including any sign of a real attacker or an unplanned outage
  • How the team handles sensitive data it reaches, and how that data is destroyed afterward
  • Contacts on both sides who can be reached at any hour, each with a backup

The authorization letter, often called a get-out-of-jail letter, is signed by an executive with the authority to approve the test. Testers carry it during physical and social engineering work, so a guard, a building manager or the police can confirm the activity with a phone call.

Deconfliction is the process that stops a red team from triggering a real incident response, and stops a real attacker from hiding behind the test. The ECB’s TIBER-EU framework puts a control team in charge of the test, “a small team within the target entity whose members are the only ones there who know a test is happening.”

When your SOC spots something odd, the control team confirms whether it’s the red team before anyone decides how far to escalate. If it isn’t, the test pauses and your normal incident response takes over.

Red Team Assessment Methodology, Phase by Phase

Most red team assessment methodology write-ups differ in naming more than in substance. The version here follows the EU’s threat-led testing rules (Commission Delegated Regulation 2025/1190, published in June 2025), because they spell out what each phase has to produce.

The last column shows what the EU regulation requires for regulated tests:

Phase What happens What it produces What the EU rules require
Preparation Scope, flags, rules of engagement and the control team are agreed Signed scope and rules of engagement A scope specification document approved by the testing authority
Threat intelligence Research into who would realistically target you, and how Attack scenarios based on real threat actors A targeted threat intelligence report
Active testing The red team runs the scenarios against your live defenses Flags reached or missed, a timestamped log of every action At least 12 weeks of active red team testing
Closure The blue team learns of the test and both sides compare notes Red team report, blue team report, replay session Red team report within 4 weeks; blue team report, replay and purple teaming within 10 weeks
Remediation Fixes are planned, made and checked Remediation plan, then a retest A remediation plan and an attestation

Many providers sell this work as adversary simulation, and the same phases apply whatever the proposal calls it.

Threat Intelligence and Scenarios

Threat intelligence turns “who would attack us?” into specific scenarios: which groups go after your sector, how they usually get in and which of your people and systems are visible from outside.

Active Testing

Active testing is the part people picture: phishing, exploiting exposed services, moving between machines, escalating privileges and heading for the flags. The team logs every action with a timestamp, so your SOC can later check each step against its own alerts and the testers have a record if something unrelated breaks.

Closure and Replay

Once active testing ends, the control team tells the blue team that a test took place. Both sides then compare the red team’s log with what the SOC actually saw, and the gaps between the two become the most useful findings you’ll get.

What Red Teams Find: Alerts Nobody Acted On

Public red team results are rare, which makes CISA’s advisories on its own assessments the best public record.

In a 2022 assessment that CISA’s 2023 advisory describes, the team got in through spearphishing emails, used Active Directory data to phish its way onto a third host, moved to a misconfigured server and from there compromised the domain controller. Multifactor authentication prompts stopped it reaching one sensitive business system, which shows you which control was actually holding.

Drawn as a route, the 2022 assessment shows exactly where the attack stopped and what stopped it.

CISA's 2022 red team assessment route, from phished workstations to the domain controller, stopped by MFA

The 2026 pair shows what separated the SOC that caught the team. Organization B kept a baseline of normal activity and tuned its alerts, so one medium-severity alert stood out. Organization A ran several SOCs and several EDR tools whose staff didn’t talk to each other, with no procedure for escalating an alert.

Even at Organization B, the team reached domain control from the foothold it was given, through cleartext credentials in an XML file and a service account with rights over almost 1,000 accounts.

Commercial results look much the same. In an attack simulation for a US food producer with \$6 billion in revenue and 10,000 employees, UnderDefense’s testers ran two phishing waves: “Out of 331 emails in total, 43 employees clicked the malicious links,” and in the second wave 40 of 529 recipients entered their credentials.

The same engagement found “6 days of silent C2 activity,” meaning command-and-control traffic the client’s SOC never flagged, along with a 10 GB file exfiltrated to Dropbox. The phishing numbers show how people behave; the six days show how long an attacker who got through would have been left alone.

Case study

How 10,000 employees responded to an attack simulation

A US food producer with $6 billion in annual revenue and 10,000 employees hired UnderDefense’s red team for a purple team assessment, including two phishing waves against its staff.

Attack path

  • Started from a standard user account on a domain-joined workstation, connected over VPN.
  • Escalated to Domain Admin through a misconfigured Active Directory Certificate Services template.
  • Exfiltrated a 10 GB file to Dropbox and encrypted 500+ files without triggering an alert.
6 daysSilent C2 activityNever flagged by the SOC
43 / 331Clicked the phishing linkFirst wave
40 / 529Entered their credentialsSecond wave

What’s in a Red Team Assessment Report

A red team assessment report is what you’re actually paying for, and Annex V of the EU regulation sets out what one must contain at a minimum.

The Red Team Report

The first block covers the attack itself: the targeted functions and systems, a summary of each scenario, the flags reached and not reached, the attack paths and techniques that worked and those that failed, any deviations from the plan and any leg-ups granted. The second block records every action the testers saw the blue team take to reconstruct or contain the attack.

The third block is findings, meaning each vulnerability with its criticality, a root cause analysis of the successful attacks and remediation recommendations ranked by priority. Under DORA the testers must deliver the report within 4 weeks of active testing ending, which is a reasonable deadline to write into any contract.

Failed paths matter as much as successful ones. A route that MFA or network segmentation blocked is evidence that a control works, and short of a real attack it’s the only such evidence you’ll get.

The Blue Team Report and Scoring

Regulated tests add a second report, written by your defenders. For each attack step it lists what was detected and the log entries behind each detection, followed by the blue team’s own root cause analysis, lessons learned and topics for a joint purple team session.

Outside regulated programs you can ask for the same report, which turns the red team’s narrative into a scorecard. No standard defines the measures below, so agree the definitions with your provider before testing starts:

Measure How it’s measured What it tells you
Detection coverage Share of attack steps that produced at least one alert or logged event Which techniques your tooling can see at all
Time to detect From the red team’s first action in a step to the first alert on it How long an attacker operates unnoticed
Time to escalate From the first alert to a person declaring an incident Whether alerts reach someone with authority to act
Time to contain From escalation to the attacker losing access Whether your response actually stops anything
Response to deliberate noise Whether anyone reacts when the team tries to get caught Whether detection works even when handed a chance

That last measure comes from CISA’s method. In the second phase of its assessments, the red team “attempts to trigger a security response” on purpose. In the 2022 test, against an organization CISA called mature, even that went unanswered.

Red Team Assessment Reporting for Different Readers

Good red team assessment reporting is split by audience. Your board needs one page showing which objectives an attacker could reach, how long it would go unnoticed and what the fix costs. Your CISO needs the attack narrative and the scorecard, your SOC needs every timestamped step mapped to MITRE ATT\&CK, and engineering needs reproducible findings with evidence.

Compared with a typical penetration testing report, a red team report adds the narrative, the flags and the blue team’s side of events, and it drops most of the low-severity noise.

Ask for a replay session, too. DORA’s rules expect the blue team and testers to “replay the attack and review the steps taken” together, and that session is usually where the findings finally make sense to the people who have to fix them.

Lining up a pentest and unsure what to scope? Get a session on scope, timing, and what the report needs to prove to auditors or the board

TIBER-EU and DORA TLPT: Red Teaming Under Regulation

For EU financial entities, red teaming may not be optional. The Digital Operational Resilience Act, adopted in 2022, requires identified financial entities to “carry out at least every 3 years advanced testing by means of TLPT,” meaning threat-led penetration testing, and its technical standards set out how that red teaming must run.

DORA’s standards mirror TIBER-EU, the framework the ECB published in May 2018 and updated in 2024 to align with DORA. The ECB’s 2025 guidance documents cover each step behind it: scoping, threat intelligence, test planning, both reports, remediation and procurement.

Providers face rules too, from CVs and certifications for the control team to professional indemnity insurance, and the test closes with an attestation other EU authorities can recognize. TIBER-EU isn’t limited to banks: the ECB says any critical sector can use it.

If your entity is in scope, the three-year clock and a remediation plan after each test are fixed costs. UnderDefense runs DORA threat-led penetration testing against real-world attack scenarios and ends it with a prioritized remediation report, the starting point for that plan.

What Drives Red Team Pricing

Red team quotes vary widely, because scope moves them more than anything else. UnderDefense’s 2023 penetration testing cost guide lists red teaming at an average of \$10,000 to \$50,000, the highest starting point of any test type in its comparison.

Duration is the biggest cost driver: 12 weeks of active testing under DORA, plus preparation and closure, is a very different project from a single-flag engagement measured in weeks.

People and physical scope add cost quickly. Phishing campaigns, phone pretexting and site entry need different specialists, legal review and sometimes travel.

Two quotes for what sounds like the same engagement can cover very different amounts of work.

Two red team assessment quotes compared: both include testing and a report, only one includes closure work

Closure work is easy to leave out of a quote. The replay, the purple team session and any retest can each be priced separately, so check what’s included before you compare two proposals.

How to Choose a Red Team Provider

Start with their red team assessment reporting, since a sample report tells you more than any sales call. Then ask these questions before you sign:

  • Can we see a redacted report from a past engagement, including the blue team scorecard?
  • Who exactly will be on our engagement, and what have they tested before?
  • How do you handle deconfliction, and who do we call at 2 a.m.?
  • Do you map every action to MITRE ATT\&CK and hand over the timestamped log?
  • Is a retest of our fixes included, along with the replay and the purple team session?
  • What happens if you find evidence of a real attacker during the test?

Be wary of proposals that promise a set number of findings or guarantee a break-in. A good red team can’t promise either, and its value lies in how precisely it tells you what your defenders saw.

Testers who think like attackers are what you’re paying for. UnderDefense started in 2017 as a penetration-testing company and runs red teaming alongside its other offensive testing. Clutch reviewers rate it 4.9 out of 5 (66 reviews, September 2026), and a legal-tech security VP who hired it for gray-box penetration testing wrote that the issues its testers found showed they went beyond running tools.

Turning Findings Into Fixes

A red team report is worth only the fixes it produces. DORA requires a formal remediation plan after every test, and the habit works outside regulation too: every finding gets an owner, a deadline and a priority taken straight from the report.

Split the work in two. Configuration and credential findings, like the cleartext passwords and overprivileged service account in CISA’s 2026 assessments, go to engineering and can be retested within weeks. Detection findings go to whoever writes your detection rules, and every attack step that produced no alert becomes a use case to build, test against the recorded red team activity and keep.

Each finding goes down one of two lanes, and each lane has its own owner and its own test.

Red team assessment findings split into configuration fixes for engineering and detection gaps for rule writers

The awkward findings are the steps that did raise an alert nobody acted on. CISA’s first recommendation after the 2026 pair was to “establish and continuously maintain a baseline and reduce alert noise,” so a real anomaly stands out the way it did for Organization B.

That is where the UnderDefense Agentic AI SOC earns its place: every alert your SIEM and EDR raise, whether that’s Splunk and CrowdStrike or Microsoft Sentinel and SentinelOne, gets investigated with the evidence attached before an analyst decides, and UnderDefense reports about 2-minute alert-to-triage. Your team spends its time on the incidents that matter, and the medium-severity alert from your next test gets read.

Expect a calibration period, though: several G2 reviewers say tuning the Agentic AI SOC to their stack took weeks to a few months.

A second, shorter test against the same flags is the only proof that the fixes changed what an attacker can actually do.

Start Small: One Flag, a Control Team, a Retest Date

Keep a first red team assessment small, and settle the unglamorous parts early: a named control team, a signed authorization letter and a contract that lists the report contents and the replay session.

Run it after your last pentest findings are closed, and book the retest date before testing starts so the fixes have a deadline. When the report lands, the attack steps that raised an alert nobody acted on are the ones worth watching investigated end to end on your own stack.

Got red team steps that fired an alert and went nowhere? Bring them to a demo

Frequently Asked Questions

How long does a red team assessment take?

Length depends on scope and on whether regulation applies. Under DORA, active red teaming alone runs at least 12 weeks. A commercial engagement scoped to one or two flags is usually much shorter, so ask the provider for a phase-by-phase timeline.

What is red team assessment in plain terms?

Red teaming is a hired, authorized attack on your organization with a specific goal, such as reaching a customer database, carried out quietly. The aim is to show whether your security team notices and responds, and which defenses held. The red team assessment report then explains the route taken, what was detected and what to fix first.

What is the difference between a red team, a blue team and a purple team?

The red team plays the attacker. The blue team is your defenders, usually the SOC and incident responders, who are tested without being warned. Purple teaming is a joint session where both sides work together after a test, or sometimes during one, to replay attacks and build detections.

Can red teaming disrupt normal operations?

Yes, which is why the rules of engagement list off-limits systems, blackout windows and stop conditions. Testers avoid destructive actions and prove each flag with a screenshot or a marker file, and the control team can pause the test at any point.

How often should you run a red team test?

EU financial entities in scope for DORA must run threat-led testing at least every three years, and their supervisor can adjust that interval based on risk. Outside regulation, a full red team security assessment every year or two, with smaller purple team exercises in between, is a sensible rhythm.

Book a 30-minute
pentest scoping session

  • Review your attack surface and what belongs in scope
  • Choose between a pentest, an assumed-breach test or a full red team
  • Get a test plan with timeline, rules of engagement and report format
Nazar Tymoshyk

Nazar Tymoshyk

CEO and the driving force behind UnderDefense

Nazar Tymoshyk is a visionary cybersecurity expert with extensive industry experience, holding a Ph.D. in Information Security, an MBA, and a degree in Computer/Information Technology Administration and Management.

Nazar’s contributions to cybersecurity have earned him recognition as a respected leader in the field. His insights have been featured in leading publications, including The Wall Street Journal, TechCrunch, and TechRepublic.

As the founder of UnderDefense, Nazar has demonstrated exceptional leadership, growing the company into a recognized provider of advanced cybersecurity solutions known for its innovative approach and strong commitment to client success. His mission is to transform how businesses approach cybersecurity by delivering tailored solutions for every stage of growth.

Nazar’s dedication to national cybersecurity also led him to serve in CERT-UA, where he played a key role in strengthening Ukraine’s cyber defense capabilities.

See MAXI investigate a real incident

  • Watch a signal go to verdict in two minutes
  • See the human decision gate in action
  • Get a tailored MDR rollout plan

10 Best Agentic AI SOC Platforms for 2026

10 Best Agentic AI SOC Platforms for 2026

Compare 8 best agentic SOC platforms for 2026. Pricing, integrations, compliance, and POC frameworks scored across 500+ MDR deployments. Evaluate now.