# 20. Welcome to 2026: What Do We Actually Do With Frontier Models?

Share
# 20. Welcome to 2026: What Do We Actually Do With Frontier Models?

A practical guide for security leaders who are done waiting for permission to engage with the above question.


Context: Frontier models, the current generation of large language models and multimodal AI systems from Anthropic, OpenAI, Google, Meta, Mistral, and others, have crossed a threshold in 2025-2026 that makes the "wait and see" posture no longer tenable. These are not toys. They are not a future concern. They are active participants in the threat landscape right now, on both sides of the line.

On the offensive side: adversaries are using frontier models to generate convincing phishing at scale, accelerate vulnerability research, write functional malware variants, conduct more sophisticated social engineering, and reduce the skill floor for complex attacks. The barrier to entry for capable attacks has dropped significantly, and it will keep dropping.

On the defensive side: frontier models offer genuine, demonstrable capability improvements in threat detection, alert triage, threat intelligence analysis, code review, documentation, and analyst augmentation. The organizations getting ahead are not the ones with the biggest AI budgets, they are the ones that have asked and answered the practical questions clearly.

The problem most security leaders face is not a shortage of vendor pitches. Every vendor in Topics 1 through 19 of this series now has an "AI-powered" story (see Topic 1 for how to interrogate those claims). The actual problem is that leaders at "regular" companies don't have a clear framework for thinking about what frontier models mean for their specific security program, and where the genuine leverage points are versus the noise.

This topic is structured differently from the others. It is less about evaluating a specific vendor and more about giving you a clear framework for thinking through your own frontier model posture, with questions you should be asking of yourself, your team, and any vendor or model provider you engage with.

The core question for every security leader in 2026: "Where are frontier models already affecting my threat landscape, where can I use them to improve my defenses, and what governance do I need to do both responsibly?"


14 Questions Every Security Leader Should Be Asking About Frontier Models

1. "Do we have a clear picture of how frontier models are already being used against us, not just how we might use them ourselves?"
Why: Most organizations are thinking about AI adoption from the inside out, how can we use it, what tools can we buy. The threat has already moved. AI-generated phishing, voice cloning for BEC (see Topic 9), AI-accelerated vulnerability scanning against your attack surface, automated credential stuffing at scale, these are current operational realities, not future scenarios.
What good looks like: Security leadership has a current threat assessment specifically addressing AI-augmented attack vectors relevant to their sector. Tabletop exercises include AI-enabled attack scenarios. Detection tuning has been updated to address AI-generated content.
Red flag: "We're monitoring the space" with no specific changes to detection, awareness training, or process controls.

2. "Where in our security operations does the human analyst bottleneck most damage our response capability, and could a frontier model help there first?"
Why: Frontier models are not evenly useful. They are transformationally useful in specific, well-defined tasks: alert triage and enrichment, threat intelligence summarization, incident timeline construction, first-draft runbook generation, log pattern explanation, and security code review. They are much less useful in tasks requiring deep organizational context, political judgment, or novel adversary reasoning. Starting where the leverage is highest produces faster, more defensible ROI.
What good looks like: A specific, named workflow in your SOC or security program where analyst time is spent on work a frontier model could handle as well or better, with a clear plan to pilot it.
Red flag: "We're evaluating AI across the board" with no specific prioritized use case.

3. "What data are we comfortable sending to a frontier model, and what governance exists to enforce that boundary?"
Why: This is the foundational governance question, and most organizations have not answered it explicitly. Frontier models used via commercial APIs (OpenAI, Anthropic, Google) process the data you send them. Depending on the API terms, that data may be used for model training, may be retained, or may be accessible to the vendor under certain circumstances. Your security data, including alerts, logs, threat intelligence, incident details, and network telemetry, is among the most sensitive data your organization holds. Before any frontier model deployment touches operational security data, you need explicit answers about where it goes and who can access it.
What good looks like: A data classification policy that explicitly addresses which data categories can be sent to which model types (API, on-prem, private cloud). Approved tools list enforced by policy and technical controls. DLP tuned to catch unauthorized model usage.
Red flag: Analysts freely using commercial AI tools with production security data, with no policy governing what they can share.

4. "Are we treating our AI tool pipeline as a supply chain risk, or just as a productivity tool?"
Why: Every AI model, API, plugin, and integration you adopt becomes part of your software supply chain (see Topic 19). Models can be manipulated via prompt injection if they process attacker-controlled data. Third-party AI plugins extend the attack surface. Model providers themselves are high-value targets. The 2025 emergence of prompt injection attacks against AI-integrated security tools demonstrated that tools using LLMs to process alerts, emails, or threat feeds can be manipulated if the input is attacker-controlled. Your AI tools are not just productivity aids, they are attack surfaces.
What good looks like: AI tools inventoried and treated like any other third-party software. Prompt injection considered in any deployment where a model processes external or attacker-influenced data. Model provider security posture reviewed as part of vendor risk management.
Red flag: AI tools adopted by the security team under the radar, without going through standard third-party risk or change management processes.

5. "Can we detect when AI is being used against us, or are our detection models still calibrated for pre-AI attack patterns?"
Why: The signals that historically indicated a phishing email, a social engineering attempt, or a reconnaissance activity have changed. AI-generated phishing has no typos, matches the recipient's communication style, and references contextually accurate details. AI-accelerated vulnerability scanning moves faster than traditional scanning signatures catch. Detection rules and user awareness training that were accurate two years ago may now create a false sense of coverage. This is not a hypothetical: research has demonstrated that current email security tools catch AI-generated spear-phishing at significantly lower rates than traditionally crafted attacks.
What good looks like: Detection engineering team has explicitly reviewed and updated content for AI-augmented attack variants. Awareness training messaging has been updated to move from "look for typos" to "verify context and intent out-of-band." Security team has tested their own controls against AI-generated attack simulations.
Red flag: Awareness training still teaches people to look for grammatical errors as a phishing indicator.

6. "Do we have a position on agentic AI, and what access controls govern what AI agents can do in our environment?"
Why: Agentic AI, systems that can take sequences of actions autonomously rather than just answering questions, is the next step already being deployed. Security vendors are building "autonomous SOC analyst" tools that can query systems, run scripts, create tickets, and take containment actions without human approval. Your own developers and operations teams may be deploying agentic AI tools that have API access to production systems. The governance question is acute: what can an AI agent do, with what access, with what oversight, and who is accountable when it takes a wrong action?
What good looks like: Explicit policy on agentic AI tools. All agentic tools inventoried with their permissions documented. Least-privilege access for AI agents enforced. High-impact actions (isolate a host, disable an account, make a network change) require human approval regardless of the agent's confidence. Audit logging for all agent actions.
Red flag: Agentic tools deployed with broad permissions because "it needs access to do its job," without explicit governance on what that access covers.

7. "What is our organization's AI governance framework, and does security have a seat at the table where those decisions are made?"
Why: Frontier model deployment decisions are being made across your organization right now, in marketing, HR, finance, legal, product, and operations, often without meaningful security review. The data governance, privacy, and security implications of how frontier models are deployed in business functions are security concerns. A customer service chatbot trained on customer data, a finance AI with access to payment systems, an HR tool processing employee records, all of these are in your threat model whether or not they appear in your security inventory.
What good looks like: Security represented in the AI governance committee or equivalent. A formal AI use case review process that includes security and privacy assessment. An inventory of approved AI tools maintained and enforced.
Red flag: Security learns about AI deployments in other departments after they're live, or not at all.

8. "How are we measuring the security of frontier model outputs, not just their usefulness?"
Why: Frontier models hallucinate. They present confident, fluent, wrong information. In a security context, a model that confidently misidentifies a benign alert as a critical threat, or that generates a remediation script with a subtle security flaw, or that summarizes a threat intelligence report inaccurately, can cause real harm. The same rigor applied to other security tools, does it work, how do we know, what's the failure mode, should apply to AI-augmented security operations.
What good looks like: Explicit validation processes for AI-generated security outputs. Analysts trained to treat AI output as a starting point, not a conclusion. Red-teaming of your own AI-augmented workflows to identify failure modes. Metrics tracking where AI outputs required human correction.
Red flag: "The AI said so" as an acceptable justification for a security decision without independent verification.

9. "Are we building internal AI literacy on the security team, or waiting for vendors to explain it to us?"
Why: The security leaders and teams getting the most from frontier models are not waiting for vendor enablement sessions. They are building genuine understanding of how these models work, where they fail, and how to use them effectively. This is not about becoming ML engineers. It is about understanding enough to evaluate vendor claims (Topic 1), to use models effectively in daily work, to spot when AI output is unreliable, and to have credible conversations with the organization's AI governance function.
What good looks like: Security team members have hands-on experience with frontier models in their actual work. Team has developed internal guidance on which tasks benefit from AI augmentation and how to use it safely. Leadership can evaluate AI vendor claims without relying entirely on the vendor.
Red flag: Security team's primary exposure to frontier models is watching vendor demos.

10. "What does 'machine-speed' actually mean in our environment, and are we prepared for attack chains that close faster than our current response processes?"
Why: The compression of attack timelines is the most operationally significant consequence of AI augmentation on the offensive side. Breach-to-ransomware timelines that historically took days are now measurable in hours in AI-augmented attack chains. Lateral movement that required skilled manual effort can now be partially automated. Vulnerability exploitation windows are shrinking as AI-accelerated scanning identifies and exploits weaknesses faster. If your current detection-to-response process assumes you have hours to respond, and the attack chain closes in 45 minutes, you have a structural gap.
What good looks like: SOC response playbooks reviewed against realistic AI-augmented attack timelines. Pre-approved containment actions that do not require multi-level approval for the most time-critical responses. Automation in place for the highest-volume, most time-sensitive response actions (endpoint isolation, account suspension, network block).
Red flag: Response processes designed around the assumption of hours of dwell time, with no review since AI-augmented attack timelines became a documented reality.

11. "Do we have a clear story for the board and executive team on frontier model risk, without either overstating it or dismissing it?"
Why: Frontier model risk is real, but the communications challenge is significant. Overstate it and you sound alarmist, potentially triggering hasty decisions. Dismiss it and you've failed to brief leadership on a material risk. The board-level conversation needs to cover three things clearly: how frontier models are changing the threat landscape the organization faces, what the organization is doing in response, and what governance exists for the organization's own AI deployments. Most security leaders don't yet have a crisp, calibrated version of this conversation ready.
What good looks like: A concise, non-technical board briefing that addresses offensive AI risk, defensive AI opportunity, and governance posture. Updated at least annually. Security leader comfortable fielding AI questions from board members without resorting to jargon or vendor slides.
Red flag: Board awareness of AI risk limited to what they read in the news, with no structured briefing from the security function.

12. "Where are the frontier model vendors themselves in our threat model, and do we understand the security implications of the models we're relying on?"
Why: The major frontier model providers, Anthropic, OpenAI, Google, Meta, Mistral, are themselves high-value targets. A compromise of a model provider's API, a poisoning of a model's training data, or a subtle capability degradation in a model you rely on for security operations all have direct security implications. This is the supply chain risk (Topic 19) applied to AI. You should know which models are in use in your environment, via which integrations, and what your contingency is if a model provider has a security incident or changes terms materially.
What good looks like: AI model providers treated as Tier 1 or Tier 2 vendors in your TPRM program (Topic 16). Awareness of which security tools in your stack call which model APIs. Contingency consideration for model provider outage or incident.
Red flag: No inventory of which external AI APIs your security tools call, or which data those calls include.

13. "What are we doing with the compliance and regulatory angle on AI, and are we ahead of it or behind it?"
Why: The regulatory landscape for AI is moving fast and unevenly. The EU AI Act is in phased implementation. US federal AI guidance from NIST, CISA, and sector regulators is proliferating. Healthcare, financial services, critical infrastructure, and defense sectors all have emerging AI-specific requirements or guidance documents that intersect with security. The organizations that are building auditable, explainable, governed AI programs now will have a materially easier compliance posture when requirements crystallize. Those that aren't will face retroactive remediation.
What good looks like: Legal and compliance teams briefed on AI regulatory trajectory relevant to your sector. AI deployments documented with the same rigor as other regulated tools. Security program able to demonstrate human oversight, explainability, and audit trails for AI-augmented decisions.
Red flag: "We'll deal with compliance when the regulations are finalized" with no current governance documentation.

14. "What does our frontier model roadmap actually look like, and who owns it?"
Why: Most organizations have an uncoordinated collection of AI experiments, vendor tools with embedded AI, and individual productivity uses, but no coherent roadmap for where they want to be in 12 and 24 months. The organizations getting ahead are not necessarily the ones moving fastest. They are the ones moving deliberately, with clear use case prioritization, explicit governance, defined success metrics, and an owner accountable for the program. Security should not be a passive recipient of whatever AI the organization deploys. It should have a seat in shaping the roadmap and a specific roadmap of its own for AI augmentation of the security function.
What good looks like: A named owner for AI governance within the security function. A prioritized use case list with owners, timelines, and success metrics. A defined review cadence to assess progress and adjust priorities as the technology evolves.
Red flag: AI strategy described as "we're exploring options" with no owner, no roadmap, and no metrics.


The Meta-Test

  • Q3, Q4, and Q12 are the governance and supply chain test. If you don't know what data your AI tools are sending where, what access your AI agents have, or how your model providers are secured, you have adopted AI faster than your governance.
  • Q1, Q5, and Q10 are the threat reality test. Frontier models are already in the threat landscape. The organizations that have explicitly updated their threat model, detection content, and response timelines for AI-augmented attacks are the ones that will not be caught flat-footed.
  • Q2 and Q9 are the practical leverage test. Not all AI use is equal. Starting with the highest-leverage, lowest-risk use cases, alert enrichment, TI summarization, runbook drafting, and building genuine team literacy produces durable results. Starting with the flashiest pitch produces impressive demos and disappointing programs.

Honest Observations: Where to Actually Start in 2026

The question we hear most from security leaders at "regular" companies is not "should we use AI?" That debate is over. The questions are "where do we start, what do we do next, and how do we do this without creating more problems than we solve?"

Here is the honest framework, drawn from what we have seen work:

1. The threat side comes before the tool side.

Before you buy a single AI-powered security tool or deploy a single AI model in your security operations, spend 90 minutes with your team answering this question honestly: how is AI already changing the attacks we face? AI-generated phishing targeting your employees, AI-accelerated vulnerability scanning of your external attack surface, AI-assisted BEC against your finance team, deepfake voice calls targeting your executives (see the deepfakes work), these are not hypothetical. Update your threat model first. Everything else follows from there.

2. The highest-ROI early deployments are almost always analyst augmentation, not automation.

The organizations seeing the clearest early returns are those using frontier models to augment analyst work, not replace it. A model that enriches an alert with relevant context, summarizes a 200-page threat intelligence report into three key points, drafts the first version of an incident timeline, or explains a complex log pattern in plain English, that model is saving hours per analyst per week with minimal risk of harm if it gets something wrong, because a human reviews it.

The "autonomous SOC" vision is further out, requires more governance, and has more potential downside. Start with augmentation. Build trust in the tools. Expand scope as confidence warrants.

3. Data governance is not optional and it is not IT's problem.

The security team needs to own the data governance question for its own AI deployments. What can go to a commercial API? What requires an on-premises or private cloud model? What can never leave a controlled environment? These are security decisions, not just privacy ones. The classification framework should exist before tools are deployed, not after analysts have been using them for six months.

If you are in a regulated industry, healthcare, financial services, critical infrastructure, legal input on this question is not optional. The data flowing through your security tools is often sensitive by definition.

4. Build your 12-point program from the foundation up, not the top down.

The 12 priorities Sid outlined at the start of this topic represent a mature, comprehensive AI security posture. Not every organization can pursue all twelve simultaneously. The practical sequencing for most "regular" companies:

First 90 days, the foundation:

  • Update your threat model for AI-augmented attacks (Q1, Q5)
  • Inventory all AI tools currently in use across the security team and the broader organization (Q7)
  • Establish data governance: what can go where, enforced by policy (Q3)
  • Run one tabletop scenario involving an AI-augmented attack vector

Months 3-6, early value:

  • Deploy frontier model augmentation in your highest-bottleneck analyst workflow (Q2)
  • Update awareness training to address AI-generated attacks (Q5)
  • Brief the board with a calibrated AI risk and governance narrative (Q11)
  • Add AI model providers to your TPRM review cycle (Q12)

Months 6-12, building the program:

  • Establish governance for agentic AI tools as they appear in vendor pitches and internal tooling (Q6)
  • Build your AI security roadmap with an owner and metrics (Q14)
  • Review detection engineering content against AI-augmented attack patterns (Q5, Q10)
  • Connect your AI governance to the broader enterprise AI governance function (Q7)

Year 2 and beyond:

  • Machine-speed response automation for the highest-frequency, highest-confidence containment actions
  • Behavior-centric, identity-first defenses hardened against AI-enabled social engineering
  • AI threat intelligence operationalized into detection content
  • Compliance posture documented and defensible as regulatory frameworks crystallize

5. The sequencing principle that matters most: govern what you have before you add more.

The single most common AI security mistake is adding AI capabilities faster than you can govern them. An organization with 20 AI tools, unclear data flows, no agentic access controls, and no audit trail is more exposed than one with 3 AI tools that are well-understood, properly scoped, and genuinely improving security outcomes.

Resist the vendor pressure to buy the AI platform before you have the governance foundation. The technology will still be there in six months. The breach that happens because an AI agent had excessive permissions or because your AI tool was processing attacker-controlled data will not wait.

6. The question that simplifies everything.

When a vendor claims their tool is "AI-powered," when a colleague suggests using a frontier model for a security task, when you are evaluating whether to expand AI in your program, ask this question: "What specific decision or action does this AI capability improve, for whom, and how will we know if it's working or not?"

If there is a clear answer, engage seriously. If the answer is vague, impressive-sounding, or deferred, apply the scrutiny from Topic 1 of this series and do not buy until there is a clear answer.

The organizations that come out of this period of AI disruption in stronger security postures will be the ones that asked that question early and answered it honestly.


Prepared collaboratively with: Chris Roberts, Sid, GPT 5.5, Kimi 2.6, Claude (Anthropic) | 2026 | Working Series: What to Ask, What to Demand, What to Walk Away From.


Read more