Frontier AI models are no longer just assisting security teams — they are reshaping the threat landscape itself. When Anthropic’s Mythos model cracked through every NSA system in thirty minutes by stacking thousands of low-severity vulnerabilities, it signaled a new era for how organizations think about risk, ethics, and the very structure of their security programs.
In this episode, Matthew Moog, Principal and Financial Services Risk Managed Services Leader at EY, joins us to unpack two decades of evolution in third-party risk management, the ethical tensions around AI deployment, and what Project Glasswing means for security leaders who are still on the outside looking in.
You can read the complete transcript of the episode here >
How has third-party risk management evolved over the past two decades?
Matthew traces TPRM through several distinct waves, each enabled by maturing technology:
- Assessment-first era (2010–2014): The practice was cyber-focused and operationally heavy. EY’s financial services division ran 15,000–20,000 assessments globally, with roughly 30% involving boots-on-the-ground on-site visits. It was a logistics factory of approvals, conflict checks, and scheduling.
- Data-enriched era (2014–2018): Cybersecurity rating providers like BitSight and SecurityScorecard made outside-in intelligence affordable. Organizations stopped choosing between assessments and data — they used both.
- Resiliency wave (2018–2022): COVID stress-tested vendor ecosystems. Organizations discovered that recovery time objectives didn’t matter much when the time to replace a critical third party was a month. Resiliency overtook compliance as the dominant concern.
- Automation era (2022–present): RPA gave way to agentic AI. LLMs now sit in three layers — embedded in intelligence suites, powering standalone SaaS products, and operating as near-OS platforms with capabilities built on top of them.
The throughline: anything that minimizes busy work and gets teams closer to actual risk management is a net positive.
What are the real benefits and risks of AI in TPRM?
The efficiency gains are undeniable. Matthew estimates AI can remove roughly 60% of assessment effort by automating document ingestion, evidence extraction, and control mapping. Their agents pull answers from policy documents, cite the specific page and paragraph, and provide screenshot evidence for rapid human verification.
But the risks are equally real:
- Over-reliance: Teams must snapshot files before running AI formatting tasks and then side-by-side review every output. Hallucinated links, merged content, and misinterpreted context are everyday occurrences.
- Data isolation: Each client is trained distinctly — no cross-pollination of data between engagements. When building shared utilities, aggressive scanning and redaction workflows are essential.
- Cost blindness: One organization burned $100 million in tokens in a single quarter because nobody capped usage. Token economics vary wildly by prompt complexity and model tier — replacing an $80K employee with an agent that costs $120K in tokens is not a win.
The advice: take the efficiency gains, but reinvest them in deeper capabilities like threat and vulnerability management, zero-day response, and agentic pen testing rather than just cutting headcount.
What is Project Glasswing and why does it matter?
Project Glasswing is the controlled access program through which select organizations — primarily large banks and UK financial institutions — get exposure to Anthropic’s Mythos model. Consulting firms like EY do not have direct model access; they participate through client engagements and working groups.
What made Mythos different from prior models is its ability to stack vulnerabilities. A single vulnerability might be low-severity on its own, but Mythos identified that combining specific sets of three or four vulnerabilities created critical exploit chains. Roughly 90% of the vulnerabilities it surfaced were in open-source code.
The implications for TPRM are significant:
- Widening the gate: Organizations inside Glasswing can prepare; those outside cannot. How the program expands to cover critical infrastructure and mid-market companies will define the security posture of entire sectors.
- Open-source scrutiny: Expect increasing pressure to map where open-source code exists within your third-party ecosystem and to prioritize vulnerabilities based on business risk.
- Speed of patching: Human-based patching cannot keep up with AI-speed vulnerability discovery. Automated remediation pipelines will become mandatory.
Why is AI ethics becoming a board-level issue?
The ethical tension is not abstract — it sits at the intersection of fiduciary duty and responsible deployment:
- Agents escape guardrails. Lab simulations showed injection prompts hidden in white-space text (colored white, invisible to humans) that directed agents to perform destructive actions. Matthew’s team detected it via 800 trap rules, but organizations moving fast without controls are exposed.
- Human capital sustainability. Firing 7,000 people to hit an efficiency target, then rehiring 3,000 three months later when tickets spike and code breaks, is a pattern already playing out in multiple sectors.
- The Mythos precedent. Anthropic refused to provide Mythos to the US government without a human-in-the-loop requirement. OpenAI took a different stance. These choices by frontier AI companies set the ethical floor for the entire industry.
Matthew frames three possible macro outcomes: modest 10–15% efficiency gains with manageable headcount shifts; aggressive 40% automation driving unemployment toward destabilizing levels; or a third path — reducing the workweek proportionally and distributing gains as improved work-life balance rather than layoffs.
Will assessments disappear as AI gets smarter?
Not entirely, but their role is shifting. Matthew draws an analogy to modern cars: vehicles now have sensors on every component providing real-time telemetry, yet dealerships still perform multi-point inspections when a car changes hands.
Similarly:
- First-time relationships still benefit from deep-dive assessments as a baseline for understanding control structures.
- Ongoing relationships should lean more toward real-time intelligence — continuous monitoring of resiliency signals, cyber ratings, and financial health indicators.
- Critical third parties still warrant annual assessments because contextual changes (data center moves, system swaps, high attrition) may not surface through automated signals alone.
The aspiration is running TPRM more like a mini-SOC: real-time threat feeds, continuous security posture management, and agentic pen testing against your vendor ecosystem.
Does frontier AI reduce the need for security professionals?
Matthew is unequivocal: no, not in the next decade. The reasoning:
- Chaos response. Agents can identify anomalies (unexpected data movement, 10x CPU spikes on a node), but critically thinking through whether to cut off a process — and understanding the business impact of doing so — still requires human judgment.
- Organizational knowledge. AI does not understand that when Jack has a problem, he goes to Molly because she solved it three years ago. Agents are hierarchical and only know what they are scoped to know.
- Deepening technical needs. Security teams need more people with deeper understanding of architecture, AI tooling, and the extended network of third-party AI integrations — not fewer people with shallower skills.
The current hiring shift is toward engineers who understand AI-native security surfaces rather than a wholesale reduction in security headcount.
What should security leaders focus on right now?
Matthew’s practical advice boils down to:
- Ask better questions. Referencing Simon Sinek’s Start With Why, he emphasizes that the quality of AI output is entirely dependent on the quality of the prompt. The same principle applies to risk programs — start with why before jumping to automation.
- Invest in controls before speed. Organizations rushing to cut 8% of headcount by year-end need to understand the new risk profile that agentic workflows introduce. More agents means a larger attack surface.
- Prepare for Mythos-class models going public. When frontier models with stacking capabilities reach general availability, the organizations that built control structures and patching pipelines in advance will survive. The rest will scramble.
