DrochaidHorizon3.ai
NodeZero/AI in offensive security
AI in offensive security

Deterministic vs non-deterministic AI in offensive security

Two kinds of AI are reshaping cybersecurity. Knowing the difference is how you choose the right tool.

In April 2026, Anthropic launched Project Glasswing and disclosed that its unreleased Claude Mythos Preview model had identified more than 10,000 high- and critical-severity vulnerabilities — including zero-days in every major operating system and browser, and flaws that had survived decades of expert human review. Anthropic declined to release the model publicly, warning that AI-augmented offensive capability at this level could proliferate to adversaries, and has briefed senior government officials on the offensive and defensive cyber capabilities the model demonstrated.

Mythos is a landmark demonstration of what probabilistic AI can do in the hands of a well-resourced research team. It is also the clearest signal the market has had that the threat side of the cybersecurity balance has shifted.

In the conversation that has followed, a wave of “AI pentesting” tools is being marketed as the defensive answer. Most of them work very differently to what they appear to do — and more importantly, differently to what NodeZero does. This page is about that difference: deterministic AI versus non-deterministic AI, the legitimate strengths of each, and why choosing the right one for the right job matters more than the marketing suggests.

What "deterministic" and "non-deterministic" actually mean

A deterministic system, given the same inputs, produces the same outputs every time. The logic is explicit. You can trace why it did what it did. You can reproduce the result tomorrow, next week, or in front of an auditor.

A non-deterministic system — which is how large language models like Claude, GPT and Gemini work — produces probabilistic output. Ask it the same question twice and you will get two related but different answers. That variability is not a defect. It is what enables these models to reason by analogy, notice patterns that rigid logic would miss, and generate novel hypotheses. It is also why they are prone to hallucination — confident output that turns out to be wrong.

Both behaviours are useful. The question is where each one belongs.

What probabilistic AI does brilliantly

Project Glasswing is the showcase. Mythos Preview found vulnerabilities that had been hiding in widely used software for decades, surviving millions of automated tests and rounds of expert human review. It did this precisely because it does not follow a script. It reasons creatively across huge bodies of code, notices things that look structurally suspicious in ways nobody had thought to encode as a rule, and chains hypotheses together the way a talented human researcher would — only faster, and at a scale no human team could match.

That creative, generative capability is a genuine advance. For vulnerability research, code review, novel exploit chain discovery, threat intelligence synthesis, and translating technical findings into language executives can act on, probabilistic AI is the right tool. It is the tool Anthropic is using. It is a tool Horizon3.ai itself uses, in scoped roles inside NodeZero, for exactly these kinds of tasks.

What probabilistic AI is not well-suited for — at least not on its own — is making exploitation decisions in someone else’s production network.

Why deterministic matters for offensive security execution

When an autonomous platform is chaining attacks across a live environment, four properties start to matter more than raw creativity.

Reproducibility

If a test finds a critical attack path on Monday and cannot reliably find it on Friday, the result is not evidence — it is an anecdote. Security teams, auditors and boards need to know that the same exposure, tested the same way, produces the same outcome.

Reliability at scale

Independent Carnegie Mellon research found that state-of-the-art LLMs — including Claude Sonnet 3.5, GPT-4o, GPT-4o mini, Gemini 1.5 Pro and Gemini 1.5 Flash, even with advanced prompting frameworks — were unable to carry out end-to-end multi-host intrusions, reaching only 1–30% of attack graph states in labs capped at 50 hosts. Horizon3 cited that finding when it reported NodeZero solving the equivalent Game of Active Directory benchmark in 14 minutes. Creativity alone does not get you across a multi-domain network reliably. Structured reasoning about the terrain does.

Safety in production

Exploits that are generated on the fly by a generative model and fired into a live environment carry real operational risk. Exploits that have been engineered, tested and validated before they ever touch a customer network do not.

Data control

Many LLM-based offensive tools call public model APIs, meaning reconnaissance data, credentials and attack paths leave the environment. For any organisation with serious data handling obligations, that is a problem that creativity cannot offset.

A reasoning-driven architecture

NodeZero is not a large language model dressed up as a pentester. It is also not a rigid rules engine. It is a reasoning-driven architecture — structured to plan, adapt and prove, and built to use the right tool for each task.

Structured to plan, adapt and prove

NodeZero operates like a skilled adversary — guided by goals, shaped by feedback, driven by outcomes. It builds a cyber terrain map of your environment, plans attack paths across it, and makes decisions based on what it sees, not just what it is told.

Precision prompting, not trial-and-error

Unlike generative agents that burn tokens in endless loops, NodeZero generates deep, task-specific prompts from a graph of validated facts. That lets scoped GenAI reason efficiently about business risk, stolen data or user context — instead of flailing toward an answer and hoping it lands.

Multiple agents, one cohesive system

NodeZero is not one model. It is an integrated AI system: graph reasoning powers attack planning, classical machine learning classifies files and behaviours, deterministic logic executes exploitation — every attack module pre-built, pre-validated and tested by Horizon3’s attack team before it ships — and scoped GenAI supports the bounded jobs below.

The principle is simple and explicit: NodeZero never uses GenAI to create or execute exploits. Every action taken against your environment is deterministic, pre-validated and reproducible. The creative, probabilistic work happens where it is appropriate — reasoning about context, meaning and business impact — not where it could produce an unsafe or unverifiable action inside your network.

That discipline carries downstream. Through the NodeZero MCP Server, proven attack paths flow into the tools your teams already use — JIRA, GitHub, Argo, SOAR — so agentic workflows can push remediations, rotate compromised credentials, fire tripwire-triggered playbooks and automatically re-verify that a fix actually closed the path. Risk gets eliminated in production, not reprioritised in a spreadsheet.

Inside that architecture, generative AI is scoped to specific, bounded jobs — never to creating or firing exploits:

High-value targeting

LLMs weigh job roles, access privileges and naming patterns to tag compromised users and systems as high value — reprioritising deeper testing and raising risk scores where it counts.

Exploit Suggester and Try Harder Agent

When NodeZero stalls, scoped GenAI proposes the next-best step — mimicking the persistence of a seasoned red teamer, inside a controlled loop rather than an open-ended one.

Advanced data pilfering

A two-stage approach: classical machine learning identifies the files worth reviewing, then an LLM reads the contents to surface credentials, intellectual property, personal information or financial data an attacker could exploit.

Real-time view chatbot

Ask questions mid-test — “which KEVs are still open in production?” — and get natural-language answers grounded in real-time attack behaviour, not a generic threat feed.

GenAI for web application testing

LLMs analyse modern web applications for logic flaws such as broken access controls — the kind of issue automated scanners routinely miss.

In every case NodeZero runs inference, not training. It does not train foundation models on your data — it builds structured prompts from live findings and runs them against models like Claude, LLaMA or Mistral, choosing the best fit for each task.

The result

You get an autonomous platform that thinks like a skilled adversary — chains weaknesses, pivots, escalates, demonstrates real impact — without the failure modes that make LLM-driven pentesters unsuitable for production use.

Every exploitation attempt is reproducible. Every decision is explainable. Every finding is tied to evidence. The creative work of vulnerability research, executive storytelling and contextual analysis benefits from modern GenAI. The operational work of breaking into your network, proving impact and validating your fixes is deterministic, controlled and safe to run continuously against live systems.

There is a precise line on data, too. Exploitation data never leaves your environment. The bounded GenAI tasks — targeting, narrative writing and data analysis — run only through secure, isolated cloud infrastructure, including Amazon Bedrock, with data residency and isolation controls. That distinction matters: it is not “nothing leaves the network”, it is “the part an attacker would touch stays in your network, and the part that reasons about meaning runs in a controlled, isolated place”.

For organisations under DISP, IRAP, Essential Eight, CPS 234, SOCI/CIRMP or NZISM obligations, that combination — deterministic execution, exploitation data held in your environment, and full command logging — is what produces the reproducible, auditable record a compliance assessment actually asks for.

What Glasswing actually means for your program

Anthropic’s framing of Project Glasswing is precise. Mythos is a starting point. The capability will proliferate. Defenders have months, not years, to change how they operate.

Mythos did not break cybersecurity. It exposed a gap that had been building for years. The industry has become very good at finding vulnerabilities and very poor at determining which of them actually lead to meaningful risk. Vulnerable does not mean exploitable. Exploitable does not mean impactful. And volume-based vulnerability management — patch everything, scan everything, alert on everything — has already outrun most organisations’ ability to remediate.

The shift that matters is from finding vulnerabilities to validating exploitability, from counting issues to understanding impact, and from patching everything to breaking attack paths. That is what NodeZero does continuously, safely, and with evidence you can rely on.

The bottom line

Probabilistic AI is an extraordinary tool for discovery, synthesis and creative reasoning. Deterministic architecture is what you need when an autonomous system is making decisions in your production network. The right answer is not one or the other. It is knowing where each belongs — and choosing a platform that applies that discipline rather than blurring the line for marketing convenience.

That is how NodeZero is built. If you would like to see how it works in practice, we would like to show you.

Trusted by 7,013 organisations worldwide
Powers the NSA's CAPT program310,332 pentests4.7 Gartner

See deterministic, reproducible pentesting in action.

NodeZero thinks like an adversary and proves impact with evidence you can rely on — safely, continuously, against your live environment.