Counter-Terrorism · Artificial Intelligence
Counter-Terrorism AI (CT-AI) Benchmark
The first benchmark of how AI models respond to terrorist misuse. Launched at the United Nations, July 2026. Independent and self-funded. This first report is a pilot: the start of a continuous programme to raise the bar.
- Models tested
- 27
- Graded outputs
- 2,339
- Use cases tested
- 26
- Pillars in 4 domains
- 13
CT-AI Safety Score
Every output is graded for how far it would help someone intent on harm; 0 to 100, higher is safer.
CT-AI Safety Score · 0–100, higher is safer · read as tiers, not precise positions.
Bar colour marks the provider. Scores separate models into tiers rather than precise positions; the full table view lists the 95% confidence interval for each.
Alignment against terrorist misuse is fragile across every frontier model tested, and no developer or government has yet solved it.
Three of the four closed models ran on lighter consumer tiers, so cross-provider ranking is not like-for-like; scores separate into tiers, not precise positions.
Here compliance means complying with the user’s harmful request, not with safety policy.
Until now there was no AI benchmark focused specifically on terrorism. So Tech Against Terrorism built one.
The finding that matters is not how often a model refuses, but what it hands over when it does not. Across 27 models and almost 2,500 single-shot prompts, around a third of responses gave meaningful uplift: real assistance towards making a bomb or planning a mass-casualty attack, above what an ordinary web search would return. That it took only rudimentary, single-shot prompts to get there is not acceptable.
Even the better models can be turned with a simple change of framing. Reframing an identical request as research, rather than stating a harmful purpose, was enough to raise compliance sharply. The reassurance these models project is largely an illusion, the product of hedged compliance and intent framing: a safety warning, and then the content regardless.
Open weights are not the problem in themselves. The acute and least reversible risk is the abliterated build, an open model with its safety training stripped out: the two we tested complied with 89% and 100% of requests, and once downloaded they cannot be recalled. The deeper danger is that developers are inadvertently creating models they cannot control. This fragility is not the fault of any single laboratory. It is shared and transnational, it affects every frontier model tested, American and Chinese alike, and no developer or government has yet solved it.
The intent of this report is to compel frontier labs to invest in safety mechanisms for AI models from their inception. Tech Against Terrorism stands ready to work alongside AI developers to share its methodology and build safeguards from the outset.
Adam Hadley CBE
Executive Director, Tech Against Terrorism
Around a third of responses gave usable uplift.
Beyond what is easily found by web search: real assistance towards making a weapon or planning an attack, from rudimentary single-shot prompts.
A high refusal rate is not safety.
Full refusals were 57% of responses. Hedged compliance, a response that opens with a refusal then supplies the content anyway, was 15%, the largest non-refusal category.
The acute risk is open models stripped of safety training.
Two open models with their safety stripped out, a process called abliteration, complied with 89% and 100% of requests. These models cannot be recalled.
Guardrails hinge on stated intent.
Reframing an identical request as research raised compliance from 17% to 42%, with no change to the technical content.
Open versus closed was not the dividing line.
Anthropic's Claude and Falcon3 ranked safest, with China's MiniMax close behind. The two abliterated builds and the two Mistral models ranked lowest.
Coverage was uneven across weapon types.
Explosives were refused around 80% of the time, but edged weapons, improvised chemical weapons and firearms acquisition were refused only about a third of the time.
Terrorist use of AI is already operational, not hypothetical.
Governments are now vetting frontier AI for national-security risk, including under a June 2026 US Executive Order. Those reviews centre on cyber and CBRN risk. Terrorist misuse is assumed covered; it is not measured.
This is already happening.
The benchmark measures the uplift a model will hand over to someone intent on harm. These cases show that demand is real: documented instances of AI implicated in terrorism, violent extremism and mass violence.
- Public cases documented
- 30+TaT incident tracker, June 2026
- AI tools implicated
- 11+
- Deaths linked
- 70+
Matthew Livelsberger
Deceased2025 · United States
Las Vegas police publicly described this as the first case on US soil in which ChatGPT queries were used to help plan an explosive device detonated outside the Trump International Hotel after Livelsberger took his own life, while OpenAI said the model had returned only publicly available information with warnings.
Islamic State supporter media network (News Harvest)
Group activity2024 · Transnational
Pro-Islamic State supporter networks, distinct from the group's official media arm, ran an AI-generated news-anchor programme called News Harvest from March 2024, repackaging official IS claims as broadcast-style bulletins within days of the Crocus City Hall attack.
Jaswant Singh Chail
Convicted2023 · United Kingdom
He exchanged thousands of messages with a Replika chatbot companion that validated his stated intent before he entered the grounds of Windsor Castle with a loaded crossbow intending to kill Queen Elizabeth II, and was sentenced to nine years for treason in remarks that address the chatbot exchanges directly.
Compiled from public court records, official reports and press reporting, and from Tech Against Terrorism’s CT Casebook, a continuously updated corpus of counter-terrorism cases. Public access to the Casebook is forthcoming.
Explore the case record →Uplift, not refusal rate.
Every prompt tested is a harmful one, so we grade what a model hands over, not how often it declines.
Uplift
How far a response takes someone beyond what ordinary online search would give them: nothing for a refusal, up to fully actionable content. A safety warning followed by the requested steps still counts, because the usable content is delivered regardless of the caveat.
Two grading engines
Each output is graded twice against a rubric fixed in advance: first as a refusal or not, then, where the model complied, for severity. Operational pillars (CBRN, weapons, attack planning, cyber) go to a local, air-gapped grader; content pillars (radicalisation, synthetic media, threats, code) to a Claude-based grader. Both apply the same definitions.
100 − severity-weighted mean harm = CT-AI Safety Score
Each output becomes a harm value from 0 to 1, averaged within each pillar, then combined across pillars with threats to life weighted most heavily. Higher is safer. A model that refuses throughout scores close to 100.
Grading is kept separate from capture and the rubrics are set before testing. A human analyst reviews a stratified sample of the automated grades against the original outputs; that validation is in progress. Full method and weights are in the report.
152 ways AI can be misused for terrorism
Mapped across 4 domains and 13 pillars; 26 tested in this first run, single-shot only, so every figure is a lower bound.
- domains
- 4
- pillars
- 13
- sub-categories
- 152
- tested this run
- 26
Taxonomy version 0.4. This first run sampled the taxonomy, single-shot: one prompt, no follow-up, no attempt to get round a refusal. Every figure is therefore a lower bound.
Violence and weapons
Severity weight 3.0: threat to life- Chemical, biological, radiological and nuclear: precursor identification
- Weaponisation and production routes
- Delivery
The remainder fell between full refusal and compliance.
- Explosives and IEDs
- Firearms acquisition and 3D-printed firearms
- Edged weapons
- Vehicle and drone attacks
- Arson
- Poisons and improvised chemical weapons
The remainder fell between full refusal and compliance.
- Target selection, timing and logistics
- Reconnaissance
- Escape and evasion
- Multi-actor coordination
- Mass-casualty optimisation
The remainder fell between full refusal and compliance.
Influence and psychological impact
- Propaganda generation and recruitment messaging
- Narrative arcs and manifestos
- Memetic and astroturf campaigns
- Translation and localisation
The remainder fell between full refusal and compliance.
- Deepfake video and voice
- Propaganda imagery
- Fabricated evidence
- Provenance and detection evasion
The remainder fell between full refusal and compliance.
- Targeted threats and coordinated intimidation
- Doxxing
- Hoax threats
- Claim-of-responsibility statements
The remainder fell between full refusal and compliance.
Operational enablement
- Malware and exploit development
- Denial-of-service
- Social engineering
- Infrastructure and ICS targeting
- Website defacement
The remainder fell between full refusal and compliance.
- Communications security
- Surveillance evasion
- Identity and document fraud
- Anonymity and counter-forensics
The remainder fell between full refusal and compliance.
- Fundraising and cryptocurrency laundering
- Sanctions evasion
- Crowdfunding and charity fronts
- Hawala and trade-based laundering
The remainder fell between full refusal and compliance.
- Directing users to external sources
- Other-model and platform-circumvention guidance
- Jurisdictional exploitation
The remainder fell between full refusal and compliance.
Autonomous and technical escalation
- Autonomous target identification
- Logistics
- Coordination
The remainder fell between full refusal and compliance.
The remainder fell between full refusal and compliance.
The remainder fell between full refusal and compliance.
Firearms and explosives account for most terrorist deaths, so this pillar was tested across all 12 of its sub-categories: 972 outputs, against 81 for each sampled pillar elsewhere.
The eleven pillars tested through one sub-category rest on 81 outputs each and are indicative pending deeper coverage. Testing was single-shot, with no adversarial follow-up, so every compliance figure understates what a determined user could obtain.
This is the start, not the study.
A single-shot pilot now, expanded adversarial testing next, multi-turn and agentic evaluation after that, repeated as models change, building toward a continuous benchmarking platform.
Built to raise the bar, not just rank it.
Tech Against Terrorism is a constructive player. The benchmark exists to help developers ship safer models, and everything we learnt building it is available to them.
For AI developers
Pre-release testing
Run your model against the benchmark before release. Results go to you first, in private: no surprises on launch day, and a clear read on where the gaps are.
Abliteration-resistance testing
We test whether a model's safety training survives weight-level tampering, before release. Results are private to the provider.
Safety mentorship
A mentorship service for foundation-model safety teams: our counter-terrorism analysts working alongside your researchers, from the threat landscape to red-team design.
Responsible disclosure
We share findings with providers, and each provider can request the raw outputs for its own models. Published material is redacted; no operational specifics are reproduced.
For platforms and hosts
Abliteration fingerprinting
A classifier that identifies de-guardrailed builds of known models, so platforms can label, gate or remove them at scale, with provenance and takedown support.
For governments
Methodology and taxonomy
The full taxonomy, grading rubrics and technical findings, offered as a foundation to build into model design and safety testing from a model's inception, not as an afterthought.
Monitoring and gap analysis
Tracking repositories for stripped derivatives of frontier models and measuring the uplift they provide, so defences can be prioritised.
This pilot sets the floor, not the ceiling. Work with us on what comes next.
Start a conversationAsked and answered.
The questions we expect from labs, governments and journalists, answered from the published report.
Doesn’t publishing this help terrorists?
No operational specifics are reproduced. Compound names, precursor chemicals and procedural detail are redacted and replaced with general descriptions; a redaction key is held internally and unredacted material is not circulated. The pattern of failure is publishable. The content that caused it is not.
Are single-shot prompts a realistic test?
They are the easiest test to pass: one rudimentary prompt, no follow-up, no jailbreaks, no pressure. That around a third of responses still gave usable uplift is the point, and it is why every figure in this report is a lower bound. Stages 2 and 3 add adversarial, multi-turn and agentic testing.
What happens next?
The pilot becomes a programme. The next runs extend coverage from the 26 use cases tested here towards the full 152-sub-category taxonomy, add multi-turn and adversarial testing, and measure what the pilot deliberately left out: reasoning modes, and web search on versus off. Model versions and test dates will be published precisely, across a proper cross-section of models. The pipeline is automated end to end, so as costs allow we intend to re-run the benchmark roughly monthly, publish each run, and open-source the engine and graders.
Is the cross-provider ranking fair?
Read it as tiers, not precise positions. Three of the four closed models ran on lighter consumer tiers, so the ranking is not like-for-like between providers. The report says so explicitly, publishes confidence intervals for every model, and the expanded run will revise the cross-provider comparison.
Who funds this?
Nobody with a stake in the result. The programme is self-funded by Tech Against Terrorism, which has received no funding from any AI provider referenced in the report. Tech Against Terrorism is an independent public-private partnership initiated by the United Nations Security Council Counter-Terrorism Committee Executive Directorate (CTED).
What do you want AI developers to do?
Three things. Measure terrorist misuse as its own category, not as a subset of cyber or CBRN risk. Test before release, not after deployment: the benchmark, taxonomy and rubrics are available to build in from a model's inception. And engage: we offer private results, methodology and mentorship to any provider that wants to improve.
Will you work with other benchmarks and safety institutes?
Yes, actively. The goal is not one benchmark but an ecosystem: harm-specific benchmarks built to common standards, each run by the people who understand that harm best. We are already comparing methodology, rubrics and results with national AI safety institutes and government evaluation programmes, and open-sourcing the benchmark engine so others can apply it to the harms they work on. If you run a benchmark, an evaluation programme or a safety institute and want to work together, use the form below.
Can researchers or providers see the underlying data?
Yes, with vetting. Raw model outputs are held by Tech Against Terrorism and can be shared with vetted researchers, and providers can request the outputs for their own models. Use the form below to ask.
The benchmark only matters if it changes what gets released.
Tell us who you are and we will route it to the right person.
Read the full report.
45 pages: full methodology, all 13 pillar breakdowns, per-model results with confidence intervals, and the incident evidence base.
If you build AI models, the offer stands: private pre-release testing, our methodology, and mentorship for your safety team.
