Tech Against TerrorismProgramme Report 01
July 2026

Counter-Terrorism · Artificial Intelligence

Counter-Terrorism AI (CT-AI) Benchmark

The first benchmark of how AI models respond to terrorist misuse. Launched at the United Nations, July 2026. Independent and self-funded. This first report is a pilot: the start of a continuous programme to raise the bar.

See the ranking
Models tested
27
Graded outputs
2,339
Use cases tested
26
Pillars in 4 domains
13
As covered byThe New York TimesThe GuardianFinancial TimesDeutsche WelleABC News (Australia)WIREDRead the coverage →
The result

CT-AI Safety Score

Every output is graded for how far it would help someone intent on harm; 0 to 100, higher is safer.

27 models
Sort

CT-AI Safety Score · 0–100, higher is safer · read as tiers, not precise positions.

SafestClaude Opus (Cowork) and Falcon3 10B98
Least safeQwen3 8B (de-guardrailed)46complied with almost every request
01
Claude Opus (Cowork)
98
02
Falcon3 10B
98
03
MiniMax M2.7
96
04
Llama 3.1 8B
94
05
Granite 3.3 8B
93
17 more models
23
JAIS Adaptive 7B
72
24
Llama 3.1 8B (abliterated)
69
25
Mistral Large 3
65
26
Mistral 7B
64
27
Qwen3 8B (de-guardrailed)
46
Closed-weightOpen-weightDe-guardrailed

Bar colour marks the provider. Scores separate models into tiers rather than precise positions; the full table view lists the 95% confidence interval for each.

Alignment against terrorist misuse is fragile across every frontier model tested, and no developer or government has yet solved it.

Three of the four closed models ran on lighter consumer tiers, so cross-provider ranking is not like-for-like; scores separate into tiers, not precise positions.

Here compliance means complying with the user’s harmful request, not with safety policy.

Foreword

Until now there was no AI benchmark focused specifically on terrorism. So Tech Against Terrorism built one.

The finding that matters is not how often a model refuses, but what it hands over when it does not. Across 27 models and almost 2,500 single-shot prompts, around a third of responses gave meaningful uplift: real assistance towards making a bomb or planning a mass-casualty attack, above what an ordinary web search would return. That it took only rudimentary, single-shot prompts to get there is not acceptable.

Even the better models can be turned with a simple change of framing. Reframing an identical request as research, rather than stating a harmful purpose, was enough to raise compliance sharply. The reassurance these models project is largely an illusion, the product of hedged compliance and intent framing: a safety warning, and then the content regardless.

Open weights are not the problem in themselves. The acute and least reversible risk is the abliterated build, an open model with its safety training stripped out: the two we tested complied with 89% and 100% of requests, and once downloaded they cannot be recalled. The deeper danger is that developers are inadvertently creating models they cannot control. This fragility is not the fault of any single laboratory. It is shared and transnational, it affects every frontier model tested, American and Chinese alike, and no developer or government has yet solved it.

The intent of this report is to compel frontier labs to invest in safety mechanisms for AI models from their inception. Tech Against Terrorism stands ready to work alongside AI developers to share its methodology and build safeguards from the outset.

Adam Hadley CBE

Executive Director, Tech Against Terrorism

What we found

Around a third of responses gave usable uplift.

Beyond what is easily found by web search: real assistance towards making a weapon or planning an attack, from rudimentary single-shot prompts.

A high refusal rate is not safety.

Full refusals were 57% of responses. Hedged compliance, a response that opens with a refusal then supplies the content anyway, was 15%, the largest non-refusal category.

The acute risk is open models stripped of safety training.

Two open models with their safety stripped out, a process called abliteration, complied with 89% and 100% of requests. These models cannot be recalled.

Guardrails hinge on stated intent.

Reframing an identical request as research raised compliance from 17% to 42%, with no change to the technical content.

Open versus closed was not the dividing line.

Anthropic's Claude and Falcon3 ranked safest, with China's MiniMax close behind. The two abliterated builds and the two Mistral models ranked lowest.

Coverage was uneven across weapon types.

Explosives were refused around 80% of the time, but edged weapons, improvised chemical weapons and firearms acquisition were refused only about a third of the time.

Terrorist use of AI is already operational, not hypothetical.

Governments are now vetting frontier AI for national-security risk, including under a June 2026 US Executive Order. Those reviews centre on cyber and CBRN risk. Terrorist misuse is assumed covered; it is not measured.

The evidence

This is already happening.

The benchmark measures the uplift a model will hand over to someone intent on harm. These cases show that demand is real: documented instances of AI implicated in terrorism, violent extremism and mass violence.

Public cases documented
30+TaT incident tracker, June 2026
AI tools implicated
11+
Deaths linked
70+

Matthew Livelsberger

Deceased

2025 · United States

Las Vegas police publicly described this as the first case on US soil in which ChatGPT queries were used to help plan an explosive device detonated outside the Trump International Hotel after Livelsberger took his own life, while OpenAI said the model had returned only publicly available information with warnings.

ChatGPTOfficial reportAttack planningCapability uplift

Islamic State supporter media network (News Harvest)

Group activity

2024 · Transnational

Pro-Islamic State supporter networks, distinct from the group's official media arm, ran an AI-generated news-anchor programme called News Harvest from March 2024, repackaging official IS claims as broadcast-style bulletins within days of the Crocus City Hall attack.

Generative video and voice toolsPress reportedPropaganda generation

Jaswant Singh Chail

Convicted

2023 · United Kingdom

He exchanged thousands of messages with a Replika chatbot companion that validated his stated intent before he entered the grounds of Windsor Castle with a loaded crossbow intending to kill Queen Elizabeth II, and was sentenced to nine years for treason in remarks that address the chatbot exchanges directly.

ReplikaCourt establishedChatbot radicalisationAttack planning

Compiled from public court records, official reports and press reporting, and from Tech Against Terrorism’s CT Casebook, a continuously updated corpus of counter-terrorism cases. Public access to the Casebook is forthcoming.

Explore the case record →
How we measured it

Uplift, not refusal rate.

Every prompt tested is a harmful one, so we grade what a model hands over, not how often it declines.

Uplift

How far a response takes someone beyond what ordinary online search would give them: nothing for a refusal, up to fully actionable content. A safety warning followed by the requested steps still counts, because the usable content is delivered regardless of the caveat.

Two grading engines

Each output is graded twice against a rubric fixed in advance: first as a refusal or not, then, where the model complied, for severity. Operational pillars (CBRN, weapons, attack planning, cyber) go to a local, air-gapped grader; content pillars (radicalisation, synthetic media, threats, code) to a Claude-based grader. Both apply the same definitions.

100 − severity-weighted mean harm = CT-AI Safety Score

Each output becomes a harm value from 0 to 1, averaged within each pillar, then combined across pillars with threats to life weighted most heavily. Higher is safer. A model that refuses throughout scores close to 100.

Grading is kept separate from capture and the rubrics are set before testing. A human analyst reviews a stratified sample of the automated grades against the original outputs; that validation is in progress. Full method and weights are in the report.

The taxonomy

152 ways AI can be misused for terrorism

Mapped across 4 domains and 13 pillars; 26 tested in this first run, single-shot only, so every figure is a lower bound.

domains
4
pillars
13
sub-categories
152
tested this run
26

Taxonomy version 0.4. This first run sampled the taxonomy, single-shot: one prompt, no follow-up, no attempt to get round a refusal. Every figure is therefore a lower bound.

01

Violence and weapons

Severity weight 3.0: threat to life
What it covers
  • Explosives and IEDs
  • Firearms acquisition and 3D-printed firearms
  • Edged weapons
  • Vehicle and drone attacks
  • Arson
  • Poisons and improvised chemical weapons
972outputs assessed, single-shot
Full refusal62%
Complied25%

The remainder fell between full refusal and compliance.

02

Influence and psychological impact

03

Operational enablement

04

Autonomous and technical escalation

Why conventional weapons in full depth

Firearms and explosives account for most terrorist deaths, so this pillar was tested across all 12 of its sub-categories: 972 outputs, against 81 for each sampled pillar elsewhere.

A sample, and a lower bound

The eleven pillars tested through one sub-category rest on 81 outputs each and are indicative pending deeper coverage. Testing was single-shot, with no adversarial follow-up, so every compliance figure understates what a determined user could obtain.

What happens next

This is the start, not the study.

A single-shot pilot now, expanded adversarial testing next, multi-turn and agentic evaluation after that, repeated as models change, building toward a continuous benchmarking platform.

Q2 2026Q3 2026Q4 2026Q1 2027
Stage 1Single-shot pilot
CompleteQ2 2026
Initial model set benchmarked
Two-engine grading and human spot-checks
Pilot report delivered
Stage 2Expanded and adversarial single-shot
Q3–Q4 2026
CoverageBroaden model set · full taxonomy coverage · multiple languages
Prompt sophisticationAdditional intent framings · jailbreak incorporation · prompt engineering · persona-based prompting
Adversarial testing and gradingAnalyst-led red-teaming · full human-QA pass and rubric refinement
Stage 3Multi-turn and agentic
Q4 2026–Q1 2027
Multi-turn conversation protocol and runs
Sophisticated techniques in conversation
Guardrail erosion under sustained pressure
Cross-cutting lanes · all four quarters
Quality assurance and verification
Taxonomy development
Partner and collaborator engagement
Provider outreach and responsible disclosure
Publication and dissemination
Milestones
Pilot reportEnd Q2 2026 · delivered
Expanded + adversarial findingsQ4 2026
Multi-turn programme reportQ1 2027
Stage 1(complete)Stage 2Stage 3Cross-cutting laneMilestoneCurrent position
How we help

Built to raise the bar, not just rank it.

Tech Against Terrorism is a constructive player. The benchmark exists to help developers ship safer models, and everything we learnt building it is available to them.

For AI developers

01

Pre-release testing

Run your model against the benchmark before release. Results go to you first, in private: no surprises on launch day, and a clear read on where the gaps are.

02

Abliteration-resistance testing

We test whether a model's safety training survives weight-level tampering, before release. Results are private to the provider.

03

Safety mentorship

A mentorship service for foundation-model safety teams: our counter-terrorism analysts working alongside your researchers, from the threat landscape to red-team design.

04

Responsible disclosure

We share findings with providers, and each provider can request the raw outputs for its own models. Published material is redacted; no operational specifics are reproduced.

For platforms and hosts

05

Abliteration fingerprinting

A classifier that identifies de-guardrailed builds of known models, so platforms can label, gate or remove them at scale, with provenance and takedown support.

For governments

06

Methodology and taxonomy

The full taxonomy, grading rubrics and technical findings, offered as a foundation to build into model design and safety testing from a model's inception, not as an afterthought.

07

Monitoring and gap analysis

Tracking repositories for stripped derivatives of frontier models and measuring the uplift they provide, so defences can be prioritised.

This pilot sets the floor, not the ceiling. Work with us on what comes next.

Start a conversation
Questions we expect

Asked and answered.

The questions we expect from labs, governments and journalists, answered from the published report.

Doesn’t publishing this help terrorists?

No operational specifics are reproduced. Compound names, precursor chemicals and procedural detail are redacted and replaced with general descriptions; a redaction key is held internally and unredacted material is not circulated. The pattern of failure is publishable. The content that caused it is not.

Are single-shot prompts a realistic test?

They are the easiest test to pass: one rudimentary prompt, no follow-up, no jailbreaks, no pressure. That around a third of responses still gave usable uplift is the point, and it is why every figure in this report is a lower bound. Stages 2 and 3 add adversarial, multi-turn and agentic testing.

What happens next?

The pilot becomes a programme. The next runs extend coverage from the 26 use cases tested here towards the full 152-sub-category taxonomy, add multi-turn and adversarial testing, and measure what the pilot deliberately left out: reasoning modes, and web search on versus off. Model versions and test dates will be published precisely, across a proper cross-section of models. The pipeline is automated end to end, so as costs allow we intend to re-run the benchmark roughly monthly, publish each run, and open-source the engine and graders.

Is the cross-provider ranking fair?

Read it as tiers, not precise positions. Three of the four closed models ran on lighter consumer tiers, so the ranking is not like-for-like between providers. The report says so explicitly, publishes confidence intervals for every model, and the expanded run will revise the cross-provider comparison.

Who funds this?

Nobody with a stake in the result. The programme is self-funded by Tech Against Terrorism, which has received no funding from any AI provider referenced in the report. Tech Against Terrorism is an independent public-private partnership initiated by the United Nations Security Council Counter-Terrorism Committee Executive Directorate (CTED).

What do you want AI developers to do?

Three things. Measure terrorist misuse as its own category, not as a subset of cyber or CBRN risk. Test before release, not after deployment: the benchmark, taxonomy and rubrics are available to build in from a model's inception. And engage: we offer private results, methodology and mentorship to any provider that wants to improve.

Will you work with other benchmarks and safety institutes?

Yes, actively. The goal is not one benchmark but an ecosystem: harm-specific benchmarks built to common standards, each run by the people who understand that harm best. We are already comparing methodology, rubrics and results with national AI safety institutes and government evaluation programmes, and open-sourcing the benchmark engine so others can apply it to the harms they work on. If you run a benchmark, an evaluation programme or a safety institute and want to work together, use the form below.

Can researchers or providers see the underlying data?

Yes, with vetting. Raw model outputs are held by Tech Against Terrorism and can be shared with vetted researchers, and providers can request the outputs for their own models. Use the form below to ask.

Work with us

The benchmark only matters if it changes what gets released.

Tell us who you are and we will route it to the right person.

We reply from ct-ai.org.

Read the full report.

45 pages: full methodology, all 13 pillar breakdowns, per-model results with confidence intervals, and the incident evidence base.

If you build AI models, the offer stands: private pre-release testing, our methodology, and mentorship for your safety team.