Security & Assurance/AI Red Teaming

Find how your AI can be abused before launch.

We attack your AI the way a real user would, find how it can be manipulated, leak sensitive information or misuse connected tools, and show what needs to be fixed before launch.

How does an AI agent test for security vulnerabilities?

Agile Labs tests how the AI agent responds to malicious prompts, untrusted content and attempts to bypass its controls. We test whether an attacker can access sensitive information, misuse connected tools or make the agent take actions it should not.

Confirmed vulnerabilities are documented with reproducible attacks, fixes and tests that can be run again after remediation.

12 of 12Published prompt-injection and jailbreak defences bypassed by adaptive attacks, most above 90%.Frontier-lab joint study, Oct 2025
Zero clicksData exfiltrated from a production assistant by a planted email, past a deployed classifier.CVE-2025-32711
Over 70%Attack success for instructions hidden in a tool’s own description, with model refusal under 3%.Tool-poisoning research, 2025

ULTHEALTH

Every confirmed attack closed, then kept closed.

An AI health application tested for what it could be made to do, then engineered so each confirmed attack was closed and kept closed.

  • Prompt injection
  • Tool abuse
  • Data exfiltration
  • Reproduction
Read the case study
AI-generated guidance drawn from a private health record

How we find what the AI can be made to do

We model where the system is exposed, then attack it the way a real attacker would, and report only what we can make happen twice.

Two is safe, three is exposed

An assistant is exposed when it combines access to private data, untrusted content and a way to send information out. Any two may be safe. All three together give a hostile instruction a path through the system.

The attack may never touch the chat box

Prompt injection does not have to come from a user. It can sit inside a PDF, email, ticket, webpage or calendar invitation the AI later reads.

Not what it says, what it does

The question is whether the model can be persuaded to call a tool, retrieve private information or act using someone else’s access, not whether it can be made to say something it should not.

Reproducible, or not reported

A confirmed finding carries the attack, the transcript, a severity, how often it worked and the fix. Findings are mapped to the OWASP Top 10 for LLM Applications and MITRE ATLAS so the attack class is clear.

Adversarial Testing Report

View the specimen report

The process and timeline

From defining what we can attack to verifying the fixes.

1AUTHORISATIONBefore day one

Set the boundaries

Agree in writing which systems we can test, what we can attempt and where testing must stop.

Written authorisationScope and permitted tests
Blast radiusAgreed before anything runs
Rollback planAnd a named contact
2THREAT MODEL2–3 days

Map what the AI can reach

Establish what data, tools and systems the AI can access, and which actions it can take.

Private dataWhat the model can read
Untrusted contentDocuments, email, tickets
Tools and functionsWhat it can call
Outbound pathsHow data could leave
3THE ATTACK1–2 weeks

Attack the system

Use automated attacks first, then test by hand for prompt injection, data exposure, tool abuse and other weaknesses.

Automated sweepThousands of known strings
Retrieval pathInstructions planted in a document
Tool abuseWhose credentials it uses
Multi-turnAttacks across many messages
4THE FINDINGS3–5 days

Prove every finding

Report only confirmed vulnerabilities, with the attack, evidence, severity, repeatability and required fix.

FindingSeverityRepeats
No provenance separationCriticalStructural
Document exfiltrates dataCritical18 / 20
Tool called with forged fieldsHigh14 / 20
System prompt recoveredMedium6 / 20
5RETESTAfter the fixes land

Test the fixes

Re-run confirmed attacks after remediation and turn them into regression tests for future releases.

Fixes land
Attacks re-tested
Cases wired into CI
Regression suiteRuns on every release

More on adversarial testing

What organisations need to know about testing AI systems against attack.

Why a system prompt is not a defence

Twelve published prompt-injection and jailbreak defences were bypassed by adaptive attacks in a single frontier-lab study, most above ninety per cent. A filter is a control with a false-negative rate, not a boundary.

Read more

The attack that never touches the chat box

Instructions planted in a document the system will later retrieve. The user asks an ordinary question and the planted text runs, which is why the guardrail never fires.

Read more

What a guardrail actually catches

Published detector accuracy runs from about sixty per cent to the high nineties depending entirely on the corpus tested, which is why the number that matters is measured on the client’s own traffic.

Read more

When organisations call us

Call us when

An AI assistant or agent is about to go live, and the product or engineering owner wants independent security testing before people depend on it.The system reads documents, email, tickets or other untrusted content while also having access to private information, APIs or tools that can perform actions.An assistant or agent is already in use and the team needs to know what an attacker or malicious user could make it reveal or do.A customer security review, board, auditor or regulator wants evidence that the system has been tested for prompt injection, data exfiltration, tool abuse and other AI-specific attacks.

Not us when

The system only generates content and has no access to private data, retrieval sources or tools, where the main concern is the quality or appropriateness of its answers.A prototype that is not intended for real users — the test belongs to the version that will face real users.

Find the security gaps that matter.

Find out what your AI can be made to do before someone else does.