Security & Assurance/AI Monitoring & Response

Know when your AI changes after it goes live.

We continuously monitor usage, cost, security and answer quality, investigate when something changes and report what needs attention each month.

How does Agile Labs monitor an AI system in production?

Agile Labs monitors every model call to track usage, cost, security events and answer quality. When something moves outside thresholds, we flag it and investigate what changed.

Each month, we report what happened, what needs attention and what action was taken.

US$47,000The bill from an agent loop that ran unwatched for eleven days.Published incident postmortem, 2025
US$14,000Model calls billed to a stolen key in a single day.Published incident postmortem, 2025
ThreeServing faults one AI lab admitted had degraded output with no version change.Provider postmortem, 2025

Lumimory

The AI is observable after it ships.

Lumimory records how its AI is being used and what happens around each request, giving the team a traceable record when something needs to be investigated.

  • Monitoring
  • Cost control
  • Incident response
  • Audit log
Read the Lumimory story
Lumimory AI audit log

Four key signals we monitor

We focus on four signals that show when the AI system is behaving differently from what was agreed.

Spend

Track the cost of every model call and flag unusual increases as they happen, rather than waiting for the provider’s bill.

Security events

Track refusals, jailbreak attempts and known attack patterns. Each alert has a named person responsible for investigating and responding.

System changes

Track changes to models, tools, prompts, data sources and providers against the approved inventory. Model deprecations are tracked at least ninety days ahead so a provider change does not become an emergency.

Answer quality

Run the same graded questions every week to detect changes in performance. Real failures become new test cases, with manual attacks repeated every quarter.

Monthly AI Oversight Report

View the specimen report

The process and timeline

From making every AI interaction visible to knowing when something changes and what happened.

1FIRST MONTH

Make every model call visible

Route AI traffic through one control point to monitor usage, cost, security events and answer quality.

Virtual keysOne per team
Budgets and limitsSet at the key
Retry capsA loop fails closed
RedactionBefore anything is stored
2CONTINUOUS

Detect changes as they happen

Monitor agreed signals continuously and alert the responsible person when a threshold is crossed.

Spend deviation51× the trailing averagePaged
Named on-callClaims team, business hours
Agreed actionCap the key, then investigate
3EVERY WEEK

Check answer quality

Run the same graded questions each week to detect changes in how the AI performs.

WeekScore
Week 195.5%
Week 295.2%
Week 391.4%
Week 495.1%
4EVERY QUARTER

Test the response

Run the incident process and attack the system again to verify the controls and response still work.

Switch a model off
Take the fallback route
Check the evidence trail
Re-attackA person, not the suite
5EVERY MONTH

Report what happened

Report AI usage, cost, security events, answer quality, incidents and changes over the month.

What ran, and what appeared
Spend against the baseline
What was blocked
What changed upstream
One decisionSo the meeting is a working session

More on AI oversight

What organisations need to know about running AI in production.

Why a monthly spend cap will not catch a runaway agent

A retry loop running at fifty times the average does its damage in minutes. A ceiling notices at the end of the month, which is why alarms fire on deviation from a trailing baseline instead.

Read more

How to tell when a provider changed the model

Providers revise models under stable API names. One lab has published a postmortem admitting three serving faults that degraded output with no version change. Asking the same questions weekly is the only way to see it.

Read more

Who gets paged when an AI system misbehaves

An alert nobody acts on trains everyone to ignore the next one. Every threshold names a person, their hours and the action before it is switched on.

Read more

When organisations call us

Call us when

The AI bill moved and nobody can explain whyModel spend has increased unexpectedly, but the provider invoice arrives after the activity that caused it and there is no useful per-team record.A complaint arrived before the metric didA user found a bad answer or unexpected behaviour before the organisation’s own monitoring detected a change in quality.A customer asks who watches the AIA customer security review asks how AI use is monitored, who responds to alerts and what evidence is kept.The board or regulator expects continuous oversightAn annual assessment is no longer enough. The organisation needs an ongoing record showing how the AI system is behaving and what happens when something changes.The AI reached production before anyone owned itThe system is already being used, but responsibility for cost, quality, attacks, model changes and incident response has never been made explicit.

Not us when

The organisation needs a one-off assessment of which AI tools are in use, what data they can reach and where the exposure sits. That is an AI Exposure Assessment.
Read about AI Exposure Assessment →

Find the security gaps that matter.

Know when your AI starts behaving differently and what to do about it.