Case Studies
AI Reliability in Practice
Real engagements. Anonymised where required. Each case study documents the problem, what we did, and the measurable outcome.
AI Auto-Reply App — 4 Critical Findings Uncovered Before Launch
We ran our full 5-layer assessment on an AI auto-reply app using our evaluation tooling. The AI pipeline had no injection sanitization, no scope restriction in default mode, and no business knowledge grounding. Four critical findings. Two layers passed. Documented with payloads and model responses.
4
Critical findings with documented evidence
63%
Risk score across 70 test cases
2 of 5
Layers passed (Context, Output)
4 Critical Findings in a Healthcare RAG Chatbot — Before It Shipped
A clinical AI team engaged us one week before their patient-facing RAG chatbot was due to go live. Our 374-vector red team found 4 critical findings — including PHI leakage via retrieval manipulation — before a single patient used the system.
4
Critical findings resolved pre-launch
48hr
Full report with evidence delivered
0
Patient data incidents post-launch
Catching Silent Model Drift Before It Hit a Quarterly Risk Report
A fintech team running AI-assisted loan decisioning had no production monitoring in place when their LLM provider pushed a silent update. We detected a 12% quality drift within 48 hours — a change that would have gone unnoticed for weeks without instrumentation.
12%
Quality drift detected in 48hr
24hr
Time from alert to rollback decision
14
Consecutive on-baseline weeks since
EU AI Act Readiness in 5 Days for a B2B SaaS AI Assistant
A B2B SaaS company preparing for EU market expansion needed documented AI governance compliance before their legal counsel could confirm EU AI Act readiness. We delivered a full gap analysis and prioritised remediation roadmap in 5 working days.
8
Compliance gaps identified and prioritised
5 days
Full audit and roadmap delivered
4 wks
Ahead of legal sign-off deadline