AI.GuardAI.Guard

Continuous red teaming

A static security review goes stale the moment a model is upgraded. AI.Guard runs continuous, scheduled adversarial probes against your endpoints, scores their resistance, and detects regression as upstream models change underneath you.

Attack catalog & generation

The catalog is continuously updated from published research, observed in-the-wild attacks, and findings from customer pilots. Base attacks (like bias, PII extraction, and toxicity attacks) are generated synthetically, then modified and expanded using multiple strategies for hardening the attacking prompts. Once the LLM answers all red-teaming prompts, its responses are evaluated against metrics. Responses that fall below the defined standards expose the model's weak spots and set the direction for further hardening and defense.

Attack familyRepresentative examplesSeverityDetection approach
Direct prompt injectionDAN, role-override, instruction-leak, system-prompt extractionCriticalClassifier + structural heuristics
Indirect injection via RAGPoisoned document, hidden HTML, encoded instructions in retrieved contentCriticalSource-tagged content boundary checks
Jailbreak variantsRoleplay framings, hypothetical reframings, low-resource language attacksHighMultilingual classifier ensemble
Tool / function-call abuseCoercing tool calls with unsafe arguments, parameter smugglingHighSchema validation + intent scoring on arguments
Training-data extractionDivergence prompts, repetition attacks, completion-leak probesMediumOutput entropy + memorization fingerprints
Data exfiltration via outputEncoded PII, side-channel formatting, markdown image beaconsHighOutput DLP + outbound link policy
Denial-of-wallet / cost abuseLong context floods, model-shopping, retry stormsMediumToken quotas + per-tenant rate limits

Scheduled probe cadence

Probes run against your staging endpoints continuously and against production on a controlled schedule with synthetic-traffic markers, so real users are never affected and findings never pollute your business metrics.

Day 0

1. Baseline

Catalog every reachable LLM endpoint, capture current policy posture, and run an initial 1,200-probe adversarial sweep across staging.

Hourly

2. Continuous probes

Lightweight scheduled attacks against non-prod traffic. Detect drift after model upgrades or prompt-template changes within minutes.

Weekly

3. Deep campaigns

Full attack-catalog sweep with multilingual, indirect-injection, and tool-abuse variants. Compares pass-rate against baseline.

Monthly

4. Drift & report

Score deltas, regressions, and policy-tuning recommendations delivered as a signed report. Tickets opened automatically for new gaps.

Scoring framework

Each endpoint receives a per-family resistance score plus an aggregate posture grade. Scores are weighted by severity and by realistic attacker access.

Resistance score

Per attack family, the share of probes correctly blocked or safely neutralized. Reported with confidence intervals.

Drift detection

Any drop in resistance after an upstream model revision, prompt template change, or policy update is flagged within the next probe cycle.

Signed reports

Monthly artifact suitable for board / regulator review. Cryptographically signed, references the exact policy and model versions tested.

Sample report excerpt

A stylized excerpt from a monthly assurance report. Real reports include per-app breakdowns, regression deltas, and remediation guidance.

Assurance report
acme-prod-eu · cycle 2026-04
signed · sha256:7b3c…a91
Aggregate posture gradeA-
Direct injection98.4% +0.2%
Indirect injection (RAG)94.1% -1.6%
Jailbreak variants96.7% +0.0%
Tool-call abuse99.0% +0.1%
Training-data extraction100% +0.0%
Output exfiltration97.3% -0.4%
Notable regression

Indirect-injection score dropped 1.6 pts after the 2026-04-12 upstream model revision. Recommended action: enable the rag.boundary.strict rule for affected apps. Estimated recovery: +1.4 pts.