Enforcement catalog
Each inspection layer is independently configurable. Below is what they do — and a live mock of inline redaction in flight.
Inline Inspection — /v1/chat/completions
received
Original — as received
“Summarize this support ticket from Jane Roe (jane.roe@example.com, SSN 000-00-0000) and draft a reply.”
Original values remain visible only inside the protected boundary.
Masked — forwarded upstream
“Summarize this support ticket from [PII_NAME_001] ([PII_EMAIL_001], SSN [PII_SSN_001]) and draft a reply.”
3 reversible tokens · 8.2 ms
Unmasked — returned to caller
“Ticket from Jane Roe (jane.roe@example.com) — suggested reply: …”
Response tokens are rehydrated before the application receives the answer.
policy: pii-strict-v3test data: example.com · never-issued SSNreversible: yes
Prompt-injection & jailbreak defense
- Classifier ensembleMultiple fine-tuned classifiers vote on injection score; ensemble reduces single-model evasion.
- Full attack taxonomyMulti-turn conversations, impersonation, code words and obfuscation, instructions embedded in data, formatting tricks (markdown, HTML), and injection through structured fields.
- Indirect injection coverageRAG context is analyzed; injected instructions inside retrieved documents are scored separately from user intent.
- Multilingual coverageDetection trained on attacks in 18 languages including low-resource ones used to bypass single-language classifiers.
DLP & reversible tokenization
- Detailed masking categoriesPII (names, addresses, IDs), Financial (cards, transactions, debt), Medical (diagnoses, prescriptions), Business (salaries, supplier data), and Documents (SSN, real-estate).
- Reversible tokensSensitive values are replaced with deterministic tokens upstream. The original value never leaves your perimeter.
- Loss-less rehydrationTokens are reversed in the response on the way back so the application sees the original data, transparently.
- Flexible policiesInstant on/off rules, client-specific policies per team/environment, and Natural-Language detection rules (no query language required).
Output filtering & guardrails
- Toxicity, bias, regulated contentConfigurable thresholds per class. Defaults aligned to OWASP LLM Top 10 and NIST AI RMF categories.
- Hallucination guardrails for toolsTool / function-call arguments are validated against your declared JSON schema before being executed.
- Citation enforcementFor RAG applications, responses can be required to cite from the retrieved set; uncited claims are flagged.
- Outbound exfil detectionEncoded data, suspicious markdown image beacons, and base64-blobs in outputs are detected and stripped.
Policy-as-code
- Versioned, reviewed policiesPolicies live in your repo as YAML/Rego. Changes go through PR review; old versions remain queryable in audit.
- Per-app, per-user scopingBind policies to apps, environments, user roles, or upstream provider. Hot-reload propagates in seconds.
- Shadow modeTest new rules against live traffic without enforcing. Observe hit-rate and false-positive cost before promotion.
- Exception workflowJustified bypass requests are logged with caller identity and routed to the AppSec owner for retroactive review.
Quotas & rate limiting
- Per-tenant, per-app, per-userIndependent token, request, and cost budgets with sliding windows.
- Model-aware accountingDifferent budgets per upstream model — cap expensive frontier models per role.
- Burst & smoothingToken-bucket with configurable burst. Friendly 429s with retry-after honored downstream.
- Denial-of-wallet protectionAnomaly detection on call patterns blocks cost-abuse loops before they exhaust budgets.
Content provenance
- Watermark detectionWhere upstream providers emit content watermarks, the gateway can require them on response and flag missing signals.
- Outbound provenanceOptional C2PA-style provenance signing on application-generated content for downstream verification.
- Model & policy attestationEvery response carries a signed header noting model id, policy version, and gateway build for downstream attribution.
- Drift signalingUpstream model identifier changes are surfaced; pin to a specific version and get notified on silent rollouts.

