Preprint

AEGIS: Adversarial Enforcement and Guardrail Interception for Healthcare AI Agents

Andrew Espira; Umut Baris Basol · 2026-05-07
Zenodo preprint
ORCID: https://orcid.org/0009-0002-9196-8094 · DOI: 10.5281/zenodo.21723990

Download PDF Canonical record Code

Abstract

Medical LLM agents remain vulnerable to adversarial input. We present AEGIS, a three-layer inline security proxy for healthcare AI agents combining an ONNX DistilBERT classifier, a locked-down LLM auditor, and HIPAA PHI redaction, with confidence-gated escalation. Across 2,165 samples, Layer 1 achieves 99.49% accuracy and 0.25% attack success rate with 0% PHI leak rate.

Cite

Andrew Espira; Umut Baris Basol (2026). AEGIS: Adversarial Enforcement and Guardrail Interception for Healthcare AI Agents. Zenodo preprint. https://doi.org/10.5281/zenodo.21723990

← All publications