Preprint
AEGIS: Adversarial Enforcement and Guardrail Interception for Healthcare AI Agents
Andrew Espira; Umut Baris Basol · 2026-05-07
Zenodo preprint
ORCID: https://orcid.org/0009-0002-9196-8094 · DOI: 10.5281/zenodo.21723990
Abstract
Medical LLM agents remain vulnerable to adversarial input. We present AEGIS, a three-layer inline security proxy for healthcare AI agents combining an ONNX DistilBERT classifier, a locked-down LLM auditor, and HIPAA PHI redaction, with confidence-gated escalation. Across 2,165 samples, Layer 1 achieves 99.49% accuracy and 0.25% attack success rate with 0% PHI leak rate.
Cite
Andrew Espira; Umut Baris Basol (2026). AEGIS: Adversarial Enforcement and Guardrail Interception for Healthcare AI Agents. Zenodo preprint. https://doi.org/10.5281/zenodo.21723990
← All publications