Research outputs by Andrew Espira
(ORCID 0009-0002-9196-8094).
Preprint
Open-weight LLM judges systematically misclassify partial-compliance refusals as COMPLIANT. Matched-pair ablations establish the executed partial action as a causal driver across four locally served models (7B–14B). In a keyword audit of reasoned competence-probe traces, 124 of 1…
Preprint
As medical AI agents accelerate administrative and clinical workflows, their speed advantage introduces a structural risk: agents can reason incorrectly faster than humans can intervene. We present SENTINEL, a mid-reasoning interception framework that audits agent reasoning trace…
Preprint
Medical LLM agents remain vulnerable to adversarial input. We present AEGIS, a three-layer inline security proxy for healthcare AI agents combining an ONNX DistilBERT classifier, a locked-down LLM auditor, and HIPAA PHI redaction, with confidence-gated escalation. Across 2,165 sa…
Conference presentation
VGAC — Visualize, Gate, Advise, Calibrate — is a calibration-first approach to GPU-cluster queue intelligence. Empirical evidence from Amazon EKS and AWS ParallelCluster Slurm shows that ECE, not AUROC, is the deployment-relevant metric, and that discrimination mostly transfers a…
Conference paper / software artifact
Reproducible scientific-software artifact accompanying the ISS26 VGAC work: calibration-aware pipeline, sample data for four cluster environments, trained calibrators, benchmarks, and a notebook that regenerates every figure.…
Technical note
Andrew Espira · 2026-04-22 · NCAR / UCAR OpenSky Technical Notes
OpenSky archival record for the VGAC / ISS26 GPU-cluster queue intelligence work, hosted by NSF NCAR and UCAR for long-term scholarly access.…