Key Takeaways
- Epic Systems used Anthropic’s Claude Mythos to uncover configurations that could expose patient records without generating audit-trail entries.
- The exercise shows how AI agents can support security testing while creating new questions about access controls, supervision, and evidence retention.
- Healthcare customers will increasingly expect AI vendors to document testing boundaries, remediation processes, and oversight practices.
Epic Systems has used Anthropic’s Claude Mythos AI agent to stress-test its medical-records systems, exposing configurations that could allow access to patient information without creating an audit-trail entry, according to The New York Times.
That finding matters because audit trails are a central part of healthcare security. They help hospitals determine who accessed a record, what information was viewed, and whether activity was authorized. A configuration that permits access without leaving that evidence could complicate incident investigations, compliance reviews, and internal monitoring.
The scale raises the stakes. Epic reportedly maintains records for approximately 325 million patients worldwide. A security weakness in technology operating at that level can create risk across hospitals, clinics, health plans, and the partners connecting to their systems.
Claude Mythos was deployed defensively. Rather than waiting for a human tester to identify every potentially dangerous configuration, Epic gave the agent a role in probing systems for weaknesses. AI agents can explore combinations of permissions, workflows, and settings at a speed that may be difficult to reproduce through manual testing alone.
Deploying AI agents for security testing raises new questions about oversight. Enterprises considering this approach need clear boundaries around what an AI security agent can inspect, which actions it can take, and what data it can retain. Testing should occur in controlled environments where possible, with production access limited and independently logged. If the agent’s own activity is not captured, the testing process could reproduce the visibility gap it is intended to find.
Healthcare AI adoption is expanding, though deployments during 2025 remained concentrated in relatively lower-risk areas, including ambient documentation, medical coding, and administrative automation. Epic and Microsoft were among the platforms most frequently used or considered.
Moving from documentation support to autonomous security testing represents a shift in risk profiles. While clinical notes and billing workflows carry substantial privacy concerns, an agent tasked with finding exploitable system paths requires unusually broad system visibility. That access can be useful, but it can also become a target or a source of unintended behavior if credentials, tools, or prompts are poorly controlled.
Data fragmentation complicates automated testing. ISG reported that 58% of healthcare organizations identified data usability as their leading data-and-AI concern, while 40% cited data integration difficulties. Medical information often spans electronic health records, imaging platforms, claims systems, laboratory applications, and third-party services. An AI agent testing across those environments may encounter inconsistent permissions and logging practices.
However, this same architectural complexity strengthens the case for automated testing. Human security teams often struggle to examine every interaction among applications, interfaces, user roles, and vendor connections. An agent can systematically search that expanding attack surface, provided Epic and its customers verify findings rather than treating model output as conclusive.
The broader market is heading toward deeper AI integration. Mordor Intelligence tracks AI adoption across healthcare information systems, reflecting growing commercial interest in applying models to operational and clinical data. As those capabilities become embedded in core platforms, security evaluation will need to cover both conventional software flaws and agent-specific risks, including excessive permissions, prompt manipulation, data leakage, and unapproved tool use.
For Epic customers, procurement reviews may now extend beyond model accuracy and productivity claims. Health systems can ask how Claude Mythos was isolated, whether test data included protected health information, how discovered weaknesses were prioritized, and whether remediation was independently validated. They can also request records showing when agents accessed systems and which actions followed.
Epic’s experiment points toward a practical use for agentic AI, but not a hands-off one. AI can help find obscure weaknesses before attackers do. The durable value will come from combining that speed with constrained access, human review, reliable logging, and evidence that identified problems were actually fixed.
⬇️