This is Part 3 of a five-part series on cybersecurity, AI governance, and the future of biological research. Part 1 laid out the governance gap and why it isn't new, and Part 2 put documented names, dates, and numbers on it. This part turns from what has already happened to where the technology, and the risk, is headed next. Part 4 lays out what to build in response, and Part 5 looks at how Senserva helps close the gap at the infrastructure layer.
Previously in this series: Part 1 introduced the three-layer governance gap, infrastructure, data, and AI model, and traced its history through Asilomar and the Human Genome Project. Part 2 backed that framing with documented cases: a generative model that produced a nerve agent candidate list in under six hours, a proxy variable that quietly encoded racial bias into a widely used healthcare algorithm, and oncology models that scored well internally and then failed once they left the environment that trained them.
The autonomous research agent is no longer a research project
On June 30, 2026, Anthropic launched Claude Science, a research workbench designed to support scientific work the same way Claude Code supports software engineering. It can autonomously execute meaningful scientific work from high-level instructions, interface with more than sixty scientific databases, and maintain a traceable record of how every result was produced (MIT Technology Review, June 30, 2026).
Claude Science is not an isolated case, and that matters for the governance argument this series is making. The same basic pattern, an AI system that plans, queries scientific databases, runs analysis, and hands off tasks with minimal human intervention, is showing up across the industry. Google DeepMind's AI co-scientist, built on Gemini, is designed to generate and refine novel research hypotheses alongside biomedical researchers. Microsoft Discovery chains together data retrieval, simulation, and hypothesis generation for materials science and chemistry research rather than answering one prompt at a time. FutureHouse, a nonprofit focused on building an AI Scientist for biology, has built a suite of research agents, including PaperQA for literature synthesis, aimed at the same kind of autonomous research loop. Sakana AI's AI Scientist project has pursued an even more end-to-end version of this idea, attempting to automate hypothesis generation, experimentation, and paper writing in one pipeline. Narrower, task-specific tools like DeepMind's AlphaFold, AlphaProteo, and GNoME predict protein structures, design new proteins, and discover new materials respectively, and often get folded into this same conversation even though they are not general-purpose research agents. Epic, whose EHR many research-affiliated health systems already run, is building toward the same trajectory outside the research bench: its 2026 roadmap introduced Agent Factory, an AI agent-building platform shipping with 120 out-of-the-box agents, and previewed a future version of its Analyst Build Assistant tool that moves from recommending configuration changes to making routine changes, troubleshooting, and self-testing on its own (Diaz, 2026). The pattern this series has been tracking in dedicated research tools is arriving inside general health IT infrastructure too.
The point is not to rank these tools against each other. It is that the trajectory is the same everywhere: more autonomy, more tool access, less human review by default. A research organization evaluating any of these platforms should be asking the same chain-of-custody questions raised in Part 1 of this series, not treating vendor selection as a substitute for its own governance work.
Agentic AI is a new attack surface, and it is arriving faster than access-review cycles
Much of the research tooling entering biology labs right now is not a single model answering a single prompt. It is agentic: systems that plan, call tools, query databases, execute code, and hand tasks off to other agents, with persistent memory carried across steps. A recent review of the agentic AI literature is direct about what this means for security: giving an agent access to external tools and the ability to execute code "introduces significant security risks, which can lead to severe vulnerabilities," including command injection attacks that trigger unintended actions or expose sensitive information (Peykani et al., 2026). Multi-agent systems compound the problem, because a compromised agent can pass malicious instructions to other agents "like a virus, even when all communications are not globally shared" (Peykani et al., 2026).
None of this is exotic. It is the same access control and audit logic that already applies to service accounts and integrations in a Microsoft 365 tenant. The difference is that an autonomous research agent with database access, code execution, and the ability to call other tools is a more consequential thing to leave under-governed than a stale app integration, and it is showing up in research environments faster than most organizations' access review cycles can keep pace with. That lag is not hypothetical. At its 2026 Users Group Meeting, Epic disclosed that it has been a heavy user of Anthropic's restricted-access "Mythos" model, made available through the Glasswing program, to scan its own several-hundred-million-line codebase for vulnerabilities its own developers had missed, and said outright that the old 30-day patching standard "is no longer good enough" (Diaz, 2026). That is the vendor behind the EHR most research-affiliated health systems already run, admitting its own patch cadence cannot keep up with what AI-assisted discovery now finds, the same access-review lag this section describes, one layer further down the stack.
The same AI governance gap shows up wherever agentic AI touches direct patient care, not just the research bench. A 2026 review of AI agents across the healthcare system found that the leading documented failure modes are diagnostic hallucination in rare or ambiguous cases, a lack of interpretability that makes it hard for clinicians to trace how an agent reached a recommendation, which matters clinically because a recommendation that cannot be explained is difficult to act on or defend to a patient, whatever else it gets right, unresolved ambiguity over who is accountable when an agent's recommendation is wrong, and training data that underrepresents specific demographic groups, which degrades performance for those populations and can produce inequitable outcomes (Zhao et al., 2026). None of those four failure modes are unique to hospitals. They describe exactly the risk profile of a research agent with database access and tool-calling privileges, just wearing a clinical coat instead of a lab coat. Regulators are starting to build dedicated evaluation sandboxes for this problem, including the UK's MHRA "AI Airlock" and the EU's CORE-MD framework, both of which score AI systems on safety, effectiveness, and equity before they touch a real patient. In practice, AI Airlock runs real AI-based medical products through staged cohorts, pairing developers with regulators, the NHS, and academia in workshops that surface novel regulatory questions before a product reaches a formal approval decision (MHRA, 2024); CORE-MD instead assigns a composite score across a product's clinical association, technical performance, and clinical performance, where a higher score triggers more extensive pre-market clinical-evidence requirements (Rademakers et al., 2025). A research organization deploying agentic tools against biological data does not have an equivalent external check yet, which is exactly why internal AI governance has to fill that gap in the meantime.
The next frontier: when the storage medium is the biology itself
An emerging frontier makes the data governance question even harder to reason about. This series uses data governance for questions about who owns, controls, and can consent to the underlying information, and AI governance for questions about how a model trained on that information behaves and who is accountable for its outputs; biocomputing complicates the first far more than the second, since here it is the information itself, not a model interpreting it, whose ownership and consent status becomes unstable. Biological computing, or biocomputing, uses DNA and other biological substrates to perform computation and store information rather than simulate it. Some biocomputing architectures store data in bacterial populations, and bacteria replicate. Sirbu and Floridi point out that once information is stored in a self-replicating biological system, ownership questions that traditional data protection law was never built for start to surface: does cell division count as data replication, and does that trigger the same regulatory obligations as copying a file, and what happens to ownership and consent when a genetic mutation alters the stored information during replication, a process that "occurs probabilistically and cannot be fully controlled" (Sirbu and Floridi, 2026)?
Biocomputing is not yet a mainstream part of most research organizations' workflows, but it is a useful stress test for how far current data governance thinking actually extends. The same authors note that molecular computation at scale, where billions of operations happen simultaneously inside a test tube or a bacterial population, breaks the sequential, step-by-step verification methods that traditional software auditing relies on. If your organization's roadmap touches DNA-based data storage, computation, or archival, that is worth flagging now, before a data governance gap becomes a live incident, rather than after.
Regulators are facing a problem they have no precedent for
That absence of precedent has a name in the regulatory science literature. Singh, Paxton, and Auclair describe what they call the "Move 37 Conundrum," named for the AlphaGo move against Lee Sedol that experts called a move no human would make. As AI generates novel biological outcomes with no scientific precedent, regulators, who traditionally work by comparing a new case to established precedent, face applications they have never seen before and have no analog to evaluate them against (Singh, Paxton, and Auclair, 2025). Do they reject the work for lack of precedent, or ask for a human explanation of something a black-box model produced?
A handful of attempts to formalize an answer to exactly this problem have come directly out of the biosecurity risk assessment literature. Frameworks built by AAAS, the FBI, and the UN Interregional Crime and Justice Research Institute, and later adapted by the National Academies of Sciences and by Tucker's work on dual-use governance, try to score a converging technology across factors like adversary type, needed expertise, vulnerability, magnitude of consequence, and governability, on the theory that even a rough, illustrative scoring exercise forces the right questions onto the table before a regulator has to ask them cold (O'Brien and Nelson, 2020). No version of this has full expert consensus, and the authors of the most recent attempt are candid that it does not solve the underlying problem of assessing a technology whose consequences are still unfolding.
Regulators in different jurisdictions are also handling the reality of continuously-updating AI models very differently, which matters for any organization operating across borders and is likely to matter more, not less, as these systems become more autonomous. The U.S. FDA's Predetermined Change Control Plan pathway lets a manufacturer get advance authorization for a defined set of future algorithm updates, so a model can keep learning without a full resubmission every time. The EU's AI Act classifies most medical AI as high-risk and requires ongoing risk management specifically for performance drift and distributional shift, on top of existing device rules. India's Medical Device Rules, by contrast, currently have no equivalent adaptive-update pathway, so a continuously learning model there either stays locked at the point of approval or has to go through a full resubmission for a material change. The U.S. FDA alone has more than 20 separate AI-related initiatives across its offices, and the OECD has tracked over 1,000 AI policy initiatives across 69 countries. Fragmentation, not absence, is increasingly the more immediate governance problem. None of this fragmentation amounts to a mandatory AI-governance audit yet. Voluntary frameworks like ISO/IEC 42001, introduced in Part 1, are likely to keep arriving well ahead of any binding audit requirement in most jurisdictions, though that gap will not hold everywhere: some countries' AI-specific rules may require a formal governance audit well before ISO 42001-style certification becomes anything more than voluntary.
None of this trajectory is slowing down, and none of it is waiting for regulators to agree on how to handle it. That is exactly why the practical AI governance an organization can build internally matters more, not less, as these systems get more capable. Part 4 of this series lays out what that governance actually looks like, with concrete examples of implementation.
About Senserva: Who We Are and What We Do
Senserva is a Microsoft-focused security posture management company based in St. Paul, Minnesota. Its platform continuously scans an organization's Microsoft 365, Intune, Defender, and Entra ID environment, running several hundred built-in checks across identity, privileged access, applications and service principals, endpoint and device compliance, email, logging, and Azure subscription roles.
For every finding, Senserva provides step-by-step remediation guidance, including validated PowerShell scripts and rollback notes, and maps results to recognized compliance frameworks such as the Microsoft Cloud Security Benchmark, CISA's SCuBA baselines, NIST 800-171, CIS, ISO 27001, SOC 2, HIPAA, and PCI-DSS. Senserva is a member of the Microsoft Intelligent Security Association (MISA) and a Microsoft Security Excellence Awards finalist.
Senserva also offers Three Free Unlimited Audits covering every tenant, user, and patch: one to Find, one to Fix, and one to Prove, using the same 650+ checks described above. Registration takes only a work email, with no cost and no obligation to buy. Request an evaluation key at senserva.com/request-key.html.
Senserva also runs Senserva Watch, a free membership that tracks CVEs and Microsoft patches on an organization's behalf and sends an alert only when something being tracked changes, including a same-day notice when Microsoft's Patch Tuesday updates ship. Signing up needs no tenant connection, agent, or call, just an email address, at senserva.com/senserva-watch.html.
Already a Senserva Watch member, or ready to join? Sign up at senserva.com/when-the-lab-meets-the-cloud.html to get the complete five-part series sent to your inbox, plus everything else membership includes: alerts when a CVE or KB you follow changes, a fresh audit credit every quarter, and 20% off when you buy.
Learn more at senserva.com.
Sources cited in this part
- Huckins, Grace. "Claude Science Is Anthropic's Newest Flagship Product." MIT Technology Review, June 30, 2026.
- Diaz, Naomi. "Epic's Roadmap: 12 Takeaways From Judy Faulkner." Becker's Hospital Review, August 18, 2026.
- Peykani, Pejman, Sanly Ghanidel, Iman Javadi-Sisi, Vaclav Snasel, and Seyedali Mirjalili. "A Holistic Review of Agentic AI Frameworks, Applications, and Research Trajectories." Archives of Computational Methods in Engineering (2026).
- Zhao, Lina, et al. "AI agent in healthcare: applications, evaluations, and future directions." npj Artificial Intelligence 2 (2026): 31.
- Medicines and Healthcare products Regulatory Agency (MHRA). "MHRA Launches AI Airlock to Address Challenges for Regulating Medical Devices that Use Artificial Intelligence." GOV.UK, May 9, 2024.
- Rademakers, Frank E., et al. "CORE-MD clinical risk score for regulatory evaluation of artificial intelligence-based medical device software." npj Digital Medicine 8 (2025): 90.
- Sirbu, Renée A., and Luciano Floridi. "An Analysis of the Governance, Ethical, Legal, and Social Implications of Biocomputing." Science and Engineering Ethics 32 (2026): 8.
- Singh, Rominder, Mark Paxton, and Jared Auclair. "Regulating the AI-enabled ecosystem for human therapeutics." Communications Medicine 5 (2025): 181.
- O'Brien, John T., and Cassidy Nelson. "Assessing the Risks Posed by the Convergence of Artificial Intelligence and Biotechnology." Health Security 18, no. 3 (2020): 219-227.