All posts

When the Lab Meets the Cloud, Part 2 of 5: Documented Cases Where the Guardrails Weren't There Yet

Senserva Watch

Join Senserva Watch and get Three Free Unlimited Audits with our full Claude MCP, a fresh audit credit every quarter, alerts when a CVE or KB you follow changes, critical security updates when they land, the Patch Tuesday wire on release day, and 20% off when you buy.

Join Senserva Watch

One email address. No tenant connection, no agent, no call.

This is Part 2 of a five-part series on cybersecurity, AI governance, and the future of biological research. Part 1 laid out why the governance gap between AI-enabled biotechnology and its oversight is dangerous, and why it isn't the first time science has faced this pattern. This part puts names, dates, and numbers on the problem. Part 3 looks at where the underlying technology is headed next, Part 4 lays out what to build in response, and Part 5 looks at how Senserva helps close the gap at the infrastructure layer.

Previously in this series: Part 1 framed the problem as a governance gap running three layers deep, the infrastructure, the data, and the AI model, and showed that science has run this exact experiment before: a detailed, enforceable safety framework came out of Asilomar in 1975, while a similar gathering on AI in 2017 produced principles too broad for anyone to act on. This part puts documented names, dates, and numbers behind that history.

Misuse and discovery share the same underlying capability

A 2024 study used machine learning to identify 161,979 previously unrecognized viruses hidden in public sequence databases, a scale of discovery that would have taken a human team years to approach manually (Hou et al., 2024). The technique is a genuine advance for surveillance and vaccine development. It is also a preview of how much biological signal AI can now extract from data that screening systems were never designed to interpret at that resolution.

The clearest documented case of the misuse side is not hypothetical either. In 2022, a team of researchers took a generative AI model built to predict which drug-like molecules would be safe and effective, and inverted its objective from minimizing toxicity to maximizing it. Running on a laptop, using only a public dataset of drug-like molecules and their known bioactivities, the model generated 40,000 candidate molecules in under six hours. Among them was VX, one of the most lethal nerve agents known, along with numerous other chemical warfare agents and molecules predicted to be even more toxic than existing ones (Urbina et al., 2022). No new lab work was required to produce that list. The researchers simply changed a sign in an optimization function. That case has since become the reference point across the biosecurity literature for how little friction stands between a beneficial AI drug-discovery tool and a weapons-design tool (Trump et al., 2026; Prunkl, 2024).

Current DNA synthesis screening systems compare requested sequences against databases of known biological threat agents. They cannot flag sequences designed by AI to be functionally dangerous while avoiding sequence-level detection. A landmark study published in Science in late 2025 warned that this gap is not hypothetical. Britain's AI Security Institute reported in December 2025 that major language and biology foundation models could reliably generate scientific protocols for synthesizing dangerous viruses when prompted by users with sufficient scientific background (Foreign Affairs Forum, May 2026).

Genome editing carries the same dual-use dilemma in a more specific form. Li et al. describe how the generative AI models now used to design better base editors and Cas protein variants could, in principle, be repurposed for "non-therapeutic purposes, such as human enhancement or biological weapon development" (Li et al., 2025). That means securing the models and trained weights behind a genome-editing pipeline, not just the data feeding it, is itself a biosecurity control.

The risk is not confined to what a model can design, either. It also reaches the infrastructure layer underneath it, the physical machines and devices a lab runs on. Documented attacks against DNA synthesis machines include side-channel techniques that reconstruct the sequence being printed from the machine's own acoustic and electromagnetic emissions, without the attacker ever touching the network (Faezi et al., 2019). Medical devices sitting on the same institutional infrastructure face a parallel threat. Infusion pumps, surgical lasers, and life-support equipment have been documented as vulnerable to hijacking, giving an attacker a foothold through the device itself rather than the network perimeter (TrapX Research Labs, 2018; Peccoud et al., 2018).

Patient privacy failures, documented and projected

For researchers working with human-derived samples and clinical data, a parallel set of vulnerabilities applies. The instinct to de-identify data is correct but incomplete. MIT researcher Latanya Sweeney demonstrated that combining ZIP code, birthdate, and sex was sufficient to re-identify 87 percent of the U.S. population from supposedly anonymized records (Sweeney, 2000).

That finding predates modern generative AI, and the newer risk it points toward is worse. Researchers have since documented PII Replay, where generative models inadvertently memorize and reproduce verbatim strings from training data, making supposedly safe synthetic outputs a vehicle for privacy violations (Gursoy et al., 2024). This risk is specific to the generative architectures now common in genomics, transcriptomics, and clinical natural language processing, and de-identifying the source data does nothing to prevent it once a model has memorized fragments of that data during training.

In 2024 alone, data security breaches exposed the protected health information of more than 242.9 million individuals across 663 reported breaches, according to HHS's Office for Civil Rights (HHS OCR, 2025). The Change Healthcare breach alone accounted for an estimated 190 million of those individuals, per HHS's own accounting of the incident (HHS.gov, 2025). AI adoption by physicians nearly doubled that same year (American Medical Association, 2025), but the workforce side of healthcare has not kept pace: HIMSS's 2024 Healthcare Cybersecurity Survey found that only 18 percent of respondents rated their organization's security awareness training as very effective, and 4 percent reported receiving no training at all (HIMSS, 2025).

That governance gap is no longer a future risk; it is already in litigation. A class action filed in San Diego in November 2025 accuses Sharp HealthCare of deploying an AI ambient-documentation tool, in use since April 2025, that records patient-clinician conversations and auto-drafts the resulting chart notes without the two-party consent California's wiretapping law requires, then sends that audio to a vendor's cloud environment where it reportedly sat retrievable for roughly 30 days before deletion (Fisher Phillips, 2025). The complaint alleges some of the resulting charts documented patient consent that never actually occurred. Strip away the hospital framing, and the underlying fact pattern, an AI tool capturing a sensitive human conversation and handing it to a third-party cloud pipeline whose retention and access controls have not caught up to what the tool now does, applies just as directly to a core facility's sample-intake calls or a clinical-trial consent conversation as it does to an exam room.

When the model itself is the problem

Documented failures are not limited to security breaches. In 2019, an algorithm used across US hospitals to decide which patients with complex medical conditions should be enrolled in extra care-management programs was found to have systematically denied those programs to Black patients who were just as sick as the White patients it did enroll. The cause was a proxy variable: the algorithm used total historical healthcare spending as a stand-in for medical need, and less money had historically been spent treating Black patients at the same level of illness, so the model learned that lower spending meant lower need (Obermeyer et al., 2019). Nothing about the algorithm was biological, but the underlying failure mode, a proxy variable that quietly encodes a demographic disparity from the training data, is exactly what the biological AI literature warns about when it flags cohort bias and data heterogeneity.

Oncology has already run this experiment in the clinic

If those risks sound abstract, oncology AI offers a well-documented preview of what happens when a model looks good on paper and then meets the real world. IBM Watson for Oncology, trained heavily on synthetic scenarios generated at a single institution, showed low and inconsistent agreement with tumor board recommendations once it was deployed internationally, a mismatch driven by differences in regional treatment practices, drug accessibility, and local clinical guidelines that the model never saw during training. It was eventually divested (Saha et al., 2026).

The Epic Sepsis Model, deployed across hundreds of U.S. hospitals, achieved a sensitivity of only 33 percent and an AUC of 0.63 in independent external validation, far below what its developers had reported internally, while also generating enough false alerts to produce genuine alert fatigue among clinicians (Saha et al., 2026).

A systematic review of COVID-19 imaging AI models found that none of the 415 models evaluated were fit for clinical deployment, with data leakage, site-confounded labels, and demographic bias showing up again and again as the root causes (Saha et al., 2026).

The common thread across all three is a problem that AI governance frameworks such as NIST's AI Risk Management Framework and ISO/IEC 42001 group under AI Data Shift, the term this series will use going forward for what the ML literature has historically called dataset shift: the statistical relationship between what a model sees in production and what it saw in training quietly stops holding. Covariate shift is when the incoming data itself changes, for example new imaging hardware or a different patient population than the one the model trained on. Concept shift is subtler and more dangerous: the relationship between the inputs and the correct answer changes, for example because clinical guidelines are updated or a new therapy shifts what a given lab value now means for risk. Concept shift is especially hard to catch because standard input-monitoring tools do not detect it. The only way it shows up is through the outcomes actually getting worse, which means someone has to be watching outcomes, not just watching whether the input data looks normal. A related but distinct failure worth naming separately is AI Model Shift, where the model's own behavior degrades or becomes misaligned with current operations even if the incoming data looks unremarkable, for example because it was never built to generalize past the single institution it trained on, arguably part of what happened to the Epic Sepsis Model above. Either failure carries the same governance requirement: monitoring cannot stop at the moment of deployment.

Every case above shares the same shape: a capability that worked exactly as designed, deployed into a context nobody had checked it against. Part 3 of this series looks at where the underlying technology is headed next, including the agentic systems now entering research environments that make this kind of unchecked deployment easier to do at scale. Part 4 comes back to what actually prevents these specific failures.

About Senserva: Who We Are and What We Do

Senserva is a Microsoft-focused security posture management company based in St. Paul, Minnesota. Its platform continuously scans an organization's Microsoft 365, Intune, Defender, and Entra ID environment, running several hundred built-in checks across identity, privileged access, applications and service principals, endpoint and device compliance, email, logging, and Azure subscription roles.

For every finding, Senserva provides step-by-step remediation guidance, including validated PowerShell scripts and rollback notes, and maps results to recognized compliance frameworks such as the Microsoft Cloud Security Benchmark, CISA's SCuBA baselines, NIST 800-171, CIS, ISO 27001, SOC 2, HIPAA, and PCI-DSS. Senserva is a member of the Microsoft Intelligent Security Association (MISA) and a Microsoft Security Excellence Awards finalist.

Senserva also offers Three Free Unlimited Audits covering every tenant, user, and patch: one to Find, one to Fix, and one to Prove, using the same 650+ checks described above. Registration takes only a work email, with no cost and no obligation to buy. Request an evaluation key at senserva.com/request-key.html.

Senserva also runs Senserva Watch, a free membership that tracks CVEs and Microsoft patches on an organization's behalf and sends an alert only when something being tracked changes, including a same-day notice when Microsoft's Patch Tuesday updates ship. Signing up needs no tenant connection, agent, or call, just an email address, at senserva.com/senserva-watch.html.

Already a Senserva Watch member, or ready to join? Sign up at senserva.com/when-the-lab-meets-the-cloud.html to get the complete five-part series sent to your inbox, plus everything else membership includes: alerts when a CVE or KB you follow changes, a fresh audit credit every quarter, and 20% off when you buy.

Learn more at senserva.com.

Sources cited in this part

  • Hou, Xiaoyuan, et al. "Using artificial intelligence to document the hidden RNA virosphere." Cell 187 (2024): 6929-6942.e16.
  • Urbina, Fabio, Filippa Lentzos, Cedric Invernizzi, and Sean Ekins. "Dual use of artificial-intelligence-powered drug discovery." Nature Machine Intelligence 4 (2022): 189-191.
  • Trump, Benjamin D., et al. "Governing the AI-biotech convergence." EMBO Reports 27 (2026): 259-264.
  • Prunkl, Carina. "AI meets biology: a call for community governance." Nature Methods 21 (2024): 1407-1408.
  • Foreign Affairs Forum. "Biosecurity and Oversight Gaps in Dual-Use Biotechnology, Part V." May 6, 2026.
  • Li, Z. et al. "From Code to Life: The AI-Driven Revolution in Genome Editing." Advanced Science (2025).
  • Faezi, Sina, Rozhin Yasaei, Anomadarshi Barua, and Mohammad Abdullah Al Faruque. "Oligo-Snoop: a non-invasive side channel attack against DNA synthesis machines." Proceedings of the 26th Annual Network and Distributed System Security Symposium (NDSS), 2019.
  • TrapX Research Labs. MEDJACK.4 Medical Device Hijacking. San Jose, CA: TrapX Security, 2018.
  • Peccoud, Jean, Jenna E. Gallegos, Randall Murch, W. Gregory Buchholz, and Sanjay Raman. "Cyberbiosecurity: from naive trust to risk awareness." Trends in Biotechnology 36, no. 1 (2018): 4-7.
  • Sweeney, Latanya. "Simple Demographics Often Identify People Uniquely." Carnegie Mellon University, Data Privacy Working Paper 3. Pittsburgh, 2000.
  • Gursoy et al. ML4H 2023 Symposium. arXiv:2403.01628, 2024.
  • U.S. Department of Health and Human Services, Office for Civil Rights. Annual Report to Congress on Breaches of Unsecured Protected Health Information for Calendar Year 2024. 2025.
  • U.S. Department of Health and Human Services. "Change Healthcare Cybersecurity Incident Frequently Asked Questions." HHS.gov, updated January 24, 2025.
  • American Medical Association. "2 in 3 Physicians Are Using Health AI: Up 78% from 2023." AMA, February 26, 2025.
  • Healthcare Information and Management Systems Society (HIMSS). 2024 Healthcare Cybersecurity Survey. March 2025.
  • Fisher Phillips. "New Class Action Targets Healthcare AI Recordings." Fisher Phillips Insights, December 9, 2025.
  • Obermeyer, Ziad, Brian Powers, Christine Vogeli, and Sendhil Mullainathan. "Dissecting racial bias in an algorithm used to manage the health of populations." Science 366 (2019): 447-453.
  • Saha, Shalini, et al. "Navigating AI and machine learning in cancer research: an end-to-end translational framework." Journal of Translational Medicine 24 (2026): 827.
  • National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1, January 2023.
  • International Organization for Standardization. ISO/IEC 42001:2023, Information Technology, Artificial Intelligence, Management System. 2023.

All posts