University of Toronto researchers presented GPUBreach at Black Hat USA 2026 in Las Vegas this week, demonstrating an attack that starts in an unprivileged CUDA process and ends with a root shell on the host CPU. The attack works with IOMMU protection enabled. For operators running untrusted workloads on shared GPUs, that makes IOMMU alone an insufficient defense against this attack chain.

The important detail is how it crosses that boundary. GPUBreach first uses Rowhammer to corrupt GPU memory. It then writes into buffers the GPU is allowed to access and exploits flaws in the NVIDIA kernel driver that processes them. The IOMMU enforces its memory-access rules, but those rules don't prevent the driver from mishandling corrupted data.

As of August 6, 2026, the practical response has several parts: patch the driver vulnerabilities, enable supported memory-error protection, and review which workloads share physical GPU memory. Each addresses a different part of the attack.

From memory bit-flips to a host root shell

Rowhammer exploits the physical proximity of memory cells. Repeatedly accessing one memory row can disturb nearby rows and flip bits in memory the attacking process never directly accessed. The original 2014 research demonstrated this on DRAM, and later work extended the technique to memory such as LPDDR4 in mobile devices.

GPUBreach targets GDDR6, which is widely used in NVIDIA consumer and workstation GPUs. In controlled testing, the University of Toronto team demonstrated up to 1,171 confirmed bit-flips on a single NVIDIA RTX 3060. That result establishes a way to cause memory corruption on the tested hardware, rather than merely a theoretical possibility.

The researchers connected that corruption to host privilege escalation in three stages:

  1. Position GPU page tables. The attacker uses CUDA's Unified Virtual Memory (UVM) allocation primitives to place GPU page-table entries, or PTEs, next to a memory row susceptible to hammering. Page tables map virtual addresses to physical memory. By reverse-engineering the NVIDIA driver, the team identified contiguous 2 MB regions used for GPU page tables and developed timing side channels to detect when new regions were allocated.
  2. Corrupt the mappings. Repeated row hammering flips bits in PTEs. Corrupted frame numbers then point to additional page-table pages, allowing the attacker to gain arbitrary read and write access to GPU memory.
  3. Exploit the host driver. The compromised GPU uses direct memory access, or DMA, to write into driver-owned buffers permitted by the IOMMU. Corrupting metadata inside those buffers triggers out-of-bounds writes in the NVIDIA kernel driver, completing the path to host root.

The Cloud Security Alliance research note describes a complete attack path from GPU memory bit-flips to a CPU root shell. The demonstrated consequence goes beyond one GPU workload accessing another workload's memory: the attacker gains control of the host.

What IOMMU can and cannot stop

An Input-Output Memory Management Unit, or IOMMU, restricts which host memory regions a device can access through DMA. Earlier GPU Rowhammer disclosures, GDDRHammer and GeForge, demonstrated GPU memory bit-flip attacks whose host escalation paths could be blocked by enabling IOMMU. NVIDIA's hardening guidance pointed to IOMMU as a primary control, and shared GPU infrastructure relied on it as an important isolation boundary.

GPUBreach doesn't disable that protection. The GPU's initial DMA write lands inside an authorized buffer. The failure comes later, when trusted driver code processes attacker-corrupted metadata and writes outside the intended bounds. The IOMMU doesn't validate the meaning of buffer contents or the driver's subsequent operations.

CVE-2025-33220 and CVE-2025-33218 cover the companion kernel driver vulnerabilities used to complete the escalation. This distinction matters when choosing a fix. Driver patches can break the demonstrated route to host root without removing the underlying hardware's susceptibility to Rowhammer.

Hardware exposure and shared inference workloads

The reported testing covers 25 models, with GDDR6-equipped NVIDIA GPUs from the RTX 20-series onward described as confirmed or likely vulnerable. The RTX A6000 is explicitly confirmed. The reported exposure also includes consumer hardware ranging from the RTX 2080 through the RTX 4090. Confirmed results and likely exposure shouldn't be treated as interchangeable when assessing a particular card.

The operational concern is especially relevant to the middle tier of AI infrastructure: workstation cards such as the A6000, lower-cost cloud inventory such as that offered through Lambda Labs and Vast.ai, and inference clusters built with 3090s and 4090s during GPU shortages. With H100s costing around $40,000 a unit, those choices offered a cheaper way to add capacity. They also put consumer and workstation hardware into environments where unrelated customers may share physical resources.

That doesn't establish that any named provider has an exploitable deployment. The risk depends on the card, driver, memory protection and workload placement. Shared physical GPU memory between attacker and target is among the research team's stated preconditions, making co-tenancy a central part of the assessment.

Hopper and Blackwell datacenter GPUs, including the H100, H200, B100 and B200, use HBM3 or HBM3e memory and have system-level error-correcting code protection, or ECC, enabled by default. The GDDR6 attack channel doesn't apply to HBM architecture in the same way. GDDR7 also reportedly shows substantially greater resistance to this class of attack. Neither distinction should be read as a claim that every other memory attack is impossible.

Model corruption can matter without host root

The host escalation is only one consequence. The CSA research note reports that a single targeted bit-flip in model weights reduced inference accuracy from approximately 80% to 0.1%. That is a result from the demonstrated setting, not an expected outcome for every model or every flipped bit.

Corrupted weights can leave outputs structurally valid. A model may still return JSON, respond coherently to other inputs and pass basic availability checks while failing badly on affected tasks. Monitoring that checks only whether inference requests succeed won't necessarily detect this kind of damage. The research also reports cryptographic key leakage through the same mechanism.

For an agentic AI pipeline, the concern is that a compromised model could produce systematically wrong decisions while the surrounding service appears healthy. The reported model-corruption demonstration supports that concern, although its consequences in a particular production agent remain workload-dependent.

Runtime model-integrity verification, including cryptographic hashing of weight tensors, deserves consideration for sensitive inference workloads on shared infrastructure. Availability monitoring and weight-integrity checks answer different questions, and patching the host escalation path doesn't by itself establish that model memory is protected.

Actions for GPU infrastructure operators

Apply the NVIDIA driver patches for CVE-2025-33220 and CVE-2025-33218. These address the companion vulnerabilities that connect GPU memory compromise to host root. They don't fix the physical Rowhammer channel, but breaking that escalation path substantially reduces the consequences of the demonstrated attack.

Enable system-level ECC on GDDR6 cards that support it. ECC corrects single-bit errors and detects double-bit errors, which can interrupt the corruption phase before it damages page tables. The reported cost is approximately 6% of VRAM and a firmware configuration change. That capacity tradeoff is reasonable for valuable inference workloads and shared systems, though support must be checked for each card.

Review physical GPU sharing. Untrusted CUDA workloads sharing a card with production inference jobs require an isolation review regardless of IOMMU configuration. The assessment needs to follow actual physical memory sharing, rather than stop at separate containers, jobs or customer accounts.

Use memory architecture as a procurement criterion. For sensitive or multi-tenant workloads purchased over the next 12 months, HBM3-based hardware such as the H100 and newer datacenter alternatives is a reasonable baseline recommendation. It avoids relying on GDDR6 defenses alone, although it doesn't eliminate the need for driver maintenance and isolation controls.

The CSA notes that three independent research teams have converged on GPU Rowhammer as a viable attack channel within the past 18 months. Its assessment is that this convergence strengthens the evidence for the risk and may shorten the time between disclosure and exploitation. That supports prioritizing mitigation without claiming that widespread exploitation has already occurred.

GPUBreach is being presented at Black Hat USA 2026 this week. The researchers disclosed it to NVIDIA on November 11, 2025, and the work was accepted to the 47th IEEE Symposium on Security and Privacy. The research and supporting material are available at gpubreach.ca.