Security Notes
Malware & Exploitation

Malware That Lives on the GPU

Most malware lives where the defender looks: on disk and in system RAM, executed by the CPU. GPU-resident malware is the idea of moving the malicious code and data off the CPU and onto the graphics card — into the GPU's own memory (VRAM), executed by the GPU's thousands of cores — precisely because that's where the antivirus and EDR are not looking. This document separates the real research from the hype, explains what is and isn't actually possible, and covers why the AI/cloud boom has made GPUs a security topic again.

12 min read 8 sections 4 model answers verified 2026-06

Last verified2026-06

Memory hook

the GPU is a blind spot, not a magic wand. Host security tools scan disk and system RAM; they are almost completely blind to VRAM and to code running on GPU cores. That blindness is the entire appeal. But a GPU can't bootstrap itself, can't make system calls, and loses its memory on power-off — so there is always a CPU-side loader to find, and that loader is the defender's foothold. The one-liner: GPU malware hides the body in VRAM, but the hands are still a normal CPU process.


What Is GPU-Resident Malware?

In plain terms: a graphics card is a second computer inside your computer. It has its own processors (thousands of small cores), its own memory (VRAM — video RAM, separate from the system RAM the CPU uses), and it can move data around on its own. Normally it draws pixels and crunches math for games and machine learning. GPU-resident malware repurposes that second computer to store and run malicious code there instead of on the main CPU.

First, an important distinction, because the term gets muddied:

TermWhat it really isLives where?
Cryptojacking / GPU minerOrdinary malware that uses the GPU to mine cryptocurrency for profitCode on disk/host; just offloads compute to the GPU
GPU-resident malware / GPU rootkitMalicious code/data stored in VRAM and executed on the GPU to evade host scanningThe payload lives in VRAM
GPU as a DMA weaponUsing the card's ability to read/write system memory directly to spy on the hostCode may be host-side; the capability is the GPU's

The overwhelming majority of "GPU malware" seen in the wild is the first row — cryptojacking — which isn't hiding in the GPU at all. The genuinely interesting, evasion-focused research is the second and third rows, and that's the focus here.


Why an Attacker Wants the GPU

Evasion — the big one.

Antivirus and EDR products scan files on disk and processes in system RAM. They have essentially no visibility into VRAM or into the instructions executing on GPU cores. Code sitting in video memory is, to almost every host security product, invisible. Memory-forensics tools like Volatility analyse a dump of system RAM — they don't see the card.

Direct memory access (DMA).

A discrete GPU is a bus-master device on the PCIe bus: it can read and write the host's system RAM directly, without asking the CPU, unless an IOMMU (the hardware that polices DMA) is configured to stop it. That gives a GPU implant a stealthy way to read host memory — scrape secrets, keystrokes, or decrypted data — and to communicate with its CPU-side loader.

Raw parallel horsepower.

Thousands of cores make GPUs ideal for the tasks attackers care about anyway: password/hash cracking, cryptomining, and bulk cryptographic operations (e.g., mass file encryption for ransomware). This is the cryptojacking motive — less about stealth, more about free compute.

It's where the value is now.

The AI boom means the most valuable data and models increasingly live on the GPU during processing — making the card itself a target, not just a hiding place.


How It Actually Works

The architecture is always a two-part split: a small CPU-side loader plus a GPU-side payload. The GPU cannot start itself — a normal process on the CPU must use the graphics driver (via a compute framework like CUDA on NVIDIA, or the vendor-neutral OpenCL) to allocate VRAM, copy the payload in, and launch it.

        HOST (CPU side)                         GPU (card)
  ┌───────────────────────────┐         ┌──────────────────────────┐
  │  Loader process            │         │  VRAM (video memory)     │
  │  - uses CUDA/OpenCL driver │  copy   │  ┌────────────────────┐  │
  │  - allocates VRAM   ───────┼────────►│  │ malicious payload  │  │
  │  - copies payload          │ launch  │  │ + stolen data      │  │
  │  - launches GPU "kernel" ──┼────────►│  └────────────────────┘  │
  │  - does I/O / network for  │◄────────┼── results / DMA reads    │
  │    the GPU                  │   DMA   │     of host RAM           │
  └───────────────────────────┘         └──────────────────────────┘
     ▲ visible to EDR                       ▲ invisible to EDR

Step by step:

Loader runs on the CPU

and opens a GPU compute context through the installed driver/runtime (CUDA or OpenCL).

Payload is copied into VRAM.

Once there, it's outside the view of host AV scanning system RAM. The malware can keep its real logic (and any stolen data) resident in video memory between bursts of activity.

A GPU "kernel" executes

the payload on the GPU cores. (Here "kernel" means a GPU compute function, not the OS kernel.)

The GPU reaches back to the host.

For anything the GPU can't do itself — networking, files, syscalls — it relies on the CPU loader. For stealthy reads of host memory it can use DMA (on discrete cards), pulling data straight out of system RAM.

Persistence is delegated to the host.

Because VRAM is wiped on power-off, the malware needs a normal host persistence mechanism (a service, cron job, run key) to reload the payload into the GPU after each reboot.


The Hard Limits (Why This Is Still Niche)

GPU malware sounds like a perfect ghost, but the constraints are severe — and they're exactly what a defender leans on:

There is always a CPU footprint.

The GPU can't bootstrap itself; a host process must allocate, copy, and launch. That loader is detectable by normal means — and it's where defenders should focus.

GPUs can't make system calls.

A GPU core can do math and touch memory, but it cannot open a socket, write a file, or invoke the OS. Anything "interesting" at the OS level still routes through the CPU — so a GPU-only implant is a brain with no hands.

VRAM is volatile.

Power off and the payload is gone. No host persistence ⇒ no survival across reboots.

Heavy dependencies.

It needs the vendor driver and a compute runtime (CUDA/OpenCL) present and usable by the malware's process — not guaranteed on every target.

Fragile and hardware-specific.

Behaviour varies across vendors (NVIDIA/AMD/Intel/Apple) and driver versions; PoCs are notoriously brittle.

The honest takeaway: GPU-resident malware is a real, demonstrated evasion technique, but it is rare, brittle, and PoC-grade. When you hear "GPU malware" in the news, it's almost always cryptojacking, not a VRAM-resident rootkit.


Real Research and PoCs

JellyFish, WIN_JELLY, and "Demon" (2015).

A team published open-source proof-of-concept GPU malware: JellyFish, a Linux GPU rootkit using OpenCL that kept code in VRAM and used GPU DMA to reach host RAM; WIN_JELLY, a Windows variant; and Demon, a GPU-based keylogger inspired by an earlier academic paper. These are the canonical references and showed the CPU-loader + VRAM-payload + DMA pattern that defines the category.

Academic keyloggers and rendering attacks.

Research such as "You Can Type, but You Can't Hide: A Stealthy GPU-based Keylogger" (2013) demonstrated monitoring the keyboard buffer from the GPU, and other work showed that leftover GPU memory could leak rendered web pages between processes — an early hint of the multi-tenant problem below.

2021 VRAM-execution PoC.

A proof of concept that allocated address space in the GPU buffer and executed code directly from VRAM via OpenCL was offered for sale on a hacking forum; it was reported (Aug 2021) to have been tested across Intel, AMD, and NVIDIA cards. It renewed attention but, again, was a demonstration of the technique, not a widespread campaign.

LeftoverLocals — CVE-2023-4969 (disclosed Jan 2024, Trail of Bits).

Not malware itself, but a vulnerability that crystallises the GPU's security problem: GPU local memory is not reliably cleared between uses, so one process could read another process's leftover data from the GPU. The researchers recovered the output of a large language model running on a co-resident process — directly relevant to AI workloads. Affected GPUs from Apple, AMD, Qualcomm, and Imagination.


The Modern Angle: Cloud, Kubernetes, and AI

For years GPU malware was an academic curiosity. The AI boom changed the calculus, because GPUs are now everywhere and shared:

GPU instances and GPU nodes.

Cloud GPU VMs (e.g., EC2 GPU instances) and Kubernetes GPU node pools for ML training/inference put high-value, internet-exposed GPUs into reach. A compromised ML training pod sits on the card already.

Multi-tenant GPU sharing → cross-tenant leakage.

To use expensive GPUs efficiently, platforms slice them — NVIDIA MIG (Multi-Instance GPU), time-slicing, and vGPU. If VRAM isn't zeroed between tenants/allocations (the LeftoverLocals class of bug), one workload can read another's residual data — model weights, inference inputs/outputs, secrets. This is the most practically worrying GPU-security issue today, and it needs no rootkit at all.

AI data lives on the GPU.

Model weights, prompts, and inference results are processed in VRAM. That makes the card a target for theft of model IP and of sensitive prompt/response data, and makes residual-memory leakage a confidentiality problem, not just a curiosity.

Supply chain.

Malicious CUDA kernels or compromised ML libraries/containers can run attacker code on GPU nodes as a matter of course — the GPU becomes an execution surface reached through ordinary software supply-chain compromise.

For a defender in a cloud/Kubernetes/AI shop, the realistic GPU threats, in order, are: (1) cryptojacking on GPU nodes, (2) cross-tenant VRAM data leakage in shared-GPU setups, and only then (3) true GPU-resident implants.


Detection and Defence

The strategy follows straight from the limits: you don't need to see into the GPU if you watch the door to it.

Hunt the loader, not the VRAM.

Every GPU implant needs a CPU-side process that opens a CUDA/OpenCL context. A non-graphics, non-ML process loading the GPU compute runtime is anomalous — baseline which processes legitimately use the GPU and alert on the rest.

Watch GPU telemetry.

Unexpected GPU utilisation, memory allocation, or temperature on a host that has no business doing GPU compute is the classic cryptojacking tell (nvidia-smi, vendor counters, cloud GPU metrics). Sudden 100% GPU on a web server = investigate.

Constrain DMA with the IOMMU.

Enabling the IOMMU (Intel VT-d / AMD-Vi) limits a device's ability to DMA arbitrary host memory, blunting the GPU-as-DMA-weapon angle. Cloud hypervisors already use this for isolation.

Zero GPU memory between tenants.

For shared/multi-tenant GPUs, ensure the platform clears VRAM between allocations and patch the LeftoverLocals class of bug (CVE-2023-4969); prefer hard isolation (MIG) over soft time-slicing for sensitive workloads.

Treat GPU nodes like any other endpoint.

Restrict who can schedule GPU workloads, scan container images, and apply the same supply-chain controls to CUDA/ML dependencies as to any other code.

Accept the forensic gap and compensate.

EDR and Volatility are blind to VRAM, and VRAM capture is volatile and hard — so detection has to lean on the host loader, GPU metrics, and network/DMA behaviour rather than on imaging the card.

Memory hook

guard the on-rampYou can't easily see code running on the GPU, so don't try to win there. Win at the on-ramp: the CPU loader that has to open a CUDA/OpenCL context, the anomalous GPU utilisation, and the DMA the IOMMU can fence off. And for the modern, more likely problem — shared GPUs in AI/cloud — the fight is zeroing VRAM between tenants, not hunting rootkits.


Interview Q&A

Q
What is GPU-resident malware, and why would an attacker use the GPU?
Model answer

It's malware that stores its code and data in the graphics card's own memory — VRAM — and executes it on the GPU's cores, instead of running entirely on the CPU. The whole point is evasion: antivirus and EDR scan disk and system RAM but are essentially blind to VRAM and to GPU execution, so the payload hides where defenders don't look. Discrete GPUs are also bus masters that can DMA into host memory, giving a stealthy way to read the host, and their parallelism is great for cracking or mining. The catch is that a GPU can't bootstrap itself or make system calls, so there's always a CPU-side loader — which is exactly where I'd hunt it.

Q
Why isn't GPU malware everywhere if it's so stealthy?
Model answer

Because the constraints are brutal. The GPU can't start itself — a normal CPU process has to allocate VRAM, copy the payload in, and launch it through the CUDA or OpenCL driver, so there's always a detectable host footprint. The GPU can't make system calls, so anything OS-level still routes through the CPU — it's a brain with no hands. VRAM is volatile, so without host persistence the payload dies on reboot. And it depends on specific drivers and runtimes and is fragile across hardware. So it's a real, demonstrated technique but stays PoC-grade. What people usually call "GPU malware" in the wild is just cryptojacking, which uses the GPU for compute but doesn't hide inside it.

Q
How would you detect or defend against it?
Model answer

I'd guard the on-ramp rather than try to see inside the card. Every GPU implant needs a CPU process to open a CUDA or OpenCL context, so I baseline which processes legitimately use the GPU and alert on anything else — a web server suddenly loading the GPU runtime is a red flag. I'd watch GPU utilisation and memory metrics for cryptojacking, enable the IOMMU to fence off malicious DMA, and treat GPU nodes like any other endpoint for supply-chain and image scanning. I'd also accept that EDR and Volatility can't see VRAM, so detection leans on the loader, GPU telemetry, and network behaviour, not on imaging the card.

Q
What's the most realistic GPU security threat in a cloud/AI environment today?
Model answer

Honestly, not rootkits — it's cryptojacking on GPU nodes first, and cross-tenant data leakage second. To use expensive GPUs efficiently, platforms share them with MIG, time-slicing, or vGPU, and if VRAM isn't zeroed between tenants one workload can read another's leftover data. That's the LeftoverLocals class of bug, CVE-2023-4969 from Trail of Bits in 2024, where researchers recovered another process's large-language-model output straight from GPU memory. In an AI shop that's a serious confidentiality problem because model weights, prompts, and responses all live in VRAM during processing. The fix is ensuring the platform clears GPU memory between allocations and preferring hard isolation like MIG for sensitive workloads — no implant required for the attack, so no implant-hunting fixes it.