Blog

CVE-2026-47483: How We Could Disrupt Thousands of Exposed GPU Servers

We found a high-severity vulnerability in NVIDIA’s GPU exporter that could let attackers crash monitoring and potentially disrupt AI workloads without authentication. We also uncovered 2,100 servers leaking telemetry from 12,000+ GPUs to anyone online.

NVIDIA DCGM Exporter is a widely used tool for monitoring GPUs across AI servers and clusters.

We discovered a high-severity vulnerability in NVIDIA DCGM Exporter that allowed an unauthenticated attacker to trigger uncontrolled resource consumption on the GPU server, leading to denial of service and information disclosure. By exhausting the exporter’s resources, an attacker could crash the monitoring service, blind operators to GPU health and activity, and potentially slow workloads running on the same host.

What made the issue especially significant was the scale of the exposure. During our research, we found more than 2,000 GPU servers exposing NVIDIA DCGM Exporter directly to the public internet without authentication. Across four scans, those systems reported more than 12,000 unique GPUs, representing an estimated $100 million in hardware.

The exposed systems included NVIDIA Blackwell Ultra B300 GPUs, H200s and H100s used for large-scale AI workloads, as well as consumer RTX 5090 and 4090 systems. Anyone who could reach these endpoints could see what hardware organizations were running, how heavily it was being used, and details about the AI infrastructure around it.

We reported the issue to NVIDIA, which assigned it CVE-2026-47483 (CVSS 8.2, High) and published a security bulletin addressing the vulnerability.

What started with a vulnerability in a GPU monitoring service gave us a rare view into how GPU infrastructure is deployed, exposed, and used across the internet.

Background

Earlier this year, we enabled NVIDIA’s DCGM Exporter on one of our own GPU clusters.

A DCGM Exporter reads telemetry directly from the GPUs on a host, including: temperature, utilization, memory usage, power consumption, error events, and more. It then exposes those metrics over HTTP, typically at :9400/metrics, so systems like Prometheus can scrape them.

A single response can contain hundreds of lines of metrics, with each GPU identified by its own unique ID, or UUID.

Exposed DCGM Exporter metrics reveal GPU activity, hardware details, and driver versions.

From these metrics alone, an external observer can learn a lot about the underlying GPU system:

  • What hardware is running? modelName reveals the exact GPU model, such as an NVIDIA H100 80GB.

  • How is it being used? GPU_UTIL, FB_USED, tensor activity, and DRAM_ACTIVE reveal idle, compute-heavy, or memory-heavy behavior. These patterns can hint at training or inference, but are not conclusive.

  • When is it busy? Repeated readings can reveal sustained activity, short bursts, and quiet periods, helping outsiders infer workload schedules and usage patterns.

  • How is the system performing? Power consumption, NVLink traffic, temperature, and XID errors reveal patterns about workload intensity, communication between GPUs, and hardware or driver issues.

Collectively, an exposed DCGM Exporter provides a surprisingly detailed profile of a GPU server: what hardware it contains, how heavily it is being used, how its GPUs communicate, and whether it is experiencing errors.

Once we realized how much these endpoints revealed, the next question was:

How many of them are exposed to the internet?

Scanning the internet for NVIDIA DCGM exporters

We started with a simple Shodan query:

Across four scans between March and May 2026, we observed roughly 2,100 hosts publicly serving DCGM Exporter metrics and exposing more than 12,000 unique GPU UUIDs. Every host returned metrics over plaintext HTTP, and none required authentication.

That exposure gave outsiders information useful for reconnaissance:

  1. Profiling the target. Hardware and software details help attackers understand the environment behind an endpoint and narrow their search for relevant known vulnerabilities.

  2. Understanding workload activity. Repeated observations can reveal busy periods, recurring usage patterns, and changes in how compute is used. Memory allocations can also remain visible during quiet compute intervals, providing clues about workloads that remain resident.

  3. Following infrastructure over time. GPU UUIDs make individual devices recognizable across observations, even when their endpoint or hostname changes. An observer can connect separate snapshots and follow changes in where the same hardware appears and how it is used.

  4. Identifying workloads and projects. Where Kubernetes workload labels are present, pod, namespace, and container names can reveal application or project names and associate them with specific GPUs. Descriptive names may even hint at the models or services being run.

But while examining the same service, we found that the same exporter could be turned against itself, remotely driven into resource exhaustion.

Vulnerability in NVIDIA DCGM Exporter

About a quarter of the exposed DCGM hosts in our scan were serving Go’s /debug/pprof/ profiling endpoints alongside /metrics.

pprof is Go’s built-in profiling interface, typically used internally for debugging CPU and memory usage. Some profiling endpoints keep requests active for a caller-specified duration, and many concurrent unauthenticated requests can drive increasing memory consumption.

At first, we assumed the exposed profiling endpoints were the result of operator misconfiguration. We then reproduced the same behavior using NVIDIA’s official DCGM Exporter container without modifying it. Deployments exposing the exporter on a routable interface could expose /debug/pprof/ as well.

With enough concurrent unauthenticated requests, the exporter could run out of memory and crash, cutting off visibility into GPU health and activity. The CPU and RAM pressure could also slow the host and interfere with training or inference workloads running beside it, particularly without strict resource limits. Hundreds of the public DCGM Exporters we observed exposed this attack surface. We tested the resource exhaustion in a controlled environment, not against those public deployments.

What Thousands of Exposed GPUs Reveal About the AI Boom

The exposed exporters handed us a rough inventory of the GPUs powering the AI boom, straight from the machines themselves. Half were data center accelerators, with models like NVIDIA H100s and H200s. The other half were consumer and workstation cards, mostly GeForce RTX 4090s and 5090s, which means a large share of the "AI infrastructure" sitting on the public internet is really gaming hardware pressed into service.

To our surprise, we also found current-generation Blackwell hardware wide open, including 312 NVIDIA B200s and 32 B300s, with their telemetry accessible without authentication.

The exposure crossed many kinds of organizations. Across our four scans, we found exposed GPU exporters on just over 2,000 hosts associated with nearly 300 organizations. Grouped by operator category, the exposed GPUs broke down like this:

  • Consumer GPU and mining farms: 35%. The largest category, with fleets dominated by GeForce cards.

  • Neoclouds: 25%. Providers offering GPU capacity as a rental service.

  • General hosting and colocation: 19%. Traditional hosting and colocation networks serving customer GPU deployments.

  • Telecoms and national ISPs: 10%. Carrier and ISP networks.

  • Hyperscalers: 6%. The large public clouds.

  • Universities and research institutes: 5%. Academic and HPC environments.

The geographic distribution was concentrated in a handful of countries. Endpoints in the United States accounted for 5,274 GPUs (44%), followed by Romania with 2,054 (17%) and China with 1,967 (16%). Smaller fleets appeared across Europe and Asia, including South Korea, Singapore, Iceland, France, Finland, and the Netherlands.

Here's what we didn't expect. Most of these GPUs were doing nothing. In every one of our four scans, roughly 60% of the GPUs reported 0% utilization. With demand for GPUs outstripping supply and prices climbing, we expected this hardware to be running flat out.

Neocloud Customers and Exposed Monitoring

We found public monitoring endpoints on infrastructure associated with Voltage Park, Northern Data, Lambda, and DigitalOcean. The largest group we documented was associated with Voltage Park: 672 public node_exporter hosts and 71 public DCGM Exporter hosts.

Voltage Park confirmed that most of the node_exporter instances and all the DCGM Exporter instances were customer-deployed. Although securing those services is the customers’ responsibility, Voltage Park’s security team investigated our report and reached out to affected customers. Its response shows how providers can help customers address exposures within the shared responsibility model.

Beyond the GPU: 12,096 Hosts Revealed the Infrastructure Behind AI

Alongside GPU telemetry, we also examined Prometheus Node Exporter, a tool for monitoring a server’s hardware and operating system. Like DCGM Exporter, it exposes metrics over HTTP for Prometheus to scrape, including CPU usage, memory usage, disk activity, and network statistics.

We found 12,096 public Node Exporter hosts reporting mlx5 metrics from NVIDIA/Mellanox InfiniBand and RoCE adapters, exposing adapter models, firmware versions, link state and fabric activity. Many also revealed hostnames, OS and kernel versions, and BIOS details. One endpoint returned:

That's enough to identify a Dell PowerEdge XE9680 running Ubuntu 22.04.5 LTS, its BIOS and adapter firmware, and an active 400 Gb/s port. Exact versions let an attacker go straight to matching known vulnerabilities. Where Node Exporter and DCGM Exporter overlapped, we could tie GPU activity to the server, software stack and network around it, all without authentication.

Exposed Node Exporter metrics reveal server hardware, OS and firmware versions, and InfiniBand link speeds.

What Should Operators and Customers Do?

Node Exporter, DCGM Exporter, and Prometheus should not be directly reachable from the public internet unless there is a specific need and appropriate access control.

Where possible, bind exporters to a loopback or private interface and restrict access using firewall rules, security groups, or other network controls so that only the monitoring infrastructure can reach them. Prometheus query APIs and target pages should be protected in the same way.

For DCGM Exporter, upgrade to 4.8.2 or later and make sure --enable-pprof is not enabled unless profiling is explicitly required. In current versions, the profiling endpoint is opt-in rather than exposed by default.

Conclusion

The findings highlight a growing security gap in AI infrastructure: companies are spending millions on GPUs while leaving critical systems exposed. Those exposures can reveal how AI environments are built and, in some cases, allow attackers to disrupt them.

Our research showed that publicly exposed monitoring services can reveal far more than basic telemetry. Across thousands of systems, we could see what GPUs were running, how they were being used, and details about the servers and networks around them.

In the case of NVIDIA DCGM Exporter, that exposure went beyond visibility and allowed attacker to exhaust resources, crash GPU monitoring, and potentially affect workloads running on the same host.

As AI infrastructure scales, organizations need to protect every layer of the stack and know exactly what they are responsible for versus what their provider is responsible for. 

Prior Work and Further Reading

This research builds on earlier work on exposed Prometheus infrastructure and Go profiling endpoints:

Continue reading

See what’s running in the data center