Hitesh Sahu
Hitesh SahuHitesh Sahu
  1. Home
  2. ›
  3. work
  4. ›
  5. …

  6. ›
  7. 5 gpu lens

Loading ⏳
Fetching content, this won’t take long…


💡 Did you know?

🍌 Bananas are berries, but strawberries are not.

🍪 This website uses cookies

No personal data is stored on our servers however third party tools Google Analytics cookies to measure traffic and improve your website experience. Learn more

Cloud-DevOps

    AI & Machine Learning

    Cloud & DevOps
    • Betting Platform


    • Connected Vehicles


    • Payment Communication Platform


    • GhostFleet — Kubernetes GPU Cluster Simulator


    • GPU-Lens — Drop-in GPU Observability for Kubernetes & SLURM


    • Squint — GPU-Aware Terminal UI for SLURM Clusters


    Full-Stack Applications

    Mobile Development

Cover Image for GPU-Lens — Drop-in GPU Observability for Kubernetes & SLURM
Fork me on GitHub
Cloud & DevOps

GPU-Lens — Drop-in GPU Observability for Kubernetes & SLURM

Open Source

2026

Creator & Maintainer

Tech Stack
DCGM
Prometheus
Grafana
Alertmanager
Kubernetes
SLURM
Helm
Docker Compose
Shell Scripting
Terraform

Summary

Drop-in GPU and scheduler observability for clusters you already have — DCGM metrics, pre-wired Grafana dashboards, and alert rules for Kubernetes and SLURM without replacing existing infrastructure.


What I Built

Project Overview

GPU-Lens delivers production-grade observability for existing GPU clusters through two core layers:

  1. GPU Health Monitoring — utilization, memory, temperature, ECC errors, and XID errors via NVIDIA DCGM
  2. Scheduler State Visibility — queue depth and node status for both Kubernetes and SLURM environments

The focus is practical deployment over theoretical examples. GPU-Lens ships pre-configured alert rules and uses validated Grafana dashboards (fetched by ID from grafana.com) rather than hand-written configurations — so it works out of the box on any cluster that already runs Prometheus.


Key Features

DCGM Exporter — Kubernetes or Bare Metal

Deploys the NVIDIA DCGM Exporter as a Kubernetes DaemonSet (via Helm) or as a Docker Compose service for bare-metal and SLURM nodes — the same metrics pipeline either way.

Pre-Wired Grafana Dashboards

Ships with validated community dashboards imported by ID — no YAML hand-editing:

DashboardGrafana IDWhat it shows
DCGM GPU metrics12239Utilization, memory, temp, ECC, XID, NVLink
Node Exporter1860CPU, memory, disk, network per host

Alert Rules Out of the Box

Pre-configured PrometheusRule / Alertmanager alerts with tunable thresholds:

AlertCondition
GPUExporterDownDCGM exporter unreachable
GPUHighTemperatureDie temp > 80°C for 5m
GPUMemoryPressureVRAM > 90% for 10m
GPUECCErrorDouble-bit ECC error detected
GPUIdleReservationGPU reserved but < 10% util for 30m

SLURM Scheduler Metrics

Optional SLURM metrics collection exposes queue depth, pending jobs, and node states to Prometheus — closing the gap between GPU health and job scheduler visibility.


Architecture

GPU Nodes
    │
    ├── DCGM Exporter (DaemonSet / Docker Compose)
    │       └── /metrics on :9400
    │
    └── Node Exporter
            └── /metrics on :9100

Prometheus ──scrapes──▶ Alertmanager ──routes──▶ PagerDuty / Slack
     │
     └──▶ Grafana
              ├── Dashboard 12239 (GPU / DCGM)
              └── Dashboard 1860  (Node Exporter)

Why It Matters

Without GPU-LensWith GPU-Lens
Setting up GPU observability from scratch takes daysSingle helm install or docker compose up
Grafana dashboards require manual JSON editingDashboards fetched by ID — always up to date
Alert thresholds are guessworkPre-tuned rules based on NVIDIA recommended limits
Kubernetes and SLURM need separate toolingSame Prometheus/Grafana stack covers both

Technology Stack

GPU Metrics: NVIDIA DCGM, dcgm-exporter

Metrics Collection: Prometheus, Node Exporter

Visualisation: Grafana (dashboards 12239, 1860)

Alerting: Alertmanager

Kubernetes Deployment: Helm 3

Bare-Metal / SLURM Deployment: Docker Compose

Infrastructure: Terraform (optional)

Language: Shell, YAML

License: Apache 2.0

← Previous

GhostFleet — Kubernetes GPU Cluster Simulator

Next →

Squint — GPU-Aware Terminal UI for SLURM Clusters

Let's work together
hiteshkrsahu@gmail.com
Munich 🥨, Germany 🇩🇪, EU
Playstore
Hitesh Sahu's apps on Google Play Store
Need Help?
Let's Connect
Navigation
  Home/About
  Skills
  Work/Projects
  Lab/Experiments
  Contribution
  Awards
  Art/Sketches
  Thoughts
  Contact
Links
  Sitemap
  Legal Notice
  Privacy Policy

Made with

NextJS logo

NextJS by

hitesh Sahu

| © 2026 All rights reserved.