Hitesh Sahu
Hitesh SahuHitesh Sahu
  1. Home
  2. β€Ί
  3. posts
  4. β€Ί
  5. …

  6. β€Ί
  7. 0 INDEX

Loading ⏳
Fetching content, this won’t take long…


πŸ’‘ Did you know?

🀯 Your stomach gets a new lining every 3–4 days.

πŸͺ This website uses cookies

No personal data is stored on our servers however third party tools Google Analytics cookies to measure traffic and improve your website experience. Learn more

Loading ⏳
Fetching content, this won’t take long…


πŸ’‘ Did you know?

🦈 Sharks existed before trees 🌳.
AI-Infrastructure

    AI-AgenticAI

    AI-DeepLearning

    AI-GenAI

    AI-Infrastructure
    • NVIDIA AI Infrastructure and Operations Fundamentals


    • AI Infra Computing : GPU, DPU, Virtualization, DGX Systems


    • AI Programming Model


    • Pinned Memory (Page-Locked Memory) in CUDA and GPU Computing


    • RAPIDS and GPU Accelerated Data Science: cuDF, cuML, CUDA, NCCL and Distributed AI Pipelines


    • NVIDIA DCGM: GPU Health, Diagnostics, and Prometheus Metrics


    • NVIDIA Base Command Manager: Provisioning and Operating GPU Clusters


    • Slurm: The HPC Workload Manager Behind AI Training Clusters


    • TensorRT and High-Performance AI Inference: CUDA, ONNX, TensorRT-LLM and GPU Optimization


    • NCCL and Distributed GPU Communication: CUDA, AllReduce, Multi-GPU and AI Cluster Networking


    • ONNX (Open Neural Network Exchange): Portable AI Models, TensorRT and Cross-Framework Inference


    • LangChain and AI Agent Orchestration: RAG, LLM Workflows, Vector Databases and Tool Calling


    • NVIDIA NeMo and Enterprise AI Platforms: Distributed LLM Training, RAG and TensorRT-LLM


    • Megatron-LM and Distributed LLM Training: Tensor Parallelism, NCCL and Trillion-Scale AI Models


    • NVIDIA Triton Inference Server: TensorRT-LLM, GPU Serving and Production AI Inference


    • NVIDIA Riva: Real-Time Conversational AI with ASR, NLP and Text-to-Speech


    • NVIDIA NGC Catalog: GPU Optimized Containers, AI Models and Enterprise AI Infrastructure


    • AI Infra Networking: GPU Clusters, InfiniBand, RoCE, and DPU Integration


    • AI Infra Storage: NVMe, Parallel File Systems, Object Storage, and GPUDirect Storage


    • AI/ML Operations


    • AI-Infrastructure Index


    AI-Machine-Learning

    AI-Math

    AWS

    Azure

    kubernetes

    Management

    Programming

    Terraform

    Z_Appendix

Cover Image for AI-Infrastructure Index
AI-Infrastructure

AI-Infrastructure Index

πŸ“™ Index of AI-Infrastructure posts

πŸ“™ AI-Infrastructure Index

πŸ“š 21 Posts
πŸ•’ Last Updated: Sun Aug 23 2026

This folder contains AI-Infrastructure-related posts.

#Blog LinkDateExcerptTags
1AI-Infrastructure IndexSun Aug 23 2026πŸ“™ Index of AI-Infrastructure posts
2NVIDIA AI Infrastructure and Operations FundamentalsFri Feb 27 2026Comprehensive guide to NVIDIA AI infrastructure covering GPU architecture, accelerated computing, training vs inference workloads, data center networking, storage design, virtualization, and operational best practices.NVIDIA AI Infrastructure GPU Computing CUDA Data Center AI Training AI Inference Networking Storage Virtualization MLOps Certification
3AI Infra Computing : GPU, DPU, Virtualization, DGX SystemsFri Feb 27 2026Comprehensive overview of modern AI infrastructure covering CPU, GPU, and DPU architectures, accelerated computing models, cluster scaling, high-speed networking (InfiniBand and RoCE), storage integration, and power and cooling considerations for AI data centers.NVIDIA CPU Architecture GPU Architecture DPU BlueField Accelerated Computing AI Infrastructure AI Training AI Inference GPU Clusters Data Center InfiniBand RoCE AI Networking Power and Cooling Storage Architecture
4AI Programming ModelFri Feb 27 2026Overview of NVIDIA's AI programming model, including core libraries (CUDA, NCCL, cuDNN), training vs inference workloads, and compute scaling models (data parallelism and model parallelism) for AI infrastructure.NVIDIA AI Infrastructure GPU Clusters Data Center AI Training AI Networking InfiniBand RoCE DPU BlueField Power and Cooling On-Prem vs Cloud Accelerated Computing
5Pinned Memory (Page-Locked Memory) in CUDA and GPU ComputingTue May 26 2026Learn how pinned memory (page-locked memory) improves CPU-to-GPU data transfer performance in CUDA, deep learning, and high-performance AI workloads using direct memory access (DMA).AI CUDA GPU Computing NVIDIA Deep Learning AI Infrastructure High Performance Computing CUDA Memory Pinned Memory Page-Locked Memory DMA AI Training Machine Learning PyTorch TensorFlow
6RAPIDS and GPU Accelerated Data Science: cuDF, cuML, CUDA, NCCL and Distributed AI PipelinesTue May 19 2026Comprehensive overview of the RAPIDS ecosystem covering GPU accelerated DataFrames, machine learning, graph analytics, CUDA execution, distributed computing with Dask and NCCL, TensorRT integration, and large-scale AI data processing pipelines on NVIDIA GPUs.NVIDIA RAPIDS CUDA cuDF cuML cuGraph CuPy GPU Computing Accelerated Computing Data Science Machine Learning Distributed Computing Dask NCCL TensorRT AI Infrastructure GPU Clusters Data Engineering Vectorized Computing AI Pipelines
7TensorRT and High-Performance AI Inference: CUDA, ONNX, TensorRT-LLM and GPU OptimizationTue May 19 2026Comprehensive overview of NVIDIA TensorRT covering ONNX model optimization, CUDA kernel fusion, FP16 and INT8 inference, TensorRT-LLM, GPU memory optimization, Triton Inference Server integration, and production-scale AI inference pipelines on NVIDIA GPUs.NVIDIA TensorRT TensorRT-LLM CUDA ONNX GPU Inference AI Inference LLM Inference Deep Learning CUDA Kernels FP16 INT8 Quantization Triton Inference Server AI Infrastructure GPU Optimization Accelerated Computing AI Serving Production AI Inference Pipelines
8NCCL and Distributed GPU Communication: CUDA, AllReduce, Multi-GPU and AI Cluster NetworkingTue May 19 2026Comprehensive overview of NVIDIA NCCL covering GPU-to-GPU communication, AllReduce operations, distributed AI training, CUDA integration, tensor synchronization, multi-node scaling, InfiniBand networking, and high performance communication for large-scale AI and HPC workloads.NVIDIA NCCL CUDA Distributed Training GPU Communication Multi-GPU AllReduce Tensor Parallelism Pipeline Parallelism AI Infrastructure HPC InfiniBand RoCE GPU Clusters Deep Learning Megatron-LM NeMo TensorRT-LLM Accelerated Computing Parallel Computing
9ONNX (Open Neural Network Exchange): Portable AI Models, TensorRT and Cross-Framework InferenceTue May 19 2026Comprehensive overview of ONNX covering portable neural network model formats, cross-framework interoperability, ONNX Runtime, TensorRT integration, GPU accelerated inference, model optimization, and production AI deployment across heterogeneous hardware platforms.NVIDIA ONNX Open Neural Network Exchange ONNX Runtime TensorRT CUDA AI Inference Deep Learning Model Deployment GPU Inference PyTorch TensorFlow Machine Learning Cross Platform AI AI Infrastructure Accelerated Computing Portable Models LLM Inference Edge AI Production AI
10LangChain and AI Agent Orchestration: RAG, LLM Workflows, Vector Databases and Tool CallingTue May 19 2026Comprehensive overview of LangChain covering AI agents, Retrieval-Augmented Generation (RAG), prompt orchestration, tool calling, memory management, vector databases, multi-step LLM workflows, and production GenAI application development.LangChain Generative AI AI Agents LLM RAG Retrieval Augmented Generation Vector Databases Prompt Engineering AI Orchestration Tool Calling AI Workflows LangGraph OpenAI LLM Applications AI Infrastructure Semantic Search AI Copilot Workflow Automation Production AI Agentic AI
11NVIDIA NeMo and Enterprise AI Platforms: Distributed LLM Training, RAG and TensorRT-LLMTue May 19 2026Comprehensive overview of NVIDIA NeMo covering large language model training, distributed GPU scaling, Megatron-LM integration, Retrieval-Augmented Generation (RAG), NeMo Retriever, TensorRT-LLM optimization, and enterprise AI deployment pipelines for production-scale generative AI systems.NVIDIA NeMo CUDA NCCL Megatron-LM TensorRT-LLM Distributed Training LLM Generative AI AI Infrastructure RAG NeMo Retriever AI Agents GPU Clusters Accelerated Computing Enterprise AI Transformer Models Triton Inference Server Deep Learning Production AI
12Megatron-LM and Distributed LLM Training: Tensor Parallelism, NCCL and Trillion-Scale AI ModelsTue May 19 2026Comprehensive overview of NVIDIA Megatron-LM covering distributed transformer training, tensor and pipeline parallelism, NCCL communication, CUDA optimization, mixed precision training, trillion-parameter scaling, and large-scale GPU accelerated language model infrastructure.NVIDIA Megatron-LM CUDA NCCL Distributed Training Tensor Parallelism Pipeline Parallelism Context Parallelism Expert Parallelism LLM Training Transformer Models GPT AI Infrastructure Accelerated Computing Deep Learning Multi-GPU GPU Clusters TensorRT-LLM NeMo Trillion Parameter Models
13NVIDIA Triton Inference Server: TensorRT-LLM, GPU Serving and Production AI InferenceTue May 19 2026NVIDIA Triton Inference Server and vLLM compared β€” PagedAttention and continuous batching mechanics, TensorRT-LLM vs vLLM vs Triton tradeoff table, when to use each for production LLM serving, and Kubernetes deployment patterns for both.NVIDIA Triton Triton Inference Server TensorRT TensorRT-LLM CUDA AI Inference LLM Serving GPU Inference Dynamic Batching AI Infrastructure Kubernetes Multi-GPU Accelerated Computing Production AI AI APIs Deep Learning GPU Scheduling Inference Optimization Model Serving vLLM PagedAttention
14NVIDIA Riva: Real-Time Conversational AI with ASR, NLP and Text-to-SpeechTue May 19 2026Comprehensive overview of NVIDIA Riva covering real-time speech AI, Automatic Speech Recognition (ASR), Natural Language Processing (NLP), Text-to-Speech (TTS), multilingual conversational AI, custom model deployment, GPU acceleration, Kubernetes deployment, and production-grade voice AI architectures.NVIDIA Riva Speech AI Conversational AI Automatic Speech Recognition ASR Text-to-Speech TTS Natural Language Processing NLP Voice AI Real-Time AI GPU Acceleration CUDA AI Infrastructure Kubernetes AI Inference Deep Learning Production AI Edge AI
15NVIDIA NGC Catalog: GPU Optimized Containers, AI Models and Enterprise AI InfrastructureTue May 19 2026Comprehensive overview of the NVIDIA NGC Catalog covering GPU optimized containers, CUDA and TensorRT environments, NeMo and Triton deployments, pretrained AI models, Kubernetes integration, NVIDIA NIM microservices, and enterprise-scale AI infrastructure for accelerated computing workloads.NVIDIA NGC NVIDIA NGC Catalog CUDA TensorRT TensorRT-LLM Triton NeMo Kubernetes GPU Containers AI Infrastructure Accelerated Computing NVIDIA NIM GPU Clusters AI Deployment Deep Learning Distributed Computing AI Platform Engineering Production AI Docker
16NVIDIA DCGM: GPU Health, Diagnostics, and Prometheus MetricsWed Jul 29 2026How NVIDIA's Data Center GPU Manager works β€” the nv-hostengine daemon, key DCGM metric field IDs, XID error codes and diagnostic levels, the DCGM Exporter DaemonSet and Prometheus integration on Kubernetes, PromQL queries and alert rules for GPU clusters, DCGM-driven KEDA autoscaling, and MIG instance monitoring.NVIDIA DCGM GPU Monitoring Prometheus Observability Kubernetes MIG MLOps
17NVIDIA Base Command Manager: Provisioning and Operating GPU ClustersWed Jul 29 2026How NVIDIA Base Command Manager (BCM) operates an entire GPU cluster β€” bare metal provisioning and node imaging, the head node vs compute node architecture, Slurm and Kubernetes workload manager integration, user and group management, the cmsh CLI and REST API, and how BCM, DCGM, and SMI fit together at different layers of the stack.NVIDIA BCM Base Command Manager Cluster Management HPC Slurm Provisioning MLOps
18Slurm: The HPC Workload Manager Behind AI Training ClustersThu Jul 30 2026Slurm from the ground up β€” the controller/node-daemon architecture, partitions and QOS, job submission with sbatch/srun/salloc, GPU allocation with GRES, multi-node MPI jobs, running containers under Slurm with enroot and pyxis, multifactor job priority, preemption, job arrays and dependencies, node health and cgroup enforcement, topology-aware scheduling, the slurmrestd API, Slurm vs Kubernetes, and a set of interview questions with answers.Slurm HPC NVIDIA MPI GPU Job Scheduling enroot AI MLOps
19AI Infra Networking: GPU Clusters, InfiniBand, RoCE, and DPU IntegrationFri Feb 27 2026Networking fundamentals for AI-centric data centers β€” the four network planes, DMA and RDMA mechanics, InfiniBand vs RoCE vs Ethernet with real numbers, the GPU interconnect hierarchy from PCIe through NVLink/NVSwitch to InfiniBand, BlueField DPUs, and how the GPU and Network Operators automate all of it on Kubernetes.NVIDIA AI Infrastructure GPU Clusters Data Center AI Networking InfiniBand RoCE DPU BlueField RDMA Accelerated Computing
20AI Infra Storage: NVMe, Parallel File Systems, Object Storage, and GPUDirect StorageFri Feb 27 2026Storage architectures for AI infrastructure β€” the hot/warm/cold tiering model with real throughput numbers, GPUDirect Storage's direct path from NVMe to GPU memory, NVMe-oF, checkpoint math for large models, erasure coding for durability, and cloud vs on-prem storage tradeoffs.NVIDIA AI Infrastructure Storage NVMe Parallel File Systems Object Storage GPUDirect Storage Checkpointing On-Prem vs Cloud Accelerated Computing
21AI/ML OperationsFri Feb 27 2026Comprehensive overview of monitoring and operations for AI infrastructure, covering GPU monitoring tools (DCGM, BCM), infrastructure monitoring (Prometheus, Grafana), cluster orchestration (Kubernetes, Slurm), power and cooling monitoring, high availability, failure scenarios, security monitoring, GPU utilization optimization, capacity planning, multi-GPU scaling strategies, lifecycle management, logging systems, and alerting best practices.NVIDIA AI Operations GPU Monitoring Data Center Management Cluster Orchestration Kubernetes Job Scheduling GPU Virtualization vGPU MIG Observability MLOps
Hitesh Sahu
Written by Hitesh Sahu, a passionate developer and blogger.

Sun Aug 23 2026

Share This on

AI-Infrastructure/0-INDEX
Let's work together
hiteshkrsahu@gmail.com
Munich πŸ₯¨, Germany πŸ‡©πŸ‡ͺ, EU
Playstore
Hitesh Sahu's apps on Google Play Store
Need Help?
Let's Connect
Navigation
Β  Home/About
Β  Skills
Β  Work/Projects
Β  Lab/Experiments
Β  Contribution
Β  Awards
Β  Art/Sketches
Β  Thoughts
Β  Contact
Links
Β  Sitemap
Β  Legal Notice
Β  Privacy Policy

Made with

NextJS logo

NextJS by

hitesh Sahu

| Β© 2026 All rights reserved.