01 — Technology · NADI
See every GPU. Find every lost hour.
NADI is an on-premises monitoring, observability and alerting platform for NVIDIA data-center GPU farms. It correlates GPU health, NVLink fabric, network fabric, scheduler jobs, tenants and energy cost in one place — inside your network, with no telemetry leaving your boundary.
The problem
Your fleet is busy. That is not the same as productive.
A GPU that is thermally throttling still reports as busy. A GPU allocated to an idle session still bills. A single NVLink with rising CRC errors halves the throughput of every job crossing it. None of this appears in a standard infrastructure monitoring tool, and all of it costs GPU-hours.
The metrics to catch it all exist. DCGM exposes most of them; Fabric Manager and the subnet manager expose the rest; the scheduler knows which job was running. What does not exist, in any packaged product, is the join between them — the chain that turns a throughput drop into a specific link, on a specific GPU serial, costing a specific number of GPU-hours.
That chain is what NADI is.
Why NADI
Four claims. Everything else is a feature.
The whole chain, joined
GPU → NVLink/NVSwitch → InfiniBand/RoCE → node → rack → job → tenant → energy cost. Every link in that chain is individually available somewhere. The joins are not packaged anywhere. Trace a throughput drop to the link and the GPU serial that caused it.
Zero telemetry egress
NADI runs entirely inside your network and is fully operable air-gapped, including licence activation. No component contacts an external service. No GPU telemetry crosses your boundary — not by default, not at all.
Reports in money, not metrics
Stranded capacity in GPU-hours per week. Energy cost per job. Cost per tenant. Warranty- recoverable value with the evidence attached. Infrastructure teams justify budget in the language of finance; the tool should speak it.
In-region, in your time zone
Implementation, acceptance testing and support delivered from Malaysia, contracting through a local Sdn Bhd entity, with on-site escalation for Klang Valley, Johor and Singapore.
Capabilities
What it watches.
GPU telemetry and health
- Per-GPU SM and compute utilisation, HBM utilisation and capacity
- Core, HBM and hotspot temperatures
- Power draw, enforced limit, cap violations, cumulative energy
- SM and memory clocks against applicable maxima
- PCIe throughput and replay counters
- Profiling metrics where available — SM active, SM occupancy, tensor pipe active — to separate allocated from productive
- Rack → node → GPU topology with aggregation at every level
- Fleet, rack and node health scores from a published, inspectable formula
Reliability and fault detection
- Correctable and uncorrectable ECC counts, volatile and aggregate
- XID capture from both DCGM and the kernel log, with a built-in reference table of codes and severities
- Row-remapping status: correctable, uncorrectable, pending, failure
- Throttle and clock-limit reasons decoded by name, with time-in-state
- Fault history keyed to serial number — surviving reboots, agent restarts and re-seating
- GPU offline, node unreachable and fallen-off-the-bus detection
Fabric and network
- Per-NVLink status, bandwidth, CRC, replay and recovery counters
- NVSwitch port state and counters via NVIDIA Fabric Manager
- Link-count anomaly detection — a GPU with fewer active links than its peers
- InfiniBand port health, error and congestion counters, with host-side fallback
- RoCE port health including PFC pause frames, ECN marks and discards
- Fabric ports correlated to the nodes and jobs they serve
Workload observability
- Slurm, Kubernetes with the NVIDIA GPU Operator, and VDI broker integration
- Per-job GPU-hours and utilisation efficiency, with the methodology shown
- Queue and wait times by partition, account, user and requested GPU count
- Stranded-capacity detection — allocated but not working — quantified in GPU-hours and currency
- Hardware events attributed to the jobs that were running at the time
- Multi-node performance skew detection
Alerting
- Threshold rules on any metric, at any scope, with duration conditions
- Default rule pack for temperature, power, ECC, XID, remap, throttling, offline, utilisation, NVLink and version drift
- Routing to email, Slack and generic webhook by severity, scope, label and time of day
- Grouping, deduplication, flap detection, escalation, maintenance suppression
- Alert history, acknowledgement and silencing — all audited
Security and access
- OIDC single sign-on and LDAP/Active Directory with group-to-role mapping
- Six roles: Administrator, Operator, Analyst, Viewer, Tenant Viewer, Auditor
- Topology- and tenant-scoped visibility, enforced server-side on every request
- Scoped, expiring, revocable API tokens
- Append-only audit log with syslog and SIEM forwarding
- Read-only against the GPU — NADI never sets clocks, changes power limits or resets a device
Coexistence
Keep your stack. Add the GPU domain.
NADI is not a replacement for Prometheus, Grafana or Datadog. It is the GPU-domain system of record that feeds them. It ships a Prometheus-compatible query endpoint and a Grafana-compatible data source, so your existing dashboards keep working and your existing Alertmanager routes keep firing.
Nor is it a scheduler. It is the layer that tells you whether your scheduler’s decisions were any good.
Licensing is per GPU, with unlimited users and unlimited metric series — no per-host penalty for dense nodes. Licence enforcement never stops collection, display or alerting.
Status: in active development
NADI version 1 is being built and delivered in four phases under an anchor engagement with a regional GPU infrastructure operator, each phase gated on formal user acceptance testing. The platform is not yet generally available.
We are talking to a small number of design partners operating fleets in the 256–4,096 GPU range — on-premises or colocation, NVLink or NVSwitch, InfiniBand or RoCE, running Slurm or Kubernetes. If that is your estate and the problems above are yours, we would like to hear from you.
NADI, NADI Core, NADI Pro and NADI Enterprise are product names of Jirisan Ventures Sdn Bhd. NVIDIA, DCGM, NVLink, NVSwitch and CUDA are trademarks of NVIDIA Corporation. Prometheus and Kubernetes are trademarks of The Linux Foundation. Grafana is a trademark of Raintank, Inc. d/b/a Grafana Labs. Slurm is a trademark of SchedMD LLC. Use of these names is descriptive of interoperability and does not imply endorsement or affiliation.