What hardware telemetry is exposed to customers?

What hardware telemetry is exposed to customers?

In 2026, OCI provides a deep look into the "physicality" of your rented servers. While traditional clouds often stop at CPU and RAM percentages, OCI exposes low-level hardware telemetry, especially for high-performance Bare Metal and GPU clusters where thermal and electrical stability are critical for workload success.

Here is the hardware telemetry currently exposed to customers:

1. GPU & AI Infrastructure Telemetry

For customers running AI training or 3D rendering, OCI Stack Monitoring provides a dedicated "Cluster Network" dashboard. In 2026, you can monitor:

  • Thermal Metrics: Real-time GPU Temperature and alerts for when a chip is nearing "thermal throttling" (slowing down to cool off).

  • Power Consumption: Total Wattage usage per GPU and per Host, allowing you to correlate power spikes with specific AI training jobs.

  • Physical Health: Fan speed utilization and ECC (Error Correction Code) memory errors, which can predict a hardware failure before it crashes your job.

  • Clock Speeds: Real-time monitoring of GPU clock utilization to ensure your "rented" silicon is performing at its advertised speed.

2. Bare Metal "Health Checks"

Since you own the entire physical box, OCI exposes infrastructure-level health metrics that are usually hidden in virtualized environments:

  • Instance Accessibility Status: A hardware-level ping (ARP-based) that tells you if the server is unresponsive even if the Operating System is still technically "up."

  • File System Anomaly Detection: OCI monitors the hardware kernel logs to detect if a local NVMe drive or a boot volume has entered a "read-only" or "anomaly" state due to physical media wear.

  • Infrastructure Health Metrics: A specific namespace (oci_compute_instance_health) that alerts you if there is an ongoing issue with the underlying power or cooling in your specific Fault Domain.

3. Compute Resource Telemetry (VM & Bare Metal)

Standard metrics included in every rental (at no extra cost) include:

  • CPU Utilization: Total activity level of the cores.

  • Memory Utilization: Current RAM usage (requires the Oracle Cloud Agent to be enabled).

  • Disk & Network I/O: Aggregate throughput (Bytes read/written) and IOPS (Input/Output operations per second) across all attached volumes and virtual network cards.


Hardware Telemetry Availability (2026)

Telemetry TypeVirtual Machines (VM)Bare Metal (BM)GPU Shapes
CPU/RAM/DiskYesYesYes
Hardware HealthLimited (Virtual)Full Physical ViewFull Physical View
TemperatureNoNoYes (GPU level)
Power DrawNoNoYes (GPU level)
ECC ErrorsNoYesYes

4. 2026 Innovation: "Telemetry Streaming"

In 2026, you don't have to stay in the OCI Console to see this data. You can use the OCI Service Connector Hub to stream hardware telemetry directly into your own tools.

  • OpenTelemetry Support: OCI hardware metrics are now compatible with OTLP (OpenTelemetry Protocol), allowing you to see your GPU temperatures and power draw right alongside your application traces in tools like Datadog, Splunk, or New Relic.

5. Why this telemetry matters for your "Rental"

Having access to power and thermal telemetry changes how you manage your costs. In 2026, some advanced OCI customers use "Power-Aware Scaling." If they see their GPU power consumption dropping while their workload is still running, it often indicates a "stalled" AI process—allowing them to kill the job and stop paying for idle, expensive hardware hours.

Looking for servers Rental ?

Call Our Expert :


  • (call for rental enquiries)

Email us :