OVDC: Network Bandwidth or Latency Bottlenecks#

Overview#

When OVDC (Omniverse Derived Cache) experiences network bandwidth or latency bottlenecks, cache population is slow, throughput between OVDC and GPU nodes is degraded, or scene load times increase. OVDC requires high bandwidth and low latency to efficiently serve derived data to render workers over the network.

OVDC serves derived data to render workers over the network. The service requires high bandwidth and low latency to support efficient data transfer. It is generally recommended to have 3.3Gbps of network bandwidth available for each GPU in a cluster, with a minimum of 1Gbps per GPU.

Note

The 3.3 Gbps-per-GPU figure is operational guidance carried over from earlier DDCS-era docs and could not be confirmed against the ovderivedcache source repository as part of this review—treat it as a starting heuristic, not a verified spec. The repo’s own confirmed capacity numbers are node NICs ≥12.5 Gbps and roughly 60,000 requests/second per OVDC instance; see the OVDC Configuration Guide for the sourced sizing table.

Network bandwidth or latency bottlenecks can occur when:

  • Insufficient OVDC pods scheduled to handle network load

  • Insufficient network bandwidth between OVDC and GPU nodes

  • Network congestion from noisy neighbors in shared environments

  • Misconfigured Kubernetes networking causing suboptimal routing

  • Cross-availability zone or cross-data center traffic

When network bottlenecks occur, OVDC cannot efficiently serve data to render workers, causing slow cache population, degraded throughput, and increased scene load times.

Symptoms and Detection Signals#

Visible Symptoms#

  • Slow cache population - Cache taking longer than expected to populate with derived data

  • Degraded throughput - Reduced data transfer rates between OVDC and GPU nodes

  • Increased scene load times - Scene loads taking significantly longer than expected

Metric Signals#

The following metrics can be used to detect network bandwidth or latency bottlenecks. Review these metrics to determine if networking is a problem. See the Metrics Reference for the full metric/type/label definitions.

ovdc_bytes_returned{}
container_network_receive_bytes_total{pod=~"ovdc-.*"}
container_network_transmit_bytes_total{pod=~"ovdc-.*"}

Network saturation can be identified through several complementary metrics. The ovdc_bytes_returned metric tracks total bytes returned from all cache levels (in-memory and disk), providing application-level visibility into data throughput. When this metric shows high values relative to available network capacity, it suggests OVDC may be saturating network bandwidth serving cached content to clients.

Container-level network metrics provide direct visibility into network interface utilization. The container_network_transmit_bytes_total metric tracks bytes transmitted by OVDC pods, where high values approaching network interface limits indicate outbound network saturation as OVDC serves data to clients. Complementing this, container_network_receive_bytes_total measures inbound traffic to OVDC pods. By monitoring both transmit and receive metrics together, you can identify whether OVDC pods are network-bound and approaching the physical limits of their network interfaces. Compare these values against the network interface capacity of your VM SKUs to determine if network scaling is needed.

Root Cause Analysis#

Known Causes#

Network bandwidth or latency bottlenecks in OVDC are typically caused by insufficient OVDC pods scheduled, insufficient network bandwidth, or network congestion.

Insufficient OVDC Pods Scheduled#

OVDC must be configured with replica count equal to the number of compute nodes calculated based on network bandwidth requirements. The rule of thumb is to provision at least 3.3 Gbps of bandwidth per GPU. If there are not enough OVDC pods scheduled, the available pods may become network-bound, causing bottlenecks. Refer to the OVDC Configuration Guide for scaling guidance.

Check OVDC pod count:

# Check current OVDC pod count
kubectl get pods -n ovdc -l app.kubernetes.io/instance=ovdc

# Check StatefulSet replica configuration
kubectl get statefulset -n ovdc -l app.kubernetes.io/instance=ovdc
kubectl describe statefulset -n ovdc <ovdc-statefulset-name>

# Calculate required OVDC pods
# Required pods = (GPU count * 3.3 Gbps) / compute_node_bandwidth
# Example: 25 GPUs * 3.3 Gbps = 82.5 Gbps / 10 Gbps per node = ~9 pods required

Insufficient Network Bandwidth#

Network bandwidth may be insufficient for the workload. OVDC requires high bandwidth to serve derived data efficiently. It is recommended to have 3.3Gbps per GPU, with a minimum of 1Gbps.

Network Congestion#

In shared network environments, other workloads may consume bandwidth, causing congestion and degraded performance for OVDC traffic.

Other Possible Causes#

  1. Misconfigured Kubernetes Networking

    • Suboptimal routing causing increased latency

    • Network policies limiting throughput

    • Cross-availability zone traffic

  2. Cross-Availability Zone or Cross-Region Traffic

    • Pods in different availability zones causing higher latency

    • Cross-region traffic increasing latency significantly

    • Public internet transit instead of private networking

  3. Node-Level Network Issues

    • Node network interface problems

    • Network driver issues

    • Hardware network limitations

Troubleshooting Steps#

Diagnostic Steps for Known Root Causes#

1. Check OVDC Pod Count and Scaling#

Verify there are sufficient OVDC pods scheduled to handle network load. Refer to the OVDC Configuration Guide for scaling requirements.

# Check current OVDC pod count
kubectl get pods -n ovdc -l app.kubernetes.io/instance=ovdc

# Calculate required OVDC pods
# Required pods = (GPU count * 3.3 Gbps) / compute_node_bandwidth
# Example: 25 GPUs * 3.3 Gbps = 82.5 Gbps / 10 Gbps per node = ~9 pods required

# Check network metrics per pod
# Query: container_network_transmit_bytes_total{pod=~"ovdc-.*"}
# Query: container_network_receive_bytes_total{pod=~"ovdc-.*"}

Analysis:

  • Pod count below requirements indicates insufficient scaling.

  • High network utilization per pod suggests pods are network-bound.

  • Network metrics approaching interface limits indicate saturation.

Resolution:

  • Scale OVDC StatefulSet to match calculated pod requirements.

  • Refer to OVDC Configuration Guide for scaling guidance.

  • Update Helm values: helm upgrade ovdc omniverse/ovderivedcache --version 6.0.0 --namespace ovdc -f values.yaml

  • Monitor network metrics after scaling to verify improvement.

2. Monitor Network Metrics#

Review network metrics to identify bandwidth saturation or bottlenecks.

# Query OVDC bytes returned metric
# ovdc_bytes_returned

# Query container network metrics
# container_network_receive_bytes_total{pod=~"ovdc-.*"}
# container_network_transmit_bytes_total{pod=~"ovdc-.*"}

# Calculate network utilization
# Compare transmit/receive bytes against network interface capacity
# Check if metrics are approaching interface limits

Analysis:

  • High ovdc_bytes_returned relative to network capacity indicates potential bottlenecks.

  • Network transmit/receive bytes approaching interface limits indicate saturation.

  • High utilization across nodes suggests network congestion.

Resolution:

  • If network is saturated, scale OVDC pods to distribute load (see step 1).

  • Upgrade to VM SKUs with higher NIC speeds if bandwidth is insufficient.

  • Investigate network congestion from other workloads.

Other Diagnostic Actions#

  • Check pod placement: Verify OVDC and GPU pods are optimally placed:

    kubectl get pods -n ovdc -o wide
    kubectl get pods -n <workload-namespace> -o wide
    # Ensure pods are in same availability zone for optimal performance
    
  • Review network policies: Check if network policies are affecting performance:

    kubectl get networkpolicies -n ovdc
    kubectl describe networkpolicy <policy-name> -n ovdc
    
  • Monitor network trends: Track network performance over time:

    # Use cloud provider network metrics
    # Monitor throughput, latency, and error rates
    # Identify patterns and trends
    

Prevention#

Proactive Monitoring#

Set up alerts for:

  • Network bandwidth thresholds: Alert when network utilization exceeds 80% of available bandwidth

  • OVDC pod count: Alert when OVDC pod count is below calculated requirements

  • Network saturation: Alert when container_network_transmit_bytes_total or container_network_receive_bytes_total approach interface limits

  • High bytes returned: Alert when ovdc_bytes_returned indicates potential network bottlenecks