OVDC: Disk Space Exhaustion#

Overview#

OVDC stores derived data on persistent volumes backed by Persistent Volume Claims (PVCs). When the volume attached to OVDC is too small or fills too quickly, pods may crash, enter a crash loop, or report errors writing to disk. OVDC requires persistent storage to maintain derived data, and disk space exhaustion prevents the service from operating correctly.

The storage limit is controlled by settings.store.maxSize, which defaults to "AUTO"—the whole storage.volume.size less min(20%, 100GB) of headroom for breathing room, compaction, and WAL (Write-Ahead Log) files. If set explicitly instead of AUTO, keep it comfortably below the PVC size—nothing validates the relationship for you. When disk space is exhausted, RocksDB cannot write new data, flush write buffers, or perform compaction operations, causing write stalls, cache eviction, and service failures.

Disk space exhaustion can occur when:

  • Persistent volume allocation is insufficient for workload data growth

  • Data growth outpaces garbage collection and cleanup mechanisms

  • Retention policies are misconfigured or not functioning

  • Garbage collection is not running or is insufficient

  • Multiple OVDC pods competing for limited storage resources

  • RocksDB compaction backlog consuming excessive disk space

  • WAL files accumulating without cleanup

When disk space is exhausted, OVDC pods may fail to start, crash during operation, or report errors when attempting to write data, resulting in service unavailability and degraded rendering performance.

Symptoms and Detection Signals#

Visible Symptoms#

  • OVDC pods crash - Pods failing with disk-related errors

  • Crash loop backoff - Pods repeatedly restarting due to disk space issues

  • Errors writing to disk - No dedicated log message exists for this; see ovdc_error{kind="rocks_io_error"} under Metric Signals below

  • Service unavailability - OVDC unable to serve requests due to disk issues

  • Write stalls - RocksDB stalling writes due to insufficient disk space

  • Cache eviction - Cache eviction occurring due to disk pressure

Metric Signals#

OVDC does not log a dedicated disk-space-exhaustion message. RocksDB I/O failures, including out-of-space conditions, surface as the rocks_io_error kind on the ovdc_error metric below, not as a distinct log string.

The following metrics can be used to detect disk space exhaustion before it causes service failures. Monitor these metrics proactively to identify capacity issues early. See the Metrics Reference for the full metric/type/label definitions.

Volume Capacity Metrics#

kubelet_volume_stats_available_bytes{persistentvolumeclaim="ovdc-*", namespace="ovdc"}
kubelet_volume_stats_capacity_bytes{persistentvolumeclaim="ovdc-*", namespace="ovdc"}
kubelet_volume_stats_used_bytes{persistentvolumeclaim="ovdc-*", namespace="ovdc"}

Persistent volume utilization is tracked through kubelet metrics that monitor available, used, and total capacity. The kubelet_volume_stats_available_bytes metric reports remaining disk space, where low values indicate approaching exhaustion. Total volume size is provided by kubelet_volume_stats_capacity_bytes, which serves as the baseline for calculating usage percentage. Current disk consumption is tracked in kubelet_volume_stats_used_bytes, where high values relative to capacity indicate the volume is approaching its limits. Alert when usage exceeds 80% of capacity to allow time for remediation before exhaustion.

RocksDB Write Stalls from Disk Pressure#

Disk space constraints can trigger RocksDB write stalls, exposed as ovdc_rocks_intrinsic_gauge{name="...", cf="..."}—see RocksDB Property Gauges for the full list of name values, including __rocksdb_stalls_total_stop, __rocksdb_stalls_total_delays, and the pending-compaction-bytes variants. Stalls indicate RocksDB is throttling writes, potentially due to insufficient disk space for compaction or new writes.

Database and I/O Performance Metrics#

ovdc_error{kind="rocks_io_error"}
ovdc_io_time{io="rocksdb_write"}

ovdc_error{kind="rocks_io_error"} is the RocksDB I/O-error counter and is the closest signal to a disk-space failure—see Error Metrics for the full kind list. Sharp increases in this counter correlate with disk exhaustion scenarios. Write I/O latency captured in ovdc_io_time{io="rocksdb_write"} (label io, not io_kind) may show elevated values during disk space pressure, as writes slow down or stall when the disk approaches capacity.

Cache Performance Impact#

ovdc_miss{level="2"}
ovdc_query{level="2"}

Disk space pressure can force cache eviction, manifesting in cache performance metrics. ovdc_miss{level="2"} may show elevated cache misses as disk space constraints force removal of cached data. Similarly, ovdc_query{level="2"} (RocksDB reads that may hit disk) may increase relative to level="0" / level="1" (row-cache and block-cache-only) queries, indicating cache eviction due to disk pressure forces more disk reads instead of serving from cache.

Root Cause Analysis#

Known Causes#

Disk space exhaustion in OVDC is typically caused by insufficient persistent volume allocation, data growth outpacing cleanup mechanisms, or misconfigured retention policies.

Insufficient Persistent Volume Allocation#

PVCs may be undersized for the workload data growth. OVDC stores derived data that can be many times the size of source content. If PVCs are not sized appropriately, they will fill to capacity quickly. settings.store.maxSize: "AUTO" derives headroom from storage.volume.size automatically, but if the PVC itself is too small, even proper configuration cannot prevent exhaustion.

Check PVC sizes and usage:

# Check PVC capacity and usage
kubectl get pvc -n ovdc
kubectl describe pvc -n ovdc

# Check actual disk usage within pods
kubectl exec -n ovdc <ovdc-pod-name> -- df -h

# Check settings.store.maxSize configuration
helm get values ovdc -n ovdc -o yaml | grep -A 2 "maxSize"

# Compare PVC size to settings.store.maxSize (AUTO reserves min(20%, 100GB))

Data Growth Outpacing Cleanup#

Data growth may outpace garbage collection and cleanup mechanisms. If garbage collection is not running frequently enough, is misconfigured, or cannot free sufficient space, disk usage will continue to grow until exhaustion occurs. RocksDB compaction may also be unable to keep up with write volume, causing accumulation of stale SST files.

Check garbage collection and compaction:

# Check garbage collection configuration
helm get values ovdc -n ovdc -o yaml | grep -A 10 "garbageCollection"

# Check if GC is running
kubectl logs -n ovdc <ovdc-pod-name> | grep -i "garbage\|gc\|cleanup"

# Review RocksDB compaction settings
helm get values ovdc -n ovdc | grep -A 5 "periodic_compaction"

# Monitor disk usage trends
kubectl exec -n ovdc <ovdc-pod-name> -- df -h

Other Possible Causes#

  1. Storage Limit Misconfiguration

    • settings.store.maxSize set explicitly and too high relative to PVC size (prefer AUTO)

    • Explicit maxSize not accounting for WAL and compaction overhead

    • Multiple OVDC pods sharing storage resources without proper limits

  2. RocksDB Compaction Issues

    • Compaction not running efficiently due to disk space constraints

    • Stale SST files not being removed due to insufficient space

    • Compaction backlog causing storage growth faster than cleanup

    • Pending compaction bytes exceeding limits due to slow disk I/O

  3. Garbage Collection Configuration Issues

    • garbageCollection.minFreeCapacity set too high, preventing GC from running when needed

    • garbageCollection.deleteKeyspaceQuantile set too low, not freeing enough space

    • GC’s capacity check runs on a fixed 5-second interval (not configurable)—this is not a tunable cause

    • GC not running due to configuration errors or service issues

  4. Node-Level Storage Issues

    • Node running out of disk space affecting all pods

    • Multiple pods on same node competing for storage

    • Storage backend performance issues preventing efficient cleanup

    • Ephemeral storage limits being exceeded

Troubleshooting Steps#

Diagnostic Steps for Known Root Causes#

1. Monitor Disk Usage Metrics and Set Up Alerts#

Monitor disk usage metrics to detect exhaustion before it causes service failures.

# Check PVC capacity and usage
kubectl get pvc -n ovdc -o wide

# Get detailed PVC information
kubectl describe pvc -n ovdc <pvc-name>

# Check disk usage from within pods
kubectl exec -n ovdc <ovdc-pod-name> -- df -h

# Query Prometheus metrics for disk usage
# kubelet_volume_stats_available_bytes{pvc_name="<pvc-name>", namespace="ovdc"}
# kubelet_volume_stats_capacity_bytes{pvc_name="<pvc-name>", namespace="ovdc"}
# kubelet_volume_stats_used_bytes{pvc_name="<pvc-name>", namespace="ovdc"}

# Calculate usage percentage
# usage_percent = (used_bytes / capacity_bytes) * 100

# Check OVDC storage usage metrics
kubectl port-forward -n ovdc svc/<ovdc-service-name> 3051:3051
curl http://localhost:3051/metrics | grep -i "storage\|disk\|size"

Analysis:

  • PVCs showing high usage (>90%) indicate capacity exhaustion risk.

  • Compare kubelet_volume_stats_used_bytes against kubelet_volume_stats_capacity_bytes.

  • OVDC storage usage approaching settings.store.maxSize indicates exhaustion.

  • Rapid increases in disk usage indicate data growth outpacing cleanup.

Resolution:

  • Set up alerts for PVC usage exceeding 80% of capacity.

  • Monitor OVDC storage usage and alert when approaching settings.store.maxSize.

  • Track disk usage trends to predict when capacity increases are needed.

  • Expand PVCs if storage class supports volume expansion (see step 3).

2. Expand Persistent Volume Claims (PVCs) as Needed#

If disk space is exhausted or approaching limits, expand PVCs to provide additional capacity.

# Check current PVC sizes
kubectl get pvc -n ovdc

# Check if storage class supports volume expansion
kubectl get storageclass <storage-class-name> -o yaml | grep allowVolumeExpansion

# If expansion is supported, edit PVC to request larger size
# Note: Some storage classes require manual PVC edit, others support dynamic expansion
kubectl edit pvc <pvc-name> -n ovdc
# Change spec.resources.requests.storage to larger value

# After expansion, update settings.store.maxSize in Helm values to match
# (or leave it as "AUTO" so it re-derives from the new storage.volume.size)
helm get values ovdc -n ovdc -o yaml > current-values.yaml
# Edit: settings.store.maxSize (or leave as "AUTO")
# Example: If PVC is 100GB, AUTO derives ~90GB (100GB less min(20%,100GB) headroom)

# Apply updated values
helm upgrade ovdc omniverse/ovderivedcache --version 6.0.0 --namespace ovdc -f current-values.yaml

Analysis:

  • PVCs at or near capacity require expansion.

  • Storage classes with allowVolumeExpansion: true support dynamic expansion.

  • After PVC expansion, settings.store.maxSize should be reviewed—AUTO re-derives automatically; an explicit value must be updated by hand.

Resolution:

  • Expand PVCs if storage class supports volume expansion.

  • For storage classes without expansion support, create new larger PVCs and migrate data.

  • Update settings.store.maxSize in Helm values after expansion, or use AUTO so it re-derives automatically.

  • Apply updated values: helm upgrade ovdc omniverse/ovderivedcache --version 6.0.0 --namespace ovdc -f current-values.yaml

  • Monitor disk usage after expansion to ensure adequate capacity.

Other Diagnostic Actions#

  • Check node-level storage: Verify nodes have sufficient storage capacity:

    kubectl describe nodes | grep -A 5 "Allocated resources"
    kubectl top nodes
    # Check for node-level disk pressure
    kubectl describe nodes | grep -i "disk\|storage\|pressure"
    
  • Review storage class configuration: Verify storage class provides adequate capacity:

    kubectl get storageclass
    kubectl describe storageclass <storage-class-name>
    # Check volume expansion capabilities
    kubectl get storageclass <storage-class-name> -o yaml | grep allowVolumeExpansion
    

Prevention#

Proactive Monitoring#

Set up alerts for:

  • PVC capacity thresholds: Alert when PVC usage exceeds 80% of capacity

  • Storage limit thresholds: Alert when OVDC storage usage approaches settings.store.maxSize

  • Garbage collection failures: Alert when GC is not running or failing

  • Rapid disk growth: Alert on rapid increases in disk usage indicating potential issues

  • Node storage pressure: Alert when node-level storage is approaching limits

  • RocksDB write stalls: Alert when write stall metrics indicate disk space pressure

  • Cache eviction rates: Alert when cache miss rates increase due to disk pressure

Configuration Best Practices#

  • Size PVCs appropriately: Estimate storage requirements based on workload patterns and size PVCs with headroom (account for derived data being many times source content size)

  • Configure storage limits: Use settings.store.maxSize: "AUTO" (or an explicit value comfortably below PVC size) to allow for overhead, WAL files, and compaction

  • Enable garbage collection: Ensure GC is configured and running with appropriate thresholds (minFreeCapacity, deleteKeyspaceQuantile; the capacity check itself runs on a fixed 5-second interval)

  • Configure periodic compaction: Enable RocksDB periodic compaction to clean up stale SST files

  • Monitor disk usage trends: Track disk usage over time to predict when capacity increases are needed

  • Plan for data growth: Account for data growth as workloads scale and scenes become more complex

  • Use volume expansion: Prefer storage classes that support volume expansion for easier capacity management

Capacity Planning#

  • Estimate storage requirements: Calculate storage needs based on scene sizes, derived data multipliers (often 3-10x source content), and retention requirements

  • Plan for storage growth: Account for storage growth as workloads scale and new assets are introduced

  • Monitor storage trends: Track storage usage trends over time to predict when capacity increases are needed

  • Test retention policies: Validate GC and retention policies under expected production load

  • Review workload patterns: Analyze workload patterns to understand data growth rates and adjust capacity planning accordingly

  • Account for multiple pods: If running multiple OVDC pods, ensure total storage capacity accounts for all pods and their data growth