OVDC: RocksDB Corruption or Failures#

Overview#

When OVDC (Omniverse Derived Cache) experiences RocksDB corruption or failures, the service may fail to start, logs indicate database corruption, or data loss is observed. OVDC uses RocksDB as its persistent storage backend for derived data, and corruption can prevent the service from operating correctly.

OVDC relies on RocksDB for persistent storage of derived data. RocksDB is a persistent key-value store that provides durability and performance. When corruption occurs, RocksDB cannot read or write data correctly, causing service failures.

RocksDB corruption can occur when:

  • Unclean shutdowns preventing proper database closure

  • Disk errors or hardware failures affecting persistent volumes

  • File system corruption on persistent volumes

  • Insufficient disk space during critical operations

The only solution for corruption is to delete all PVCs and delete all pods to force a restart, which will result in data loss and require the cache to be rebuilt.

Symptoms and Detection Signals#

Visible Symptoms#

  • Service fails to start - OVDC pods unable to start due to database errors

  • Database corruption errors - Logs indicating RocksDB corruption or invalid data

  • Data loss - Missing or corrupted derived data in cache

  • Repeated pod restarts - Pods crashing repeatedly due to database errors

Log Messages#

Database Open Failures#

Where to find these logs:

  • Pod: ovdc-*

  • Location: OVDC Pod

  • Application: OVDC/RocksDB

  • Description: Errors indicating RocksDB cannot open the database

# LEVEL: Error/Fatal
# SOURCE: OVDC/RocksDB
kubernetes.pod_name: ovdc-* and
message: "*failed to open rocksdb database*"

Metric Signals#

See the Metrics Reference for the full ovdc_error metric definition.

RocksDB I/O errors are tracked through the ovdc_error metric with kind="rocks_io_error". This metric counts I/O errors reported by RocksDB during database operations. High values or sharp increases indicate disk I/O failures or data corruption that may prevent the database from reading or writing data correctly. Sustained I/O errors typically precede database corruption and service failures.

Root Cause Analysis#

Known Causes#

RocksDB corruption in OVDC is typically caused by unclean shutdowns, disk errors, or software bugs.

Unclean Shutdowns#

Unclean shutdowns occur when OVDC pods are terminated abruptly without allowing RocksDB to flush data and close the database properly. This can happen during node failures, forced pod deletions, or OOM kills. When RocksDB cannot close cleanly, WAL or SST files may be left in an inconsistent state, causing corruption.

Check for unclean shutdowns:

kubectl describe pod -n ovdc <ovdc-pod-name> | grep -A 10 "Events:"
kubectl get pods -n ovdc -l app.kubernetes.io/instance=ovdc -o wide
# Look for termination reasons: OOMKilled, Evicted, NodeShutdown

kubectl get events -n ovdc --sort-by='.lastTimestamp' | grep -i "evict\|kill\|terminate"

Disk Errors or Hardware Failures#

Disk errors or hardware failures on persistent volumes can cause data corruption. This may manifest as I/O errors, read/write failures, or corrupted data blocks. Persistent volumes backed by unreliable storage or experiencing hardware issues are more susceptible to corruption.

Check for disk errors:

kubectl logs -n ovdc <ovdc-pod-name> | grep -i "disk\|io\|error\|fail"
kubectl describe pvc -n ovdc
# Check for volume attachment issues or storage backend problems

Other Possible Causes#

  1. File System Corruption

    • File system corruption on persistent volumes

    • File system errors preventing proper writes

    • Mount issues causing data inconsistencies

  2. Insufficient Disk Space

    • Disk space exhaustion during critical operations

    • WAL files or SST files corrupted due to space constraints

    • Compaction failures due to insufficient space

Troubleshooting Steps#

Diagnostic Steps for Known Root Causes#

1. Delete All PVCs and Pods to Force Restart#

The only solution for RocksDB corruption is to delete all PVCs and delete all pods to force a restart. This will result in data loss, and the cache will be rebuilt as workloads run.

# List all OVDC PVCs
kubectl get pvc -n ovdc

# Delete all OVDC PVCs
# WARNING: This will result in complete data loss
kubectl delete pvc --all -n ovdc

# Delete all OVDC pods
kubectl delete pods -n ovdc -l app.kubernetes.io/instance=ovdc

# Verify pods are running and healthy
kubectl get pods -n ovdc -l app.kubernetes.io/instance=ovdc

Analysis:

  • Deleting PVCs removes all corrupted data.

  • Deleting pods forces recreation with new PVCs.

  • Cache will be rebuilt as workloads access OVDC.

Resolution:

  • Delete all PVCs to remove corrupted data (data will be lost).

  • Delete all pods to force restart with clean storage.

  • Monitor pod startup to ensure service is healthy.

  • Cache will rebuild automatically as workloads run (renders will be slower during rebuild).

Prevention#

Proactive Monitoring#

Set up alerts for:

  • RocksDB I/O errors: Alert when ovdc_error{kind=~"rocks_io_error"} increases

  • Pod termination events: Alert on OOM kills or forced pod terminations

  • Database corruption errors: Alert on corruption-related log messages

  • Pod startup failures: Alert when OVDC pods fail to start repeatedly