OVDC: RocksDB Corruption or Failures#
Overview#
When OVDC (Omniverse Derived Cache) experiences RocksDB corruption or failures, the service may fail to start, logs indicate database corruption, or data loss is observed. OVDC uses RocksDB as its persistent storage backend for derived data, and corruption can prevent the service from operating correctly.
OVDC relies on RocksDB for persistent storage of derived data. RocksDB is a persistent key-value store that provides durability and performance. When corruption occurs, RocksDB cannot read or write data correctly, causing service failures.
RocksDB corruption can occur when:
Unclean shutdowns preventing proper database closure
Disk errors or hardware failures affecting persistent volumes
File system corruption on persistent volumes
Insufficient disk space during critical operations
The only solution for corruption is to delete all PVCs and delete all pods to force a restart, which will result in data loss and require the cache to be rebuilt.
Symptoms and Detection Signals#
Visible Symptoms#
Service fails to start - OVDC pods unable to start due to database errors
Database corruption errors - Logs indicating RocksDB corruption or invalid data
Data loss - Missing or corrupted derived data in cache
Repeated pod restarts - Pods crashing repeatedly due to database errors
Log Messages#
Database Open Failures#
Where to find these logs:
Pod:
ovdc-*Location: OVDC Pod
Application: OVDC/RocksDB
Description: Errors indicating RocksDB cannot open the database
# LEVEL: Error/Fatal
# SOURCE: OVDC/RocksDB
kubernetes.pod_name: ovdc-* and
message: "*failed to open rocksdb database*"
Metric Signals#
See the Metrics Reference for the full
ovdc_error metric definition.
RocksDB I/O errors are tracked through the ovdc_error metric with
kind="rocks_io_error". This metric counts I/O errors reported by
RocksDB during database operations. High values or sharp increases
indicate disk I/O failures or data corruption that may prevent the
database from reading or writing data correctly. Sustained I/O errors
typically precede database corruption and service failures.
Root Cause Analysis#
Known Causes#
RocksDB corruption in OVDC is typically caused by unclean shutdowns, disk errors, or software bugs.
Unclean Shutdowns#
Unclean shutdowns occur when OVDC pods are terminated abruptly without allowing RocksDB to flush data and close the database properly. This can happen during node failures, forced pod deletions, or OOM kills. When RocksDB cannot close cleanly, WAL or SST files may be left in an inconsistent state, causing corruption.
Check for unclean shutdowns:
kubectl describe pod -n ovdc <ovdc-pod-name> | grep -A 10 "Events:"
kubectl get pods -n ovdc -l app.kubernetes.io/instance=ovdc -o wide
# Look for termination reasons: OOMKilled, Evicted, NodeShutdown
kubectl get events -n ovdc --sort-by='.lastTimestamp' | grep -i "evict\|kill\|terminate"
Disk Errors or Hardware Failures#
Disk errors or hardware failures on persistent volumes can cause data corruption. This may manifest as I/O errors, read/write failures, or corrupted data blocks. Persistent volumes backed by unreliable storage or experiencing hardware issues are more susceptible to corruption.
Check for disk errors:
kubectl logs -n ovdc <ovdc-pod-name> | grep -i "disk\|io\|error\|fail"
kubectl describe pvc -n ovdc
# Check for volume attachment issues or storage backend problems
Other Possible Causes#
File System Corruption
File system corruption on persistent volumes
File system errors preventing proper writes
Mount issues causing data inconsistencies
Insufficient Disk Space
Disk space exhaustion during critical operations
WAL files or SST files corrupted due to space constraints
Compaction failures due to insufficient space
Troubleshooting Steps#
Diagnostic Steps for Known Root Causes#
1. Delete All PVCs and Pods to Force Restart#
The only solution for RocksDB corruption is to delete all PVCs and delete all pods to force a restart. This will result in data loss, and the cache will be rebuilt as workloads run.
# List all OVDC PVCs
kubectl get pvc -n ovdc
# Delete all OVDC PVCs
# WARNING: This will result in complete data loss
kubectl delete pvc --all -n ovdc
# Delete all OVDC pods
kubectl delete pods -n ovdc -l app.kubernetes.io/instance=ovdc
# Verify pods are running and healthy
kubectl get pods -n ovdc -l app.kubernetes.io/instance=ovdc
Analysis:
Deleting PVCs removes all corrupted data.
Deleting pods forces recreation with new PVCs.
Cache will be rebuilt as workloads access OVDC.
Resolution:
Delete all PVCs to remove corrupted data (data will be lost).
Delete all pods to force restart with clean storage.
Monitor pod startup to ensure service is healthy.
Cache will rebuild automatically as workloads run (renders will be slower during rebuild).
Prevention#
Proactive Monitoring#
Set up alerts for:
RocksDB I/O errors: Alert when
ovdc_error{kind=~"rocks_io_error"}increasesPod termination events: Alert on OOM kills or forced pod terminations
Database corruption errors: Alert on corruption-related log messages
Pod startup failures: Alert when OVDC pods fail to start repeatedly