OVDC: Configure#

This guide provides detailed information on the Omniverse Derived Cache (OVDC) component and its use in a self-hosted NVCF cluster.

OVDC reduces scene load time and improves performance when properly configured and sized for the workload. Derived data generation is computationally expensive and time-consuming. OVDC trades network bandwidth for compute time by caching derived data, allowing multiple GPUs to share pre-generated content.

values.yaml has two halves, and the distinction matters: settings.* mirrors the application’s own config.toml schema and is rendered key-by-key into the chart’s ConfigMap—every key there must be one the binary deserializes, and the schema is closed (an unrecognised key fails the render). Everything else is ordinary Kubernetes-level chart configuration.

Base Configuration#

OVDC requires some configuration to be properly installed. Create a file on your local machine called values.yaml. A base configuration is provided in the following dropdown.

image:
  pullSecrets:

    - name: regcred

replicas: 1
selfAntiAffinity: true
affinity:
  nodeAffinity:
    preferredDuringSchedulingIgnoredDuringExecution:

      - weight: 100
        preference:
          matchExpressions:

            - key: node-type
              operator: In
              values:

                - compute

resources:
  requests:
    # Should cover engine.cacheSize + engine.blockCacheSize + write buffers + ~6GB TCP buffer.
    memory: 56G

storage:
  volume:
    size: 330Gi
    storageClassName: "gp3"

settings:
  logFormat: "json"
  logLevel: "info"
  store:
    maxSize: AUTO
  engine:
    cacheSize: "32G"
    blockCacheSize: "8G"
    increaseParallelism: 8
    cf:
      max_write_buffer_number: 24

Base Configuration Field Reference#

Field

Type

Required

Default

Description

image.pu llSecrets

list of {name}

No

[{name: regcred}]

Kubernetes imagePu llSecrets entries; each name must reference an existing ` kubernetes .io/dockerc onfigjson` Secret in the target namespace.

` replicas`

integer

No

1

Number of OVDC pods to deploy.

selfAnt iAffinity

boolean

No

true

When true, the scheduler prefers to place OVDC pods on nodes that do not already run one.

` affinity`

Kubernetes affinity object

No

{}

Merged over the chart’s self-an ti-affinity rule; values set here win on conflict.

resou rces.reques ts.memory

quantity string

No

56G

Should cover setti ngs.engine. cacheSize + ` settings.e ngine.block CacheSize` + write buffers + ~6GB TCP buffer.

` storage.vo lume.size`

quantity string

No

300Gi

Size of the Persistent VolumeClaim created per pod. `` volumeClaim Templates`` are immutab le—changing this on an existing release is rejected on upgrade.

`` storage.vol ume.storage ClassName``

string

No

ma naged-csi

Must support ReadW riteOnce; select a class with sufficient IOPS /throughput (see the 800 MB/s guidance above).

settings. logFormat

string enum

No

json

One of json, human, human_ no_color.

settings .logLevel

string enum

No

info

One of error, warn, info, debug, trace, off.

se ttings.stor e.maxSize

quantity string or "AUTO"

No

AUTO

RocksDB on-disk target. AUTO derives from ` storage.vo lume.size` less min(20 %, 100GB) headroom; requires storage.v olume.enabl ed: true.

setti ngs.engine. cacheSize

quantity string

No

32G

Size of the in-process row cache. 0 disables it.

` settings.e ngine.block CacheSize`

quantity string

No

8G

Size of the RocksDB LRU block cache.

sett ings.engine .increasePa rallelism

integer

No

8

RocksDB background threads for flushes and c ompactions.

`` settings.en gine.cf.max _write_buff er_number``

integer

No

24

Number of in-memory write buffers (64MB each by default) that absorb write bursts before flushing to disk.

sett ings.teleme try.prometh eusMetricsE xposition

boolean

No

true

When true, starts the Prometheus metrics HTTP service on pr ometheusMet ricsPort.

settings. telemetry.p rometheusMe tricsPort

integer

No

3051

Port the Prometheus metrics service binds.

For the complete set of settings.* keys, including RocksDB db / cf / block option maps, gRPC, TLS, JWT, and OpenTelemetry configuration, see the full values reference.

Complete Configuration Reference#

The base configuration above covers the essential settings for most deployments. For advanced configuration options or to explore all available settings, refer to the complete values file below. This reference includes all configuration options available in the OVDC Helm chart, including advanced settings for TLS, OpenTelemetry, and resource management.

1. Provision and Scale#

Proper provisioning and scaling are critical for OVDC performance. When undersized or misconfigured for the workload, simulations will slow down or even fail.

OVDC’s data directory is typically a network-attached PersistentVolumeClaim rather than local NVMe, and its sustained throughput is the binding operational constraint:

Important

Provision network-attached PVCs for at least 800 MB/s sustained throughput. Below that, background compaction by RocksDB (the embedded log-structured storage engine OVDC uses to persist cached data) cannot retire the write amplification a loaded node generates, the backlog grows, and the store spends its time in write stalls regardless of tuning. Many cloud volume types offer burst credits that exhaust under sustained load—size for the sustained figure, not the burst figure.

Reference sizing for a single OVDC replica:

Resource

Guidance

Compute SKU

Comparable to Azure E8s_v5–E20s_v5 or AWS r8g.12xlarge–24xlarge

Memory

~8 GB DRAM per vCPU

Remote disk

2x–8x the total scene data

Network

Node NICs ≥ 12.5 Gbps

Throughput

~60,000 requests/second per instance; scale replicas linearly

AWS EKS Example#

Assuming the SKU for compute nodes is r8g.12xlarge.

SKU

vCPU

RAM

NIC

r8g.12xlarge (compute)

48

384 G

20 Gbps

A single OVDC pod should be placed on each compute node. The goal is to expand network and storage-throughput capability with each pod, distributing cache load across multiple nodes.

Schedule ‘3’ OVDC Pods

yaml title="values.yaml" replicas: 3

2. Memory Configuration#

OVDC memory configuration includes the Kubernetes container resources block, along with engine-level cache settings under settings.engine. The default configuration allocates memory for the row cache, block cache, and write buffers.

resources:
  requests:
    memory: 56G

settings:
  engine:
    cacheSize: "32G"
    blockCacheSize: "8G"
    increaseParallelism: 8
    threadsHigh: 4
    threadsLow: 4
    cf:
      write_buffer_size: "64M"
      max_write_buffer_number: 24

Advanced Tuning

The settings above determine how memory is allocated by the cache. It can be controlled with these keys:

  • Row Cache (settings.engine.cacheSize): OVDC’s own in-process cache holding recently read entries, not a RocksDB cache. Defaults to 32G. 0 disables it.

  • Block Cache (settings.engine.blockCacheSize): RocksDB’s LRU block cache for metadata, filters, and file blocks. Defaults to 8G. Also backs the write buffer manager when settings.engine.useWriteBufferManager is true (the default), so the two share one memory budget instead of adding up.

  • Write Buffers (settings.engine.cf.write_buffer_size x settings.engine.cf.max_write_buffer_number): In-memory buffers that absorb write bursts before flushing to disk. The default is a 64MB buffer size with 24 buffers (~1.5GB total capacity).

resources.requests.memory should cover the row cache (cacheSize) plus the block cache (blockCacheSize) plus write buffers plus roughly 6GB of TCP buffers. The chart requests 56G by default.

3. Volume Configuration#

Reading from persistent storage, even a network-attached volume, is often faster than regenerating derived data. Therefore, OVDC persists content to a Kubernetes volume.

storage:
  volume:
    size: 330Gi
    storageClassName: "gp3"

storage.volume.size determines the size of the PersistentVolumeClaim created and attached to each pod. volumeClaimTemplates are immutable, so changing this value on an existing release is rejected on upgrade—plan the size up front.

storage.volume.storageClassName determines the performance characteristics of the persistent volume. Select a storage class that provides high IOPS and throughput suitable for database workloads (see the 800 MB/s guidance above).

settings.store.maxSize is the RocksDB on-disk target and the denominator garbage collection measures free capacity against. The default is "AUTO", which derives the target from storage.volume.size: the whole volume less min(20%, 100GB) of headroom.

Volume

Resulting target

10Gi

80.0%

300Gi

80.0%

1Ti

90.9%

10Ti

99.1%

AUTO requires storage.volume.enabled: true; the render fails otherwise. To set an explicit size instead, use a value such as "275G", which is used verbatim.

Example Kubernetes StorageClass for AWS#

If a StorageClass named gp3 does not already exist in your cluster, one can be created using the following configuration:

apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: gp3
provisioner: kubernetes.io/aws-ebs
volumeBindingMode: WaitForFirstConsumer
parameters:
  type: gp3
  iops: "5000"
  throughput: "1000"

Apply this StorageClass to your cluster:

kubectl apply -f storageclass-gp3.yaml

4. Telemetry#

OVDC exposes Prometheus metrics for monitoring cache performance, hit rates, and storage utilization, and can separately export traces, logs, and metrics through OpenTelemetry (OTLP over gRPC, the same remote-procedure-call protocol OVDC’s client API itself uses).

settings:
  telemetry:
    prometheusMetricsExposition: true
    prometheusMetricsPort: 3051

Metrics Configuration:

  • settings.telemetry.prometheusMetricsExposition: When true, an HTTP metrics service is started on prometheusMetricsPort, serving Prometheus plain-text metrics. When false, nothing listens on the port, but the Service and its scrape annotations are still created, so scrapes fail rather than being skipped.

  • prometheusAnnotation (top level): When true, adds prometheus.io/* scrape annotations to the metrics Service.

OTLP export is separate and requires a collector you provide—the chart does not deploy one. Configure it under settings.otelTraces, settings.otelLogs, and settings.otelMetrics, each with its own exporter ("otlp" or "none") and otlp.endpoint. This build only ships the gRPC OTLP exporter; http/protobuf and http/json deserialize but export nothing.

Note

There is no ServiceMonitor in this chart, and the standalone OTel-collector sidecar (otelCollector.*) was removed in 6.0.0. Run a collector separately and point the otel*.otlp.endpoint settings at it, or use the prometheus.io/* annotations with a Prometheus that supports annotation-based scraping.

Configuration Recommendations#

The following configuration provides a complete values file suitable for a small production deployment. Adjust replicas, resources.requests.memory, and storage.volume.size for your GPU count and workload per the sizing guidance above.

image:
  pullSecrets:

    - name: regcred

replicas: 3
selfAntiAffinity: true
affinity:
  nodeAffinity:
    preferredDuringSchedulingIgnoredDuringExecution:

      - weight: 100
        preference:
          matchExpressions:

            - key: node-type
              operator: In
              values:

                - compute

resources:
  requests:
    memory: 56G

storage:
  volume:
    size: 330Gi
    storageClassName: "gp3"

settings:
  logFormat: "json"
  logLevel: "info"
  store:
    maxSize: AUTO
  engine:
    cacheSize: "32G"
    blockCacheSize: "8G"
    increaseParallelism: 8
    cf:
      max_write_buffer_number: 24
  telemetry:
    prometheusMetricsExposition: true
    prometheusMetricsPort: 3051

Summary#

This guide covered the configuration options for OVDC, including scaling considerations, memory allocation, storage sizing, and monitoring setup. Proper configuration of these settings is essential for optimal OVDC performance in your self-hosted NVCF cluster.

Once you have prepared your values.yaml file with the appropriate configuration, proceed to the deployment guide to deploy OVDC using Helm.