OVDC: Helm Values#

values.yaml has two halves, and the distinction matters:

  • settings.* mirrors the application’s own config schema (rust/ovdc/src/config.toml in the OVDC repository) and is rendered key-by-key into the ConfigMap’s config.toml. Every key here must be one the binary deserializes; a missing required key panics the process at startup.

  • Everything else is Kubernetes-level chart configuration.

The container re-reads config.toml every 5 seconds and applies changes without a restart, but only for settings that are read per-use: logFormat / logLevel / logFilter, all otel* / telemetry sections (including otelTraces.level / filter and otelLogs.level / filter), garbageCollection, and store.blobCutoff / blobChunkSize. store.maxSize, store.mode, the whole engine section, grpc, grpcTls, and grpcJwt are read once at startup, so changing them requires a pod restart. A reloaded file that fails validation panics the pod, so a bad helm upgrade surfaces as CrashLoopBackOff.

Top-Level Field Reference#

  • ``image.repository`` — string, default nvidia/omniverse/ovderivedcache. Image repository under image.registry.

  • ``image.pullSecrets`` — list of {name}, default [{name: regcred}]. Kubernetes imagePullSecrets entries.

  • ``nameOverride``, ``fullnameOverride`` — string, default "". Override generated resource-name components. See the deployment guide for the naming pattern.

  • ``replicas`` — integer, default 1. Number of OVDC pods. Not hot-reloaded (pod restart).

  • ``podSecurityContext``, ``securityContext`` — Kubernetes securityContext object, default non-root uid/gid 1000 with all capabilities dropped. Pod- and container-level security settings.

  • ``resources`` — Kubernetes resources object, default requests.memory: 56G. Passed through verbatim to the container.

  • ``storage.volume.enabled`` — boolean, default true. When false, the DB is ephemeral (container filesystem)—dev/test only. Not hot-reloaded (immutable).

  • ``storage.volume.size``, ``storageClassName`` — quantity string / string, default 300Gi / managed-csi. Per-pod PersistentVolumeClaim sizing and class. Not hot-reloaded (volumeClaimTemplates are immutable).

  • ``settings.store.*`` — object, default maxSize: AUTO, blobCutoff: 8000, blobChunkSize: 4000000, mode: "log". RocksDB storage sizing and mode; see Configure. maxSize and mode are not hot-reloaded; blobCutoff and blobChunkSize are.

  • ``settings.engine.*`` — object, default cacheSize: "32G", blockCacheSize: "8G", plus RocksDB db, cf, and block option maps. RocksDB engine tuning: cacheSize, blockCacheSize, and the db, cf, block, blobsCf, blobsBlock option maps. Not hot-reloaded (pod restart).

  • ``settings.garbageCollection.*`` — object, default minFreeCapacity: 40, deleteKeyspaceQuantile: 60. Garbage collection thresholds; the capacity check itself runs on a fixed 5-second interval. Hot-reloaded.

  • ``settings.telemetry.*`` — object, default prometheusMetricsExposition: true, prometheusMetricsPort: 3051. Prometheus metrics exposition. Hot-reloaded.

  • ``settings.logFilter`` — string, default "". Optional per-module level overrides for the stdout sink, layered on top of logLevel (format: module[*]=level[,module2[*]=level2]..., for example "ovdc::grpc=debug"). Independent of otelTraces.level / filter and otelLogs.level / filter. Hot-reloaded.

  • ``settings.otel``, ``otelTraces``, ``otelLogs``, ``otelMetrics`` — object, default exporter: "otlp", endpoint: "http://localhost:4317". OpenTelemetry export configuration, one section per signal. otelTraces.level / otelLogs.level set the base level for that sink (enum, same values as logLevel); otelTraces.filter / otelLogs.filter are per-module overrides only (same format as logFilter), independent of logLevel and of each other. Hot-reloaded.

  • ``settings.grpc.*`` — object, default port: 3010 plus per-field defaults documented inline below. gRPC service tuning: timeouts, keepalives, message size limits. Not hot-reloaded (pod restart).

  • ``settings.grpcTls.*`` — object, default enabled: false. TLS for the gRPC listener; see TLS Configuration. Not hot-reloaded (pod restart).

  • ``settings.grpcJwt.*`` — object, default enabled: false. JWT verification for the gRPC listener. Not hot-reloaded (pod restart).

  • ``service.*`` — object, default loadBalancer: false. Service ports are not configurable here—they derive from settings.grpc.port and settings.telemetry.prometheusMetricsPort.

  • ``prometheusAnnotation`` — boolean, default true. Adds prometheus.io/* scrape annotations to the metrics Service.

  • ``debug`` — boolean, default false. When true, replaces the container command with sleep—development only, never enable in production.

The full field-level YAML, including every RocksDB engine.db / cf / block option and every grpc / grpcTls / grpcJwt / otel* key with inline descriptions, follows below.

# Labels added to the StatefulSet, its volume claims and the pod template.
global:
  additionalLabels:

image:
  # Registry and repository the release image is pulled from.
  registry: nvcr.io
  repository: nvidia/omniverse/ovderivedcache
  pullPolicy: IfNotPresent
  # Rendered as the pod's imagePullSecrets, so each entry is a `name:` mapping
  # referencing a kubernetes.io/dockerconfigjson Secret in this namespace.
  pullSecrets:

    - name: regcred
  # Image tag, used only when overrideTag is true. Otherwise the chart's
  # appVersion is used.
  tag: "latest"
  overrideTag: false

# Override the chart name used to build resource names. The release name still
# prefixes them: release `foo` with nameOverride `bar` yields `foo-bar`.
nameOverride: ""
# Replace the generated resource name outright, ignoring the release name.
fullnameOverride: ""

# The number of pods to deploy.
replicas: 1

# Node label constraints for the pods, e.g. `agentpool: ovdcpool`.
nodeSelector: {}

# Taints the pods tolerate, in the standard Kubernetes tolerations format.
tolerations: []

# When true, the scheduler prefers to place ovdc pods on nodes that do not
# already run one. Merged into `affinity` below; values set there win on
# conflict. Set false to drop the chart's rule entirely.
selfAntiAffinity: true

# Additional affinity rules for the pod template.
affinity: {}

# Annotations added to the pod template. Changing these rolls the StatefulSet.
podAnnotations: {}

# Pod-level securityContext, applied to every container in the pod. fsGroup
# makes the non-root default work by chowning the mounted volume at /env.
podSecurityContext:
  runAsNonRoot: true
  runAsUser: 1000
  runAsGroup: 1000
  fsGroup: 1000
  seccompProfile:
    type: RuntimeDefault

# Passed to the container as RUST_LOG and RUST_BACKTRACE. RUST_LOG is only a
# fallback for the stdout log level: settings.logLevel wins whenever it parses.
rustLog: "info"
rustBacktrace: "full"

# Container-level securityContext for the ovdc container.
securityContext:
  runAsNonRoot: true
  runAsUser: 1000
  runAsGroup: 1000
  allowPrivilegeEscalation: false
  capabilities:
    drop:

    - ALL

# Compute resources for the ovdc container, passed through verbatim. Should be
# at least engine.cacheSize + engine.blockCacheSize + write buffers + ~6GB TCP buffer.
resources:
  requests:
    memory: 56G
  #limits:
  #  memory: 248G

storage:
  volume:
    # When enabled the DB environment resides on a k8s mounted volume. When
    # false the DB environment is ephemeral (container filesystem)—dev/test only.
    enabled: true
    # Size of the requested PersistentVolumeClaim. Drives settings.store.maxSize
    # when that value is "AUTO". volumeClaimTemplates are immutable—changing
    # this on an existing release is rejected on upgrade.
    size: 300Gi
    # Azure-specific default; change to match a class available in your cluster.
    # Must support ReadWriteOnce; each replica gets its own volume.
    storageClassName: managed-csi

settings:
  # Allowed values: json, human, human_no_color
  logFormat: "json"
  # Log level for stdout. Allowed values: error, warn, info, debug, trace, off.
  logLevel: "info"
  # Optional per-module level overrides for the stdout sink, layered on top of
  # logLevel. Independent of otelTraces.level/filter and otelLogs.level/filter.
  # Format: module[*]=level[,module2[*]=level2]...  e.g. "ovdc::grpc=debug"
  logFilter: ""

  store:
    # Target max size for the application to use ("AUTO" derives it from
    # storage.volume.size: the whole volume less min(20%, 100GB) headroom).
    # Requires storage.volume.enabled: true, otherwise set an explicit size
    # such as "275G".
    maxSize: AUTO
    # Values larger than this are written as a chunked blob.
    blobCutoff: 8000
    # Size by which blob values above blobCutoff are split.
    blobChunkSize: 4000000
    # Storage engine mode. Allowed values: legacy, legacy_and_discrete, log.
    # Changing this on a populated database changes how values are read back.
    mode: "log"

  # RocksDB engine configuration (rust/ovdc_map/src/environment/rocks_adapter/config.rs).
  # Read once at startup; changes require a pod restart.
  engine:
    # When true, memtable memory is charged against the same LRU cache as
    # blockCacheSize.
    useWriteBufferManager: true
    # Size of the in-process row cache. This is OVDC's own cache, not RocksDB's. 0 disables it.
    cacheSize: "32G"
    # Size of the RocksDB LRU block cache.
    blockCacheSize: "8G"
    # RocksDB increase_parallelism: total background threads for flushes/compactions.
    increaseParallelism: 8
    # Env high/low priority background thread counts.
    threadsHigh: 4
    threadsLow: 4

    # Passed to RocksDB verbatim, merged over the binary's built-in defaults.
    # db: DBOptions (database-wide). cf/block: CFOptions/BlockBasedTableOptions
    # for the 'default' column family.
    db:
      allow_concurrent_memtable_write: true
      max_open_files: 128
      max_background_jobs: 6
      table_cache_numshardbits: 8
      fail_if_options_file_error: true
      keep_log_file_num: 8
      avoid_unnecessary_blocking_io: true
      max_log_file_size: "64M"
      stats_dump_period_sec: 0
      create_if_missing: true

    cf:
      min_blob_size: 8096
      enable_blob_files: true
      enable_blob_garbage_collection: true
      write_buffer_size: "64M"
      max_write_buffer_number: 24
      min_write_buffer_number_to_merge: 2
      memtable_whole_key_filtering: true
      level0_stop_writes_trigger: 24
      periodic_compaction_seconds: 28800
      memtable_prefix_bloom_size_ratio: 0.020000

    block:
      cache_index_and_filter_blocks: true
      block_size: 16384
      index_type: "kTwoLevelIndexSearch"
      filter_policy: "bloomfilter:10"
      partition_filters: true
      optimize_filters_for_memory: true
      data_block_index_type: "kDataBlockBinaryAndHash"
      pin_l0_filter_and_index_blocks_in_cache: true
      prepopulate_block_cache: "kFlushOnly"

    # Options for the 'blobs' column family (exists when store.mode is "log" or
    # "legacy_and_discrete"). Empty means RocksDB stock defaults, which differ
    # sharply from 'default' above.
    blobsCf: {}
    blobsBlock: {}

  # Garbage collection configuration. A background task compares DB usage
  # against settings.store.maxSize every 5 seconds.
  garbageCollection:
    # If free map space falls below this percentage, GC begins.
    minFreeCapacity: 40
    # Percentage of the keyspace GC attempts to remove once it begins.
    deleteKeyspaceQuantile: 60

  # HTTP metrics service serving Prometheus plain text metrics.
  telemetry:
    prometheusMetricsExposition: true
    prometheusMetricsPort: 3051

  # OpenTelemetry integration. otel.enabled/useEnvVars are required by the
  # schema but not independently actionable—disable export per-signal below.
  otel:
    enabled: true
    useEnvVars: true

  otelTraces:
    exporter: "otlp"
    otlp:
      protocol: "grpc"
      endpoint: "http://localhost:4317"
      timeout: "10s"
    processor:
      scheduleDelay: 5000
      maxQueueSize: 2048
      maxBatchSize: 512
    attributeValueLengthLimit: "4K"
    attributeCountLimit: 128
    eventCountLimit: 128
    linkCountLimit: 128
    eventAttributeCountLimit: 128
    linkAttributeCountLimit: 128
    samplerArg: 0.25
    sampler: "always_on"
    # Base level for events forwarded to the span exporter. Independent of
    # logLevel and otelLogs.level. Allowed values: error, warn, info, debug, trace, off.
    level: "info"
    # Per-module level overrides layered on top of `level`.
    # Format: module[*]=level[,module2[*]=level2]...
    filter: ""

  otelLogs:
    exporter: "otlp"
    otlp:
      protocol: "grpc"
      endpoint: "http://localhost:4317"
      timeout: "10s"
    processor:
      scheduleDelay: 5000
      maxQueueSize: 2048
      maxBatchSize: 512
    attributeValueLengthLimit: "4K"
    attributeCountLimit: 128
    # Base level for events forwarded to the log exporter. Independent of
    # logLevel and otelTraces.level. Allowed values: error, warn, info, debug, trace, off.
    level: "info"
    # Per-module level overrides layered on top of `level`.
    # Format: module[*]=level[,module2[*]=level2]...
    filter: ""

  otelMetrics:
    exporter: "otlp"
    otlp:
      protocol: "grpc"
      endpoint: "http://localhost:4317"
      timeout: "15s"
    processor:
      interval: "30s"
    filter: "info"

  # gRPC HTTP service. Read once at startup; changes require a pod restart.
  grpc:
    # Port for the gRPC service. Also used for the container port, both
    # Services' targetPort, and the readiness/liveness probes.
    port: 3010
    initialStreamWindowSize: "512K"
    initialConnectionWindowSize: "32M"
    connectionConcurrencyLimit: 32
    maxConcurrentStreams: 0
    maxFrameSize: 0
    timeout: "30s"
    writeTimeout: "15s"
    tcpKeepaliveAfterIdleTime: "3s"
    tcpKeepaliveInterval: "3s"
    tcpKeepaliveRetries: 5
    http2KeepaliveInterval: "3s"
    http2KeepaliveTimeout: "20s"
    # If true a tenant ID can be set through a string in the metadata (unauthenticated).
    tenantFromMetadata: true
    maxDecodingMessageSize: "5M"
    # Max time a single PUT stream can run for. Capped at 4h.
    putTimeout: "1h"

  # TLS for gRPC. Certificate/key paths are fixed by the chart at /etc/ovdc/tls
  # and are not configurable—only secretName and includeCaRoot are.
  grpcTls:
    enabled: false
    secretName: ""
    includeCaRoot: false

  # JWT verification for gRPC.
  grpcJwt:
    enabled: false
    require: false
    jwkPublicKeysetUrl: []
    jwkUpdateIntervalSecs: 60
    cacheSizeMb: 16
    verifyAud: false
    audClaim: ""
    verifyExp: false

# Service ports are not configurable here—they are derived from
# settings.grpc.port and settings.telemetry.prometheusMetricsPort.
service:
  annotations: {}
  # One LoadBalancer Service per pod (named <fullname>-<ordinal>), allowing an
  # external IP per replica. The headless Service is always created.
  loadBalancer: false
  loadBalancerSourceRanges: []
  loadBalancerIP: ""

# Adds prometheus.io scrape annotations to the metrics service.
prometheusAnnotation: true

# Values for the prometheus.io scrape annotations above.
monitoring:
  path: /metrics
  scheme: http

# When true the ovdc container's command is replaced with `sleep` (service does
# not run) and an extra <fullname>-debug pod is created with config.toml mounted
# at /config for inspection. Development only.
debug: false