OVDC: Metrics Reference#

All OVDC-emitted metrics are exposed as Prometheus plain-text on settings.telemetry.prometheusMetricsPort (default 3051), prefixed with the OpenTelemetry service.name resource attribute, which is ovdc. This list is not exhaustive—scrape /metrics directly for the authoritative, current set.

A few Kubernetes-native metrics (kubelet, kube-state-metrics, cAdvisor) are also referenced by the runbooks below; they are not emitted by OVDC itself and are included here for convenience.

Cache and Storage Metrics#

Metric

Type

Labels

Description

ovdc_query

counter

level (0, 1, or 2)

Every Get attempt, labeled by the cache tier it was served from: 0 is the in-process row cache, 1 is a RocksDB b lock-cache-only read, 2 is a full RocksDB read that may hit disk.

ovdc_miss

counter

level

A miss at that tier.

ovdc_answer

counter

level

A hit at that tier.

o vdc_get_bytes

counter

level

Bytes returned on a hit at that tier.

ovdc_put

counter

none

Put request count.

o vdc_put_bytes

counter

none

Bytes written using Put.

ovdc_b ytes_returned

counter

none

Total bytes returned across all Get responses (ap plication-level throughput).

ovdc_ro wcache_insert

counter

none

Entries added to the in-process row cache after a disk read.

` ovdc_io_time`

histogram (seconds)

io (`` rocksdb_read``, roc ksdb_read_l1, roc ksdb_read_l2, r ocksdb_write, ro cksdb_delete, or r ocksdb_merge)

RocksDB operation latency.

ovdc _gc_time_hist

histogram (seconds)

operation (gc)

Duration of a full gar bage-collection pass.

Error Metrics#

ovdc_error{kind="..."} counts errors by kind. Every value the binary currently emits:

kind value

Meaning

rocks_try_again

RocksDB TryAgain—transient, the caller should retry. Also reported to the client as Busy.

rocks_incomplete

RocksDB Incomplete—also increments ovdc_put_delay and is reported to the client as Busy.

rocks_corruption

RocksDB Corruption.

rocks_not_supported

RocksDB NotSupported.

rocks_invalid_argument

RocksDB InvalidArgument.

rocks_io_error

RocksDB IOError—disk I/O failure, including out-of-space conditions.

rocks_merge_in_progress

RocksDB MergeInProgress.

rocks_shutdown_in_progress

RocksDB ShutdownInProgress.

rocks_timed_out

RocksDB TimedOut.

rocks_aborted

RocksDB Aborted.

rocks_busy

RocksDB Busy.

rocks_expired

RocksDB Expired.

rocks_compaction_too_large

RocksDB CompactionTooLarge.

rocks_column_family_dropped

RocksDB ColumnFamilyDropped.

rocks_unknown

RocksDB Unknown.

monitor_collect_failed

The internal stats-collection task failed to read RocksDB properties for a cycle; that cycle’s ovdc_rocks_* gauges are stale.

NotFound (a normal cache miss) does not increment ovdc_error—see ovdc_miss above instead.

RocksDB Property Gauges#

Unlike the earlier assumption in this doc set, RocksDB’s internal properties are exported as Prometheus metrics—there is no need to read the RocksDB LOG file for these. A background task polls RocksDB’s native property interface every collection cycle and republishes each property as a gauge, labeled by name (the property ID) and cf (default, blobs, or - for a database-wide value that is not split by column family):

Metric

Type

Labels

Description

ovdc_rocks_in trinsic_gauge

gauge (u64/bool)

name, cf

Integer or boolean RocksDB properties.

ov dc_rocks_intrin sic_f64_gauge

gauge (float)

name, cf

Floating-point RocksDB properties.

ovdc_r ocks_capacity

gauge

name (bytes_free or p_free)

Free bytes and free-percentage against settings.s tore.maxSize.

ovdc_ rocks_interval_ bytes_per_sec

gauge

name (write or read)

Throughput sampled over the collection interval.

Common name values for ovdc_rocks_intrinsic_gauge / _f64_gauge (not exhaustive—every RocksDB property the monitor reads is republished this way):

name value

Meaning

__rocksdb_stalls_total_stop / __rocksdb_stalls_total_delays

Write stalls: hard stops and soft delays, aggregated across causes.

__rock sdb_stalls_stops_memtable_limit / __rocks db_stalls_delays_memtable_limit

Stalls from memtable count limits.

__rocksdb_stalls _stops_pending_compaction_bytes / __rocksdb_stalls_ delays_pending_compaction_bytes

Stalls from compaction backlog.

__rocksdb_stalls_w rite_buffer_manager_limit_stops

Stalls from the shared write-buffer-manager limit (see settings .engine.useWriteBufferManager).

__rocksdb_blockcache_capacity / __rocksdb_blockcache_usage

Block cache size and current usage.

__rocksdb_tables_sst_num / __rocksdb_tables_blobs_num

SST and blob file counts, per column family.

rocksdb.sstables

Per-level SST file listing, used to derive L0 file counts.

If a monitor collection cycle fails, ovdc_error{kind="monitor_collect_failed"} increments and that cycle’s gauge values are stale rather than absent.

API and gRPC Metrics#

Metric

Type

Labels

Description

ovdc_a pi_keys_total

counter

operation (grpc_get, grpc_put, and so on)

Keys processed, by RPC.

ovdc_api _errors_total

counter

none

Ap plication-level API errors (distinct from ovdc_error, which is RocksDB-layer).

ovdc_api_ request_total

counter

operation

Requests received, by RPC.

ovdc_ap i_connections

up/down counter

none

Current open connection count.

ovdc_ap i_uptime_secs

gauge

none

Process uptime in seconds.

ovdc_api _request_time

histogram (seconds)

none

End-to-end request duration.

ovdc_api_wri te_queue_time

histogram (seconds)

none

Time a write spent queued before processing.

ovdc_api_w rites_waiting

up/down counter

none

Writes currently queued.

ovdc_api _tcp_rx_total / ovdc_api _tcp_tx_total

counter (bytes)

none

Raw TCP bytes receiv ed/transmitted.

o vdc_grpc_server _call_started

counter

grpc_method

One increment per RPC start.

ov dc_grpc_server_ call_duration

histogram (seconds)

` grpc_method`, grpc_status

End-to-end RPC duration.

ovdc_grpc_ser ver_call_sent_t otal_compressed _message_size

histogram (bytes)

` grpc_method`, grpc_status

Outbound message sizes (for example, Get chunks).

ovdc_grpc_ser ver_call_rcvd_t otal_compressed _message_size

histogram (bytes)

` grpc_method`, grpc_status

Inbound message sizes.

Note

Label keys with dots in the Rust source (grpc.method, grpc.status) are sanitized to underscores in the Prometheus exposition format: grpc_method, grpc_status.

Kubernetes Metrics Referenced by the Runbooks#

Metric

Type

Labels

Source

Description

kubel et_volume_s tats_availa ble_bytes

gauge

pe rsistentvol umeclaim, `` namespace``

kubelet

Remaining disk space on the PVC.

kube let_volume_ stats_capac ity_bytes

gauge

pe rsistentvol umeclaim, `` namespace``

kubelet

Total PVC size.

`` kubelet_vol ume_stats_u sed_bytes``

gauge

pe rsistentvol umeclaim, `` namespace``

kubelet

Current disk consumption on the PVC.

contai ner_network _receive_by tes_total

counter

pod

cAdvisor

Inbound network bytes per pod.

contain er_network_ transmit_by tes_total

counter

pod

cAdvisor

Outbound network bytes per pod.

k ube_pod_sta tus_ready

gauge

pod, `` namespace``

kube-st ate-metrics

Whether a pod is ready to accept traffic.

k ube_pod_sta tus_phase

gauge

pod, n amespace, phase

kube-st ate-metrics

Pod lifecycle phase (` Pending`, ` Running`, Failed, and so on).

kube_pod_ container_s tatus_resta rts_total

counter

pod, n amespace, `` container``

kube-st ate-metrics

Container restart count.

RocksDB Terms Used in the Runbooks#

Term

Meaning

Memtable

The in-memory write buffer RocksDB fills before flushing to disk. Sized by setti ngs.engine.cf.write_buffer_size x max_write_buffer_number.

SST file

Sorted String Table—RocksDB’s immutable on-disk file format. Memtables flush to new SST files; compaction merges and rewrites them.

L0 (level 0)

The first on-disk level SST files land at after a memtable flush, before compaction merges them into higher levels. A large L0 file count is an early sign compaction is falling behind.

Compaction

The background process that merges and rewrites SST files across levels, reclaiming space from deleted/overwritten keys. Controlled by settings.engine .cf.periodic_compaction_seconds and related cf / block keys.

WAL (write-ahead log)

RocksDB’s durability log. OVDC’s default durability level (Low) skips the WAL for throughput, since cached values are always recomputable.

Write stall

RocksDB throttling or blocking writes because memtables, L0 file count, or pending compaction bytes exceed a threshold. Visible as ovdc_rocks_intrinsic_ga uge{name="__rocksdb_stalls_*"}.