OVDC: Metrics Reference#
All OVDC-emitted metrics are exposed as Prometheus plain-text on
settings.telemetry.prometheusMetricsPort (default 3051),
prefixed with the OpenTelemetry service.name resource attribute,
which is ovdc. This list is not exhaustive—scrape /metrics
directly for the authoritative, current set.
A few Kubernetes-native metrics (kubelet, kube-state-metrics, cAdvisor) are also referenced by the runbooks below; they are not emitted by OVDC itself and are included here for convenience.
Cache and Storage Metrics#
Metric |
Type |
Labels |
Description |
|---|---|---|---|
|
counter |
|
Every |
|
counter |
|
A miss at that tier. |
|
counter |
|
A hit at that tier. |
|
counter |
|
Bytes returned on a hit at that tier. |
|
counter |
none |
Put request count. |
|
counter |
none |
Bytes written using Put. |
|
counter |
none |
Total bytes returned across all Get responses (ap plication-level throughput). |
|
counter |
none |
Entries added to the in-process row cache after a disk read. |
` ovdc_io_time` |
histogram (seconds) |
|
RocksDB operation latency. |
|
histogram (seconds) |
|
Duration of a full gar bage-collection pass. |
Error Metrics#
ovdc_error{kind="..."} counts errors by kind. Every value the binary
currently emits:
|
Meaning |
|---|---|
|
RocksDB |
|
RocksDB |
|
RocksDB |
|
RocksDB |
|
RocksDB |
|
RocksDB |
|
RocksDB |
|
RocksDB |
|
RocksDB |
|
RocksDB |
|
RocksDB |
|
RocksDB |
|
RocksDB |
|
RocksDB |
|
RocksDB |
|
The internal stats-collection
task failed to read RocksDB
properties for a cycle; that
cycle’s |
NotFound (a normal cache miss) does not increment ovdc_error—see
ovdc_miss above instead.
RocksDB Property Gauges#
Unlike the earlier assumption in this doc set, RocksDB’s internal
properties are exported as Prometheus metrics—there is no need to
read the RocksDB LOG file for these. A background task polls
RocksDB’s native property interface every collection cycle and
republishes each property as a gauge, labeled by name (the property
ID) and cf (default, blobs, or - for a database-wide
value that is not split by column family):
Metric |
Type |
Labels |
Description |
|---|---|---|---|
|
gauge (u64/bool) |
|
Integer or boolean RocksDB properties. |
|
gauge (float) |
|
Floating-point RocksDB properties. |
|
gauge |
|
Free bytes and
free-percentage
against
|
|
gauge |
|
Throughput sampled over the collection interval. |
Common name values for ovdc_rocks_intrinsic_gauge /
_f64_gauge (not exhaustive—every RocksDB property the monitor reads
is republished this way):
|
Meaning |
|---|---|
|
Write stalls: hard stops and soft delays, aggregated across causes. |
|
Stalls from memtable count limits. |
|
Stalls from compaction backlog. |
|
Stalls from the shared
write-buffer-manager limit (see
|
|
Block cache size and current usage. |
|
SST and blob file counts, per column family. |
|
Per-level SST file listing, used to derive L0 file counts. |
If a monitor collection cycle fails,
ovdc_error{kind="monitor_collect_failed"} increments and that
cycle’s gauge values are stale rather than absent.
API and gRPC Metrics#
Metric |
Type |
Labels |
Description |
|---|---|---|---|
|
counter |
|
Keys processed, by RPC. |
|
counter |
none |
Ap
plication-level
API errors
(distinct from
|
|
counter |
|
Requests received, by RPC. |
|
up/down counter |
none |
Current open connection count. |
|
gauge |
none |
Process uptime in seconds. |
|
histogram (seconds) |
none |
End-to-end request duration. |
|
histogram (seconds) |
none |
Time a write spent queued before processing. |
|
up/down counter |
none |
Writes currently queued. |
|
counter (bytes) |
none |
Raw TCP bytes receiv ed/transmitted. |
|
counter |
|
One increment per RPC start. |
|
histogram (seconds) |
`
grpc_method`,
|
End-to-end RPC duration. |
|
histogram (bytes) |
`
grpc_method`,
|
Outbound
message sizes
(for example,
|
|
histogram (bytes) |
`
grpc_method`,
|
Inbound message sizes. |
Note
Label keys with dots in the Rust source (grpc.method,
grpc.status) are sanitized to underscores in the Prometheus
exposition format: grpc_method, grpc_status.
Kubernetes Metrics Referenced by the Runbooks#
Metric |
Type |
Labels |
Source |
Description |
|---|---|---|---|---|
|
gauge |
|
kubelet |
Remaining disk space on the PVC. |
|
gauge |
|
kubelet |
Total PVC size. |
`` kubelet_vol ume_stats_u sed_bytes`` |
gauge |
|
kubelet |
Current disk consumption on the PVC. |
|
counter |
|
cAdvisor |
Inbound network bytes per pod. |
|
counter |
|
cAdvisor |
Outbound network bytes per pod. |
|
gauge |
|
kube-st ate-metrics |
Whether a pod is ready to accept traffic. |
|
gauge |
|
kube-st ate-metrics |
Pod
lifecycle
phase
(`
Pending`,
`
Running`,
|
|
counter |
|
kube-st ate-metrics |
Container restart count. |
RocksDB Terms Used in the Runbooks#
Term |
Meaning |
|---|---|
Memtable |
The in-memory write buffer
RocksDB fills before flushing to
disk. Sized by
|
SST file |
Sorted String Table—RocksDB’s immutable on-disk file format. Memtables flush to new SST files; compaction merges and rewrites them. |
L0 (level 0) |
The first on-disk level SST files land at after a memtable flush, before compaction merges them into higher levels. A large L0 file count is an early sign compaction is falling behind. |
Compaction |
The background process that
merges and rewrites SST files
across levels, reclaiming space
from deleted/overwritten keys.
Controlled by
|
WAL (write-ahead log) |
RocksDB’s durability log. OVDC’s
default durability level
( |
Write stall |
RocksDB throttling or blocking
writes because memtables, L0 file
count, or pending compaction
bytes exceed a threshold. Visible
as
|