Prometheus /metrics endpoint & observability #15
Labels
No labels
area/ai
area/backend
area/frontend
area/infra
area/scheduler
area/wled
good-first-issue
priority/high
priority/low
priority/medium
type/bug
type/chore
type/ci-cd
type/docs
type/feature
type/qa
v1.0.0
v1.1.0
v1.2.0
v2.0.0
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
rbrooks/Iris-WLED#15
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Goal
Expose operational metrics so Iris can be monitored (and alerted on) in a homelab Prometheus/Grafana stack.
Why it's valuable
Iris is a set-and-forget scheduler; when a nightly push silently fails, the user wants an alert, not to notice dark lights. v1 has structured JSON logging and a
schedule_log, but no scrape target.Sketch
prometheus-client, exposeGET /metrics(env-gated).Acceptance criteria
/metricsexposes counters/gauges for pushes, scheduler jobs, WLED connectivityProposed enhancement (not in the original spec).
Done — #124 merged, CI green.
/metricsexposes counters/gauges for pushes, scheduler jobs, WLED connectivityMETRICS_ENABLED)docs/setup.mdThe metric that matters
iris_scheduler_last_success_timestamp_seconds{job}. Iris is set-and-forget, so the failure worth paging on is the nightly push that never happened — and that is an absence, not an error. Nothing is logged for a job that did not run, so there is no line to alert on:Same reasoning made a down controller report
iris_wled_up 0rather than omitting the series. An absent series cannot be alerted on with== 0; it just stops matching, which is precisely the failure the metric exists to catch.Instrumented at chokepoints
Every scheduler job already funnels through
log_runand every WLED write through_try_push, so the counters increment there — one place each, rather than aninc()beside every job that would form a second accounting able to drift from the first.Given this milestone has produced three variants of "the test agreed with the code rather than reality", I wrote the drift check explicitly: it writes N schedule-log rows and asserts the counter moved by exactly N, exercising the real
log_run. Removing the instrumentation turns five tests red — checked, not assumed.State gauges are computed at scrape time from the controller status and
schedule_log, the same sources the API reports from, so a gauge cannot disagree with the UI. They are deliberately not derived counters:schedule_logis pruned afterSCHEDULE_LOG_RETAIN_DAYS, and a counter that goes down breaks everyrate()query built on it.Auth
Unauthenticated when enabled, as every exporter is — Prometheus cannot complete an OIDC flow, the same constraint #16 hit. The difference from a feed is that this exposes counts rather than data, and nothing is labelled with a host, path or credential, so a token would add friction without protecting anything. Documented as "keep it on the container network".
Safety
Recording never raises. The instrumentation sits inside
log_runand_try_push, where an exception would take down a scheduled job for the sake of a counter. The registry is private rather than the global default, which would otherwise pad the output with process and GC collectors that say nothing about whether the lights came on.