Skip to main content
Lumovi finds the Prometheus or VictoriaMetrics in your cluster and charts what the cluster used, from the last 15 minutes to the last week. It asks through the API server with your own credentials, so nothing needs to be port-forwarded or exposed. History shows up in a few places:
  • The Metrics page. Its Usage tab ranks and compares, and its Right-sizing tab works out what each workload should request from a week of it. See Right-sizing.
  • A Metrics tab on pods, workloads and nodes.
  • The overview, whose CPU and memory cards show the last hour. See Overview.
Usage over the last six hours from Prometheus, by namespace.Usage over the last six hours from Prometheus, by namespace.

The Metrics page: six hours of CPU, by namespace.

What you need

  • A Prometheus-compatible server in the cluster. Prometheus (as kube-prometheus-stack, the Prometheus chart or kube-prometheus install it) or VictoriaMetrics, single-node or cluster.
  • cAdvisor’s metrics, scraped. CPU, memory and network charts use the kubelet’s container_* series. kube-prometheus-stack scrapes them out of the box.
  • kube-state-metrics, for restarts. It also places pods on nodes when cAdvisor’s series don’t carry a node label.
  • Permission to reach it. Your account needs get on services/proxy in the namespace the server runs in, and list on services across the cluster to find it.
Right-sizing needs a little more: 7 days of history (a day at the least), cAdvisor’s CPU throttling series (container_cpu_cfs_periods_total and container_cpu_cfs_throttled_periods_total), and kube-state-metrics to see OOM kills over the whole week.

How Lumovi finds it

You usually don’t have to do anything. Unless you’ve chosen a source, Lumovi looks through the cluster’s services for one that answers PromQL:
1

It scores every service

It skips services that come with Prometheus but don’t answer queries themselves (Alertmanager, operators, exporters, kube-state-metrics, Grafana, VictoriaMetrics’ agent and storage components, and so on), and rates the rest by their name and their app.kubernetes.io/name or app label:Services in a namespace named monitoring, prometheus, observability or victoria-metrics get 5 more. When none of a service’s ports has a name or number it expects, its first port is used. Services without ports are skipped.
2

It asks the best four

In order, it sends each a trivial query. The first that answers is used. Its name and version show on the chip at the top of every chart.
Lumovi looks once per session, and again when you save the metrics source or choose Look again where a chart would be. If yours runs under a name Lumovi doesn’t recognize, or you have several and want another, choose it yourself.

Choosing the source

Open Metrics source from the chip on any chart, or from the command palette (⌘K, then “Metrics source”). ⌘ is Ctrl on Windows and Linux.
The default. Lumovi looks for Prometheus and VictoriaMetrics among the cluster’s services, as described above.
The choice is kept per cluster.
In your cluster: the administrator sets the default for everyone with the chart’s metrics.source (auto, off, or a service like monitoring/prometheus-operated:9090). Each person can still choose another; it’s kept in their browser. See Helm values.

The Metrics page

Open it from the sidebar, or press G then U. It has two tabs: Usage, which ranks what uses the most and shows how that changed, and Right-sizing, which works out what each workload should request. This section is about Usage. For the other, see Right-sizing. Lumovi remembers the time range you pick, and the Metrics tabs start from it too. Above the chart, summary tiles: Now, Average and Peak for the total, and either Of allocatable, on average (for the whole cluster) or the busiest group. For restarts: how many there were, how many pods (or workloads…) restarted, which restarted most, and how many times. When the chart stacks up the whole cluster, with no namespace or filter picked, a line marks what the nodes can allocate. Below it:
  • A distribution. How the groups spread out. Click a band to filter the table to it.
  • Ranked by average (or Ranked by restarts): every group with its now, average and peak, and its share of the total, sortable. A workload has no peak, since its pods’ peaks don’t add up to one. It shows 50 at a time, with a button for more. Click a row to open it, on its Metrics tab when it has one.
The page follows the namespace menu, and keeps everything you pick in its address, so Back and Forward restore it.

The Metrics tab

Pods, Deployments, StatefulSets, DaemonSets, ReplicaSets, Jobs, CronJobs and Nodes have a Metrics tab in their detail panel. So do custom resources that run pods, like an Argo Rollout, or that their view relates to pods.
A pod's CPU and memory over time, against its requests and limits.A pod's CPU and memory over time, against its requests and limits.

A pod's Metrics tab: CPU and memory against its requests and limits.

A limit line appears only when every container has a limit, since without one there’s no ceiling to draw. Each chart’s header shows its total Now, Avg and Peak, where they add up. Show as table lists each series’ latest, average and peak value instead of the chart, and Show as chart goes back.

Reading the charts

  • Zoom in by dragging across a chart on the Usage tab or in a Metrics tab. The stretch you picked shows next to the time range, and stays put rather than refreshing. Reset zoom goes back to the range you picked. Right-sizing’s charts always show their week, and don’t zoom.
  • Read values by hovering, or with ← →, Home and End when the chart has focus. Esc hides the readout.
  • Show or hide a series by clicking it in the legend, when a chart has more than one. Alt-, ⌘- or Shift-click shows only that one, and Show all brings the others back.
  • Lines across a chart mark requests, limits or what’s allocatable, each with its label and value. When lines are close together, their labels go above and below them in turn, so they don’t overlap. A line far above the data, more than three times its peak, would flatten the chart, so it’s noted at the top instead, like ↑ Allocatable 27.53, above the chart.
Each range has its own resolution, and charts refresh on their own while you look:

When there’s no history

Where a chart would be, Lumovi says why, and what to do: Under Can’t read usage history, a message says what went wrong. When several services looked like Prometheus and none of them answered, it’s about the first. A chart that’s empty for a time usually means Prometheus has no samples for it: the object wasn’t running yet, or cAdvisor isn’t scraped. Restarts come from kube-state-metrics, so without it that chart stays empty.

Live usage

What’s in use right now, from metrics-server.

What Lumovi needs

The RBAC behind each feature, history included.