+++ categories = ["software"] tags = ["automation","grafana","influxdb","telegraf"] date = 2025-05-16T08:00:00-05:00 description = "" draft = false slug = "server-monitor" title = "πŸ“ˆ Server Monitoring with Grafana" author = "nicholas" +++ In this post I will develop a server monitoring dashboard that will show me the status of my server at a glance. The goal here is to replicate the functionality of a simple command-line "real-time" system information monitor e.g. `top`, `htop`, etc. via a web interface. {{< image src="images/btop.png" caption="Btop UI" >}} - A **web-based** monitoring dashboard, as opposed to e.g. `htop`, will be useful to me because it will allow me to access this dashboard from virtually any OS, anywhere at any time. No more `ssh`-ing into my machine and running all manner of monitoring-related commands. - The web interface will also make integrating with other web-based dashboards trivial, which is something I plan to do later so that I can show my server stats alongside things like calendars, reminders, etc. in a master dashboard. - The tools I plan to use will allow me to configure alerts that will notify me if certain metrics go beyond threshold values, e.g. disk space, CPU utilization > 90%, etc. ## Project Scope There are a few pieces of information I want immediately visible when I visit my dashboard: - disk temperatures - fan speeds - CPU temperatures - CPU utilization - processes & threads This is the basic set of host-level metrics that I care about at the moment, and once I am able to monitor these I will consider the project done. I will be using tools that will make additional metrics easy to add later, e.g. network traffic, container-level metrics, etc. ## Tools Here is a diagram of the tools I will use and how they relate to each other. ```mermaid graph TD subgraph "πŸ–§ Host Machine" A[πŸ“Š Host metrics] -->|πŸ“₯ Collect| B[Telegraf] B[πŸ“¦ Telegraf] -->|πŸ“Š Write| C[InfluxDB] D[Grafana] -->|πŸ“¨ Query|C[πŸ“¦ InfluxDB] C[πŸ“¦ InfluxDB] -->|πŸ“ˆ Time-series data| D[Grafana] end D[πŸ“¦ Grafana] -->|🌐 Web Browser| E[πŸ‘€ User] ``` ## Stuck with InfluxDB v2 Unfortunately, I cannot run the latest version of Influxdb: `Influxdb:3`. My CPU does not support the required AVX instruction set, so the best I can do is `Influxdb:2`. I run into this same limitation each time I try to run a modern version (>=5) of MongoDB, which I use for other projects. Nevertheless, this older version will work fine, although I will be stuck writing my data queries in `Flux` instead of `SQL`. I really hate that. ## Configuration #### Directory Structure Since I will want to be able to easily edit a `telegraf.conf` configuration file from the host machine, I will bind-mount a `telegraf` subdirectory within the root `monitor` directory. This will contain my `telegraf.conf`. Also, I will create a `.env` file to store secret values to pass to my `docker-compose.yml` file. The directory structure will look like this: ``` πŸ“monitor/ β”œβ”€β”€ docker-compose.yml β”œβ”€β”€ .env β”œβ”€β”€ πŸ“telegraf/ └── telegraf.conf ``` #### Telegraf Telegraf is configured in `telegraf.conf`. I provide my configuration below. It contains many more input plugins than are being used, but I wanted to cover my bases. The two most important sections of this configuration file are the agent and the output sections. - **`[agent]`** - defines agent-level settings for the Telegraf daemon. - `hostname = "server"` sets the `host` tag for all metrics collected by Telegraf. - **`[[outputs.influxdb_v2]]`** - defines the output plugin, or where Telegraf sends the collected metrics. Targets InfluxDB v2. - `bucket = "telegraf"` defines the InfluxDB v2 bucket where metrics will be written. Relevant to the example query I use below is the `[[inputs.smart]]` plugin. This plugin uses `smartctl` (smartmontools) to query disk health and attributes from the devices listed in `devices = [ "/dev/sda","/dev/sdb","/dev/sdc","/dev/sdd","/dev/sde","/dev/sdf" ]`. Since I added the line `attributes = true`, I will have access to individual SMART attributes e.g. **temperature**, reallocated sectors, power on hours, etc. ```toml [agent] interval = "10s" round_interval = true hostname = "server" [[outputs.influxdb_v2]] urls = ["http://server:3997"] token = "" organization = "ward" bucket = "telegraf" [[inputs.cpu]] percpu = true totalcpu = true collect_cpu_time = false report_active = true [[inputs.diskio]] devices = ["sda", "sdb", "sdc", "sdd", "sde", "sdf"] [[inputs.procstat]] pattern = ".*" pid_finder = "native" [[inputs.smart]] path_smartctl = "/usr/sbin/smartctl" use_sudo = true devices = [ "/dev/sda","/dev/sdb","/dev/sdc","/dev/sdd","/dev/sde","/dev/sdf" ] attributes = true [[inputs.mem]] [[inputs.disk]] [[inputs.net]] [[inputs.netstat]] [[inputs.system]] [[inputs.sensors]] [[inputs.kernel]] [[inputs.swap]] [[inputs.interrupts]] [[inputs.linux_sysctl_fs]] [[inputs.processes]] [[inputs.zfs]] ``` #### Docker image I am not using the official Telegraf image, but instead a third-party customized variant [`golift/telegraf`](https://github.com/golift/telegraf-docker). This is because I needed to monitor HDD temperatures, and wanted to use `smartctl` (smartmontools) to do it. Telegraf can do this, but using it inside a container makes this difficult. The [repository README](https://github.com/golift/telegraf-docker/blob/main/README.md#why) explains the issue: > Provides a telegraf docker image with added tools for monitoring disks, sensors and IPMI. This exists because the base telegraf Docker image makes it difficult to monitor some system metrics. > > Re-creates the official Telegraf docker container with the following tools added: > - smartctl (smartmontools) > - ipmitool > - nvme-cli > - sensors (lm-sensors) > - mtr (mtr-tiny) > - sudo #### env ```env GRAFANA_ADMIN_USER=admin GRAFANA_ADMIN_PASSWORD= GRAFANA_PORT=3999 INFLUXDB_PORT=3997 TELEGRAF_PORT=3996 DOCKER_INFLUXDB_INIT_USERNAME=admin DOCKER_INFLUXDB_INIT_ADMIN_TOKEN= DOCKER_INFLUXDB_INIT_PASSWORD= DOCKER_INFLUXDB_INIT_ORG=ward ``` #### Docker Compose ```yml volumes: grafana_data: influxdb_data: networks: monitor_net: name: monitor_net driver: bridge services: influxdb: image: influxdb:2 container_name: influxdb ports: - "${INFLUXDB_PORT}:8086" restart: unless-stopped volumes: - influxdb_data:/var/lib/influxdb2 environment: - DOCKER_INFLUXDB_INIT_MODE=setup - DOCKER_INFLUXDB_INIT_USERNAME=${DOCKER_INFLUXDB_INIT_USERNAME} - DOCKER_INFLUXDB_INIT_PASSWORD=${DOCKER_INFLUXDB_INIT_PASSWORD} - DOCKER_INFLUXDB_INIT_ORG=${DOCKER_INFLUXDB_INIT_ORG} - DOCKER_INFLUXDB_INIT_BUCKET=telegraf - DOCKER_INFLUXDB_INIT_ADMIN_TOKEN=${DOCKER_INFLUXDB_INIT_ADMIN_TOKEN} networks: - monitor_net telegraf: image: golift/telegraf privileged: true container_name: telegraf restart: unless-stopped volumes: - ./telegraf:/etc/telegraf:ro - /:/hostfs:ro environment: - HOST_ETC=/hostfs/etc - HOST_PROC=/hostfs/proc - HOST_SYS=/hostfs/sys - HOST_VAR=/hostfs/var - HOST_RUN=/hostfs/run - HOST_MOUNT_PREFIX=/hostfs ports: - "${TELEGRAF_PORT}:9273" networks: - monitor_net grafana: image: grafana/grafana-oss container_name: grafana restart: unless-stopped environment: - GF_SECURITY_ADMIN_USER=${GRAFANA_ADMIN_USER} - GF_SECURITY_ADMIN_PASSWORD=${GRAFANA_ADMIN_PASSWORD} #- GF_SERVER_ROOT_URL=http://grafana.uuard.com ports: - "${GRAFANA_PORT}:3000" volumes: - 'grafana_data:/var/lib/grafana' healthcheck: test: [ "CMD", "wget", "--spider", "http://localhost:3000/login" ] interval: 30s timeout: 10s retries: 3 networks: - monitor_net ``` #### Grafana Config & Examples The first two steps of configuring data collection(via **Telegraf**) and writing (via **InfluxDB**) are done. Now I need to fetch, process, and display the data to my liking (via **Grafana**). This mostly involves writing ugly Flux queries and selecting a visualization to represent the data. ##### HDD Temperature This query gives me the temperature (Β°C) of each HDD in my machine: ```flux from(bucket: "${datasource}") |> range(start: -1m) |> filter(fn: (r) => r.host == "${host}" and r._measurement == "smart_device" and r._field == "temp_c") |> group(columns: ["device"]) |> last() |> yield(name: "temperature") ``` In this query, there are two variables used: `datasource` and `host`. These were both defined in `telegraf.conf`. ```toml [agent] ... hostname = "server" [[outputs.influxdb_v2]] ... bucket = "telegraf" ``` The query returns the latest (`last`) `temp_c` values per `device`, for a single `host`, from the last `60` seconds of SMART metrics. I selected the gauge visualization to represent this. ##### CPU Core Utilization This query gives me the CPU utilization per core (100% minus IDLE). ```flux from(bucket: "${datasource}") |> range(start: -1m) |> filter(fn: (r) => r.host == "${host}" and r._measurement == "cpu" and r._field == "usage_idle" and r.cpu != "cpu-total" ) |> group(columns: ["cpu"]) |> aggregateWindow(every: 1m, fn: mean, createEmpty: false) |> map(fn: (r) => ({ r with _value: 100.0 - r._value })) |> yield(name: "core_total_usage") ``` ## Result Here is what this looks like alongside the rest of the metrics I configured: {{< image src="images/grafana-live-dashboard.png" caption="Grafana Dashboard" >}} Each of the visualizations in this image are created with Flux queries similar to the example above. I don't bother writing each one out here. Instead, I attach a [JSON file](files/master.json) that can be imported into Grafana. It contains among many other details all of the queries used in the image. Done.