2024-12-05 21:26:01 -05:00
# System Health Monitoring Daemon
2026-04-14 12:54:21 -04:00
[](https://code.lotusguild.org/LotusGuild/hwmonDaemon/actions?workflow=lint.yml)
[](https://code.lotusguild.org/LotusGuild/hwmonDaemon/actions?workflow=security.yml)
2024-12-05 21:26:01 -05:00
A robust system health monitoring daemon that tracks hardware status and automatically creates tickets for detected issues.
2026-10-03 01:06:06 -04:00

## ⚡ Quick run
Copy and paste on any Proxmox node (as root). No clone needed.
```bash
# Test it: prints the health summary and the tickets it WOULD create (nothing is sent)
python3 -c "import urllib.request; exec(urllib.request.urlopen('https://code.lotusguild.org/LotusGuild/hwmonDaemon/raw/branch/main/hwmonDaemon.py').read().decode('utf-8'))" --dry-run
# Run for real (creates tickets)
python3 -c "import urllib.request; exec(urllib.request.urlopen('https://code.lotusguild.org/LotusGuild/hwmonDaemon/raw/branch/main/hwmonDaemon.py').read().decode('utf-8'))"
# Install the systemd timer (hourly check)
curl -o /etc/systemd/system/hwmon.service https://code.lotusguild.org/LotusGuild/hwmonDaemon/raw/branch/main/hwmon.service && curl -o /etc/systemd/system/hwmon.timer https://code.lotusguild.org/LotusGuild/hwmonDaemon/raw/branch/main/hwmon.timer && systemctl daemon-reload && systemctl enable --now hwmon.timer
```
## Why this exists
A six-node Proxmox/Ceph cluster fails in small, boring ways: a drive starts reallocating sectors, a node runs hot, a pool fills up, Ceph reports slow operations. Nobody watches a dashboard all day, and a naive alert that fires every hour buries the one that matters. hwmonDaemon runs on every node, checks the hardware, and files **one ticket per real problem** in [Tinker Tickets ](https://code.lotusguild.org/LotusGuild/tinker_tickets ), with the diagnosis already written up.
```mermaid
flowchart LR
subgraph node["each Proxmox node (systemd timer)"]
smart["SMART + drive usage"] --> detect
sys["memory / CPU load / ECC"] --> detect
net["management + Ceph network"] --> detect
ceph["Ceph health + OSDs"] --> detect
lxc["LXC storage"] --> detect
detect{{"Detect + filter<br/>sustained load, standing flags"}}
end
detect -->|"dedup key: category + host + device"| api["Tinker Tickets API"]
api -->|"open ticket exists"| upd["update it"]
api -->|"closed, problem recurs"| reopen["reopen it"]
api -->|"new problem"| new["create ticket"]
```
Design choices worth calling out:
- **Dedup by hash** of category + hostname + device, so a repeating alert updates the existing ticket and a recurrence reopens a closed one.
- **Sustained, not spiky:** CPU tickets are gated on the 15-minute load average per core, not a transient spike.
- **Knows the cluster:** a standing Ceph `noout` flag (set on purpose for a node without a UPS) is not ticketed.
- **Dry-run first:** `--dry-run` prints exactly what would be filed, which is how the screenshots below were produced (on a live node, with nothing sent).
## Gallery
Output from a live node (`micro1` ) in dry-run mode:

An issue was detected (Ceph reporting slow BlueStore operations), so the daemon renders the ticket it would create, including an executive summary and the cluster status:

<sub>Rendered from real command output with `scripts/render_terminal.py` .</sub>
2024-12-05 21:26:01 -05:00
## Features
- Comprehensive system health monitoring:
- Drive health (SMART status and disk usage)
- Memory usage
- CPU utilization
- Network connectivity (Management and Ceph networks)
- Automatic ticket creation for detected issues
- Configurable thresholds and monitoring parameters
- Dry-run mode for testing
- Systemd integration for automated daily checks
2025-09-03 12:43:16 -04:00
- LXC container storage monitoring
- Historical trend analysis for predictive failure detection
- Manufacturer-specific SMART attribute interpretation
- ECC memory error detection
2024-12-05 21:26:01 -05:00
## Installation
1. Copy the service and timer files to systemd:
```bash
sudo cp hwmon.service /etc/systemd/system/
sudo cp hwmon.timer /etc/systemd/system/
```
2. Reload systemd daemon:
```bash
sudo systemctl daemon-reload
```
3. Enable and start the timer:
```bash
sudo systemctl enable hwmon.timer
sudo systemctl start hwmon.timer
```
2024-12-05 21:38:40 -05:00
### One liner (run as root)
```bash
2026-10-03 00:16:58 -04:00
curl -o /etc/systemd/system/hwmon.service https://code.lotusguild.org/LotusGuild/hwmonDaemon/raw/branch/main/hwmon.service && curl -o /etc/systemd/system/hwmon.timer https://code.lotusguild.org/LotusGuild/hwmonDaemon/raw/branch/main/hwmon.timer && systemctl daemon-reload && systemctl enable hwmon.timer && systemctl start hwmon.timer
2024-12-05 21:38:40 -05:00
```
2024-12-05 21:26:01 -05:00
## Manual Execution
2026-01-06 17:03:27 -05:00
### Direct Execution (from local file)
2024-12-05 21:26:01 -05:00
1. Run the daemon with dry-run mode to test:
```bash
python3 hwmonDaemon.py --dry-run
```
2. Run the daemon normally:
```bash
python3 hwmonDaemon.py
```
2026-01-06 17:03:27 -05:00
### Remote Execution (same as systemd service)
Execute directly from repository without downloading:
1. Run with dry-run mode to test:
```bash
2026-10-03 00:16:58 -04:00
/usr/bin/env python3 -c "import urllib.request; exec(urllib.request.urlopen('https://code.lotusguild.org/LotusGuild/hwmonDaemon/raw/branch/main/hwmonDaemon.py').read().decode('utf-8'))" --dry-run
2026-01-06 17:03:27 -05:00
```
2. Run normally (creates actual tickets):
```bash
2026-10-03 00:16:58 -04:00
/usr/bin/env python3 -c "import urllib.request; exec(urllib.request.urlopen('https://code.lotusguild.org/LotusGuild/hwmonDaemon/raw/branch/main/hwmonDaemon.py').read().decode('utf-8'))"
2026-01-06 17:03:27 -05:00
```
2024-12-05 21:26:01 -05:00
## Configuration
The daemon monitors:
- Disk usage (warns at 80%, critical at 90%)
2025-09-03 12:43:16 -04:00
- LXC storage usage (warns at 80%, critical at 90%)
2024-12-05 21:26:01 -05:00
- Memory usage (warns at 80%)
2025-09-03 12:43:16 -04:00
- CPU usage (warns at 95%)
2024-12-05 21:26:01 -05:00
- Network connectivity to management (10.10.10.1) and Ceph (10.10.90.1) networks
2025-09-03 12:43:16 -04:00
- SMART status of physical drives with manufacturer-specific profiles
- Temperature monitoring (warns at 65°C)
- Automatic duplicate ticket prevention
- Enhanced logging with debug capabilities
## Data Storage
The daemon creates and maintains:
- **Log Directory**: `/var/log/hwmonDaemon/`
- **Historical SMART Data**: JSON files for trend analysis
- **Data Retention**: 30 days of historical monitoring data
2026-01-06 16:57:16 -05:00
- **Storage Limit**: Automatically enforced 10MB maximum
- **Cleanup**: Oldest files deleted first when limit exceeded
2025-09-03 12:43:16 -04:00
2024-12-05 21:26:01 -05:00
## Ticket Creation
The daemon automatically creates tickets with:
- Standardized titles including hostname, hardware type, and scope
2025-09-03 12:43:16 -04:00
- Detailed descriptions of detected issues with drive specifications
2024-12-05 21:26:01 -05:00
- Priority levels based on severity (P2-P4)
- Proper categorization and status tracking
2025-09-03 12:43:16 -04:00
- Executive summaries and technical analysis
2024-12-05 21:26:01 -05:00
## Dependencies
- Python 3
- Required Python packages:
- psutil
- requests
2025-09-03 12:43:16 -04:00
- System tools:
2024-12-05 21:26:01 -05:00
- smartmontools (for SMART disk monitoring)
2025-09-03 12:43:16 -04:00
- nvme-cli (for NVMe drive monitoring)
## Excluded Paths
The following paths are automatically excluded from monitoring:
- `/media/*`
- `/mnt/pve/mediafs/*`
- `/opt/metube_downloads`
- Pattern-based exclusions for media and download directories
2024-12-05 21:26:01 -05:00
## Service Configuration
The daemon runs:
2026-01-06 16:57:16 -05:00
- Hourly via systemd timer (with 60-second randomized delay)
2024-12-05 21:26:01 -05:00
- As root user for hardware access
- With automatic restart on failure
2026-01-06 16:57:16 -05:00
- 5-minute timeout for execution
- Logs to systemd journal
## Recent Improvements
**Version 2.0** (January 2026):
- ✅ Added 10MB storage limit with automatic cleanup
- ✅ File locking to prevent race conditions
- ✅ Disabled monitoring for unreliable Ridata drives
- ✅ Added timeouts to all network/subprocess calls (10s API, 30s subprocess)
- ✅ Fixed unchecked regex patterns
- ✅ Improved error handling throughout
- ✅ Enhanced systemd service configuration with restart policies
2024-12-05 21:26:01 -05:00
2025-09-03 12:43:16 -04:00
## Troubleshooting
```bash
# View service logs
sudo journalctl -u hwmon.service -f
# Check service status
sudo systemctl status hwmon.timer
# Manual test run
python3 hwmonDaemon.py --dry-run
```
2024-12-05 21:26:01 -05:00
## Security Note
Ensure proper network security measures are in place as the service downloads and executes code from a specified URL.
2026-04-14 12:54:21 -04:00
## CI
| Workflow | Purpose | Triggers |
|---|---|---|
| `lint.yml` | flake8 on all `.py` files | Every push and PR |
| `security.yml` | bandit `-ll` (medium+ severity) | Every push, PR, and weekly Monday 6am |
Branch protection is enabled on `main` — the lint check must pass before any PR can merge.
Lint config: `.flake8` (max-line-length 120, F841/E501 ignored).