7.0 KiB
Monitoring und Logging
Version: 2.0
Date: 2026-08-20
Health Endpoints
/health/live — Liveness Probe
Prüft ob der Prozess lebt. Immer 200 wenn der API-Prozess läuft.
curl https://crm.media-on.de/health/live
# → {"status":"alive"}
Verwendung: Kubernetes/Coolify Restart-Entscheidung.
/health/ready — Readiness Probe
Prüft ob die App bereit ist Requests zu bedienen:
- PostgreSQL Verbindung
- Redis Verbindung
- Storage Backend
- Worker Heartbeat (falls verfügbar)
curl https://crm.media-on.de/health/ready
# → {"status":"ready","checks":{"database":"ok","redis":"ok","storage":"ok"}}
Verwendung: Load Balancer Traffic-Routing. Bei not_ready → 503.
/api/v1/health — Full Health
Vollständiger Health Check mit allen Details. Backward compatible.
curl https://crm.media-on.de/api/v1/health
# → {"status":"healthy","version":"1.0.0","checks":{...}}
System Dashboard (Admin-only)
Das System Dashboard ist die zentrale Monitoring-Oberfläche für Administratoren.
- URL:
/system-dashboard(Frontend, Admin-only) - API:
GET /api/v1/system/dashboard(Admin-only) - Alerts:
GET /api/v1/system/alerts(Admin-only)
Dashboard-Inhalte
| Bereich | Metriken |
|---|---|
| System Health | Overall status (healthy/degraded/down) |
| Database | Connections, Table Count, DB Size (bytes/human) |
| Redis | Connected Clients, Used Memory, Peak Memory, Uptime |
| Worker | Queue Length, Active Workers, Status (up/degraded/down) |
| API Stats | Total Requests, Error Count, Error Rate, Avg Response Time |
| Plugins | Total Discovered, Active Plugins (name, version, is_core) |
| Storage | Disk Usage (total/used/free), File Count, Status |
| Alert Feed | System Messages aus Communication-System |
Alert-Dispatch
Bei Problemen (DB down, Redis down, Worker down, High Error Rate) sendet das System Dashboard automatisch eine System-Message an das Communication-System (post_system_message()). Diese erscheint im Alert-Feed des Dashboards und in den Benachrichtigungen der Admins.
# Dashboard abfragen
curl -b "leocrm_session=<session>" https://crm.media-on.de/api/v1/system/dashboard | jq .
# Aktive Alerts abfragen
curl -b "leocrm_session=<session>" https://crm.media-on.de/api/v1/system/alerts | jq .
Metrics Endpoint
/api/v1/metrics — Prometheus Metrics
Prometheus-kompatible Metriken. Admin-only (403 für non-admin).
curl -H "Authorization: Bearer ..." https://crm.media-on.de/api/v1/metrics
Metriken:
leocrm_http_requests_total— HTTP Request Counterleocrm_http_request_duration_seconds— Request Duration Histogramleocrm_db_pool_size— DB Connection Pool Sizeleocrm_db_pool_checked_out— Active DB Connectionsleocrm_outbox_pending— Pending Outbox Eventsleocrm_outbox_failed— Failed Outbox Events
Alerting via Notification-System
LeoCRM nutzt das integrierte Notification-System für Alerting. Bei System-Problemen werden automatisch System-Messages an das Communication-System gesendet.
Alert-Quellen
| Quelle | Trigger | Ziel |
|---|---|---|
| System Dashboard | DB/Redis/Worker down, High Error Rate | Communication-System (Alert Feed) |
| Backup Job | Backup fehlgeschlagen | Communication-System + Audit Log |
| ARQ Worker | Job fehlgeschlagen, Queue überlastet | Communication-System |
| Outbox | Delivery failed | Audit Log + Communication-System |
Alert-Typen
| Alert | Bedingung | Severity |
|---|---|---|
| API Down | /health/live nicht erreichbar |
Critical |
| API Not Ready | /health/ready = not_ready |
Warning |
| DB Down | Health check database = down | Critical |
| Redis Down | Health check redis = down | Warning |
| High Error Rate | Fehlerrate > 5% | Warning |
| Slow Response | p95 > 2s | Warning |
| DB Pool Exhausted | Pool checked_out = pool_size | Critical |
| Outbox Backlog | Pending > 100 | Warning |
| Worker Down | Worker heartbeat fehlt | Critical |
| Backup Failed | Backup job returned error | Critical |
| Queue Overload | Queue length > 500 | Warning |
Alert-Empfang
Alerts erscheinen:
- Im System Dashboard Alert-Feed (
/system-dashboard) - In den Benachrichtigungen der Admins (Notification-System)
- Im Audit Log (für Backup/Worker/Outbox Alerts)
Incident Response Runbook
Siehe docs/incident-response-runbook.md für detaillierte Notfall-Prozeduren.
Schnell-Referenz
| Incident | Erste Maßnahme | escalation |
|---|---|---|
| Server-Ausfall | Coolify Restart → Health-Check → System-Message | Hetzner Support |
| DB-Crash | PostgreSQL Restart → Migration-Check → Backup-Restore | Coolify DB Restart |
| Redis-Crash | Redis Restart → Session-Check | Coolify Service Restart |
| Security-Breach | Logs prüfen → Password-Reset → Audit-Log Export | Incident Response Team |
| Backup Failed | Backup-Log prüfen → Manueller Retry → Storage prüfen | Admin |
| Worker Down | Worker-Container Restart → Queue prüfen | Coolify Service Restart |
Incident Response Schritte
- Erkennen — Alert im System Dashboard oder externes Monitoring
- Eingrenzen — Health-Checks, Logs, Metrics prüfen
- Beheben — Restart, Restore, Konfiguration anpassen
- Verifizieren — Health-Check grün, System Dashboard ok
- Dokumentieren — Audit Log Eintrag, Post-Mortem bei Critical
Externes Monitoring
Empfohlene Tools
- Uptime Kuma — Einfache Uptime-Überwachung
- Prometheus + Grafana — Full Metrics Dashboard
- Sentry — Error Tracking
- Coolify Health Monitoring — Eingebaut in Coolify
Coolify Health Check Konfiguration
health_check:
type: http
path: /health/ready
port: 8000
interval: 30
timeout: 10
retries: 3
start_period: 15
Uptime Kuma Setup
- Monitor URL:
https://crm.media-on.de/health/live - Expected Status: 200
- Interval: 30s
- Alert bei: 3 consecutive failures
Prometheus Scrape Config
scrape_configs:
- job_name: 'leocrm'
metrics_path: '/api/v1/metrics'
static_configs:
- targets: ['crm.media-on.de']
authorization:
type: Bearer
credentials: '<admin-token>'
Strukturiertes Logging
Alle API-Requests werden strukturiert geloggt:
{
"method": "POST",
"path": "/api/v1/contacts",
"status": 200,
"duration_ms": 15.3,
"tenant_id": "bfe4d09e-...",
"event": "api_request",
"level": "info",
"timestamp": "2026-08-20T14:00:00Z"
}
Log-Level:
info— Normale API-Requestswarning— Langsame Requests, Permission deniederror— 500er Fehler, Exceptions
Error Tracking
# app/core/monitoring.py
record_error(
error_code="DB_CONNECTION_FAILED",
message="Database connection lost",
context={"host": "postgres", "port": 5432},
trace_id="abc-123",
)
Fehler werden mit trace_id getrackt für Korrelation über Services hinweg.