Files
leocrm/docs/monitoring.md
T

7.0 KiB

Monitoring und Logging

Version: 2.0
Date: 2026-08-20

Health Endpoints

/health/live — Liveness Probe

Prüft ob der Prozess lebt. Immer 200 wenn der API-Prozess läuft.

curl https://crm.media-on.de/health/live
# → {"status":"alive"}

Verwendung: Kubernetes/Coolify Restart-Entscheidung.

/health/ready — Readiness Probe

Prüft ob die App bereit ist Requests zu bedienen:

  • PostgreSQL Verbindung
  • Redis Verbindung
  • Storage Backend
  • Worker Heartbeat (falls verfügbar)
curl https://crm.media-on.de/health/ready
# → {"status":"ready","checks":{"database":"ok","redis":"ok","storage":"ok"}}

Verwendung: Load Balancer Traffic-Routing. Bei not_ready → 503.

/api/v1/health — Full Health

Vollständiger Health Check mit allen Details. Backward compatible.

curl https://crm.media-on.de/api/v1/health
# → {"status":"healthy","version":"1.0.0","checks":{...}}

System Dashboard (Admin-only)

Das System Dashboard ist die zentrale Monitoring-Oberfläche für Administratoren.

  • URL: /system-dashboard (Frontend, Admin-only)
  • API: GET /api/v1/system/dashboard (Admin-only)
  • Alerts: GET /api/v1/system/alerts (Admin-only)

Dashboard-Inhalte

Bereich Metriken
System Health Overall status (healthy/degraded/down)
Database Connections, Table Count, DB Size (bytes/human)
Redis Connected Clients, Used Memory, Peak Memory, Uptime
Worker Queue Length, Active Workers, Status (up/degraded/down)
API Stats Total Requests, Error Count, Error Rate, Avg Response Time
Plugins Total Discovered, Active Plugins (name, version, is_core)
Storage Disk Usage (total/used/free), File Count, Status
Alert Feed System Messages aus Communication-System

Alert-Dispatch

Bei Problemen (DB down, Redis down, Worker down, High Error Rate) sendet das System Dashboard automatisch eine System-Message an das Communication-System (post_system_message()). Diese erscheint im Alert-Feed des Dashboards und in den Benachrichtigungen der Admins.

# Dashboard abfragen
curl -b "leocrm_session=<session>" https://crm.media-on.de/api/v1/system/dashboard | jq .

# Aktive Alerts abfragen
curl -b "leocrm_session=<session>" https://crm.media-on.de/api/v1/system/alerts | jq .

Metrics Endpoint

/api/v1/metrics — Prometheus Metrics

Prometheus-kompatible Metriken. Admin-only (403 für non-admin).

curl -H "Authorization: Bearer ..." https://crm.media-on.de/api/v1/metrics

Metriken:

  • leocrm_http_requests_total — HTTP Request Counter
  • leocrm_http_request_duration_seconds — Request Duration Histogram
  • leocrm_db_pool_size — DB Connection Pool Size
  • leocrm_db_pool_checked_out — Active DB Connections
  • leocrm_outbox_pending — Pending Outbox Events
  • leocrm_outbox_failed — Failed Outbox Events

Alerting via Notification-System

LeoCRM nutzt das integrierte Notification-System für Alerting. Bei System-Problemen werden automatisch System-Messages an das Communication-System gesendet.

Alert-Quellen

Quelle Trigger Ziel
System Dashboard DB/Redis/Worker down, High Error Rate Communication-System (Alert Feed)
Backup Job Backup fehlgeschlagen Communication-System + Audit Log
ARQ Worker Job fehlgeschlagen, Queue überlastet Communication-System
Outbox Delivery failed Audit Log + Communication-System

Alert-Typen

Alert Bedingung Severity
API Down /health/live nicht erreichbar Critical
API Not Ready /health/ready = not_ready Warning
DB Down Health check database = down Critical
Redis Down Health check redis = down Warning
High Error Rate Fehlerrate > 5% Warning
Slow Response p95 > 2s Warning
DB Pool Exhausted Pool checked_out = pool_size Critical
Outbox Backlog Pending > 100 Warning
Worker Down Worker heartbeat fehlt Critical
Backup Failed Backup job returned error Critical
Queue Overload Queue length > 500 Warning

Alert-Empfang

Alerts erscheinen:

  1. Im System Dashboard Alert-Feed (/system-dashboard)
  2. In den Benachrichtigungen der Admins (Notification-System)
  3. Im Audit Log (für Backup/Worker/Outbox Alerts)

Incident Response Runbook

Siehe docs/incident-response-runbook.md für detaillierte Notfall-Prozeduren.

Schnell-Referenz

Incident Erste Maßnahme escalation
Server-Ausfall Coolify Restart → Health-Check → System-Message Hetzner Support
DB-Crash PostgreSQL Restart → Migration-Check → Backup-Restore Coolify DB Restart
Redis-Crash Redis Restart → Session-Check Coolify Service Restart
Security-Breach Logs prüfen → Password-Reset → Audit-Log Export Incident Response Team
Backup Failed Backup-Log prüfen → Manueller Retry → Storage prüfen Admin
Worker Down Worker-Container Restart → Queue prüfen Coolify Service Restart

Incident Response Schritte

  1. Erkennen — Alert im System Dashboard oder externes Monitoring
  2. Eingrenzen — Health-Checks, Logs, Metrics prüfen
  3. Beheben — Restart, Restore, Konfiguration anpassen
  4. Verifizieren — Health-Check grün, System Dashboard ok
  5. Dokumentieren — Audit Log Eintrag, Post-Mortem bei Critical

Externes Monitoring

Empfohlene Tools

  • Uptime Kuma — Einfache Uptime-Überwachung
  • Prometheus + Grafana — Full Metrics Dashboard
  • Sentry — Error Tracking
  • Coolify Health Monitoring — Eingebaut in Coolify

Coolify Health Check Konfiguration

health_check:
  type: http
  path: /health/ready
  port: 8000
  interval: 30
  timeout: 10
  retries: 3
  start_period: 15

Uptime Kuma Setup

  1. Monitor URL: https://crm.media-on.de/health/live
  2. Expected Status: 200
  3. Interval: 30s
  4. Alert bei: 3 consecutive failures

Prometheus Scrape Config

scrape_configs:
  - job_name: 'leocrm'
    metrics_path: '/api/v1/metrics'
    static_configs:
      - targets: ['crm.media-on.de']
    authorization:
      type: Bearer
      credentials: '<admin-token>'

Strukturiertes Logging

Alle API-Requests werden strukturiert geloggt:

{
  "method": "POST",
  "path": "/api/v1/contacts",
  "status": 200,
  "duration_ms": 15.3,
  "tenant_id": "bfe4d09e-...",
  "event": "api_request",
  "level": "info",
  "timestamp": "2026-08-20T14:00:00Z"
}

Log-Level:

  • info — Normale API-Requests
  • warning — Langsame Requests, Permission denied
  • error — 500er Fehler, Exceptions

Error Tracking

# app/core/monitoring.py
record_error(
    error_code="DB_CONNECTION_FAILED",
    message="Database connection lost",
    context={"host": "postgres", "port": 5432},
    trace_id="abc-123",
)

Fehler werden mit trace_id getrackt für Korrelation über Services hinweg.