# Monitoring und Logging > **Version:** 2.0 > **Date:** 2026-08-20 ## Health Endpoints ### `/health/live` — Liveness Probe Prüft ob der Prozess lebt. Immer 200 wenn der API-Prozess läuft. ```bash curl https://crm.media-on.de/health/live # → {"status":"alive"} ``` Verwendung: Kubernetes/Coolify Restart-Entscheidung. ### `/health/ready` — Readiness Probe Prüft ob die App bereit ist Requests zu bedienen: - PostgreSQL Verbindung - Redis Verbindung - Storage Backend - Worker Heartbeat (falls verfügbar) ```bash curl https://crm.media-on.de/health/ready # → {"status":"ready","checks":{"database":"ok","redis":"ok","storage":"ok"}} ``` Verwendung: Load Balancer Traffic-Routing. Bei `not_ready` → 503. ### `/api/v1/health` — Full Health Vollständiger Health Check mit allen Details. Backward compatible. ```bash curl https://crm.media-on.de/api/v1/health # → {"status":"healthy","version":"1.0.0","checks":{...}} ``` --- ## System Dashboard (Admin-only) Das System Dashboard ist die zentrale Monitoring-Oberfläche für Administratoren. - **URL:** `/system-dashboard` (Frontend, Admin-only) - **API:** `GET /api/v1/system/dashboard` (Admin-only) - **Alerts:** `GET /api/v1/system/alerts` (Admin-only) ### Dashboard-Inhalte | Bereich | Metriken | |---------|----------| | **System Health** | Overall status (healthy/degraded/down) | | **Database** | Connections, Table Count, DB Size (bytes/human) | | **Redis** | Connected Clients, Used Memory, Peak Memory, Uptime | | **Worker** | Queue Length, Active Workers, Status (up/degraded/down) | | **API Stats** | Total Requests, Error Count, Error Rate, Avg Response Time | | **Plugins** | Total Discovered, Active Plugins (name, version, is_core) | | **Storage** | Disk Usage (total/used/free), File Count, Status | | **Alert Feed** | System Messages aus Communication-System | ### Alert-Dispatch Bei Problemen (DB down, Redis down, Worker down, High Error Rate) sendet das System Dashboard automatisch eine System-Message an das Communication-System (`post_system_message()`). Diese erscheint im Alert-Feed des Dashboards und in den Benachrichtigungen der Admins. ```bash # Dashboard abfragen curl -b "leocrm_session=" https://crm.media-on.de/api/v1/system/dashboard | jq . # Aktive Alerts abfragen curl -b "leocrm_session=" https://crm.media-on.de/api/v1/system/alerts | jq . ``` --- ## Metrics Endpoint ### `/api/v1/metrics` — Prometheus Metrics Prometheus-kompatible Metriken. Admin-only (403 für non-admin). ```bash curl -H "Authorization: Bearer ..." https://crm.media-on.de/api/v1/metrics ``` Metriken: - `leocrm_http_requests_total` — HTTP Request Counter - `leocrm_http_request_duration_seconds` — Request Duration Histogram - `leocrm_db_pool_size` — DB Connection Pool Size - `leocrm_db_pool_checked_out` — Active DB Connections - `leocrm_outbox_pending` — Pending Outbox Events - `leocrm_outbox_failed` — Failed Outbox Events --- ## Alerting via Notification-System LeoCRM nutzt das integrierte Notification-System für Alerting. Bei System-Problemen werden automatisch System-Messages an das Communication-System gesendet. ### Alert-Quellen | Quelle | Trigger | Ziel | |--------|---------|------| | System Dashboard | DB/Redis/Worker down, High Error Rate | Communication-System (Alert Feed) | | Backup Job | Backup fehlgeschlagen | Communication-System + Audit Log | | ARQ Worker | Job fehlgeschlagen, Queue überlastet | Communication-System | | Outbox | Delivery failed | Audit Log + Communication-System | ### Alert-Typen | Alert | Bedingung | Severity | |-------|-----------|----------| | API Down | `/health/live` nicht erreichbar | Critical | | API Not Ready | `/health/ready` = not_ready | Warning | | DB Down | Health check database = down | Critical | | Redis Down | Health check redis = down | Warning | | High Error Rate | Fehlerrate > 5% | Warning | | Slow Response | p95 > 2s | Warning | | DB Pool Exhausted | Pool checked_out = pool_size | Critical | | Outbox Backlog | Pending > 100 | Warning | | Worker Down | Worker heartbeat fehlt | Critical | | Backup Failed | Backup job returned error | Critical | | Queue Overload | Queue length > 500 | Warning | ### Alert-Empfang Alerts erscheinen: 1. Im System Dashboard Alert-Feed (`/system-dashboard`) 2. In den Benachrichtigungen der Admins (Notification-System) 3. Im Audit Log (für Backup/Worker/Outbox Alerts) --- ## Incident Response Runbook Siehe `docs/incident-response-runbook.md` für detaillierte Notfall-Prozeduren. ### Schnell-Referenz | Incident | Erste Maßnahme | escalation | |----------|---------------|------------| | **Server-Ausfall** | Coolify Restart → Health-Check → System-Message | Hetzner Support | | **DB-Crash** | PostgreSQL Restart → Migration-Check → Backup-Restore | Coolify DB Restart | | **Redis-Crash** | Redis Restart → Session-Check | Coolify Service Restart | | **Security-Breach** | Logs prüfen → Password-Reset → Audit-Log Export | Incident Response Team | | **Backup Failed** | Backup-Log prüfen → Manueller Retry → Storage prüfen | Admin | | **Worker Down** | Worker-Container Restart → Queue prüfen | Coolify Service Restart | ### Incident Response Schritte 1. **Erkennen** — Alert im System Dashboard oder externes Monitoring 2. **Eingrenzen** — Health-Checks, Logs, Metrics prüfen 3. **Beheben** — Restart, Restore, Konfiguration anpassen 4. **Verifizieren** — Health-Check grün, System Dashboard ok 5. **Dokumentieren** — Audit Log Eintrag, Post-Mortem bei Critical --- ## Externes Monitoring ### Empfohlene Tools - **Uptime Kuma** — Einfache Uptime-Überwachung - **Prometheus + Grafana** — Full Metrics Dashboard - **Sentry** — Error Tracking - **Coolify Health Monitoring** — Eingebaut in Coolify ### Coolify Health Check Konfiguration ```yaml health_check: type: http path: /health/ready port: 8000 interval: 30 timeout: 10 retries: 3 start_period: 15 ``` ### Uptime Kuma Setup 1. Monitor URL: `https://crm.media-on.de/health/live` 2. Expected Status: 200 3. Interval: 30s 4. Alert bei: 3 consecutive failures ### Prometheus Scrape Config ```yaml scrape_configs: - job_name: 'leocrm' metrics_path: '/api/v1/metrics' static_configs: - targets: ['crm.media-on.de'] authorization: type: Bearer credentials: '' ``` --- ## Strukturiertes Logging Alle API-Requests werden strukturiert geloggt: ```json { "method": "POST", "path": "/api/v1/contacts", "status": 200, "duration_ms": 15.3, "tenant_id": "bfe4d09e-...", "event": "api_request", "level": "info", "timestamp": "2026-08-20T14:00:00Z" } ``` Log-Level: - `info` — Normale API-Requests - `warning` — Langsame Requests, Permission denied - `error` — 500er Fehler, Exceptions ### Error Tracking ```python # app/core/monitoring.py record_error( error_code="DB_CONNECTION_FAILED", message="Database connection lost", context={"host": "postgres", "port": 5432}, trace_id="abc-123", ) ``` Fehler werden mit `trace_id` getrackt für Korrelation über Services hinweg.