Files
leocrm/docs/monitoring.md
T

3.2 KiB

Monitoring und Logging

Health Endpoints

/health/live — Liveness Probe

Prüft ob der Prozess lebt. Immer 200 wenn der API-Prozess läuft.

curl https://crm.media-on.de/health/live
# → {"status":"alive"}

Verwendung: Kubernetes/Coolify Restart-Entscheidung.

/health/ready — Readiness Probe

Prüft ob die App bereit ist Requests zu bedienen:

  • PostgreSQL Verbindung
  • Redis Verbindung
  • Storage Backend
  • Worker Heartbeat (falls verfügbar)
curl https://crm.media-on.de/health/ready
# → {"status":"ready","checks":{"database":"ok","redis":"ok","storage":"ok"}}

Verwendung: Load Balancer Traffic-Routing. Bei not_ready → 503.

/api/v1/health — Full Health

Vollständiger Health Check mit allen Details. Backward compatible.

curl https://crm.media-on.de/api/v1/health
# → {"status":"healthy","version":"1.0.0","checks":{...}}

Metrics Endpoint

/api/v1/metrics — Prometheus Metrics

Prometheus-kompatible Metriken. Admin-only (403 für non-admin).

curl -H "Authorization: Bearer ..." https://crm.media-on.de/api/v1/metrics

Metriken:

  • leocrm_http_requests_total — HTTP Request Counter
  • leocrm_http_request_duration_seconds — Request Duration Histogram
  • leocrm_db_pool_size — DB Connection Pool Size
  • leocrm_db_pool_checked_out — Active DB Connections
  • leocrm_outbox_pending — Pending Outbox Events
  • leocrm_outbox_failed — Failed Outbox Events

Externes Monitoring

Empfohlene Tools

  • Uptime Kuma — Einfache Uptime-Überwachung
  • Prometheus + Grafana — Full Metrics Dashboard
  • Sentry — Error Tracking
  • Coolify Health Monitoring — Eingebaut in Coolify

Alerting Regeln

Alert Bedingung Severity
API Down /health/live nicht erreichbar Critical
API Not Ready /health/ready = not_ready Warning
DB Down Health check database = down Critical
Redis Down Health check redis = down Warning
High Error Rate Fehlerrate > 5% Warning
Slow Response p95 > 2s Warning
DB Pool Exhausted Pool checked_out = pool_size Critical
Outbox Backlog Pending > 100 Warning
Worker Down Worker heartbeat fehlt Critical

Coolify Health Check Konfiguration

health_check:
  type: http
  path: /health/ready
  port: 8000
  interval: 30
  timeout: 10
  retries: 3
  start_period: 15

Uptime Kuma Setup

  1. Monitor URL: https://crm.media-on.de/health/live
  2. Expected Status: 200
  3. Interval: 30s
  4. Alert bei: 3 consecutive failures

Prometheus Scrape Config

scrape_configs:
  - job_name: 'leocrm'
    metrics_path: '/api/v1/metrics'
    static_configs:
      - targets: ['crm.media-on.de']
    authorization:
      type: Bearer
      credentials: '<admin-token>'

Strukturiertes Logging

Alle API-Requests werden strukturiert geloggt:

{
  "method": "POST",
  "path": "/api/v1/contacts",
  "status": 200,
  "duration_ms": 15.3,
  "tenant_id": "bfe4d09e-...",
  "event": "api_request",
  "level": "info",
  "timestamp": "2026-07-29T16:00:00Z"
}

Log-Level:

  • info — Normale API-Requests
  • warning — Langsame Requests, Permission denied
  • error — 500er Fehler, Exceptions