Files
crm-system/docs/runbook-restore.md
T

14 KiB

CRM System — Restore & Incident Runbook

Scope: Single source of truth for diagnosing, restoring and maintaining the production CRM System at https://crm.media-on.de:443 (Coolify, FastAPI + PostgreSQL).

Audience: On-call DevOps / backend engineer with Coolify UI access, the server's SSH key, and read-only access to the Forgejo repository.


0. Architecture in 30 seconds

Layer Tech Where it lives
Frontend 13 static HTML pages (Alpine + Tailwind) Built into the image at app/webui/, served by FastAPI's StaticFiles mount
Backend FastAPI 0.115 (async) on uvicorn 1 worker Docker container on Coolify, internal port 8000
Database PostgreSQL 16 Coolify-managed crm-postgres resource (internal DNS)
Auth JWT (HS256, 24h) via python-jose, bcrypt-hashed passwords app/services/auth.py
Reverse proxy Traefik (managed by Coolify) Terminates TLS on :443 with Let's Encrypt
Source Git repo on Forgejo, branch master https://forge.media-on.de/leopoldadmin/crm-system

Domain gotcha (read once, remember forever): The Coolify Domain field must contain the port, e.g. https://crm.media-on.de:443. Without it, Let's Encrypt issuance silently fails and Traefik returns 404. See COOLIFY_SETUP.md § 0.


1. Health-endpoint checks

Run these from anywhere with internet access. The expected status is 200 and the response is JSON.

CRM=https://crm.media-on.de:443

# Root-level health (used by Docker HEALTHCHECK in the Dockerfile)
curl -fsS -w "\nHTTP %{http_code}  (%{time_total}s)\n" "$CRM/health"

# Versioned health (mounted under the /api/v1 router)
curl -fsS -w "\nHTTP %{http_code}  (%{time_total}s)\n" "$CRM/api/v1/health"

# Frontend SPA entry — must return text/html, NOT 404
curl -fsSI "$CRM/index.html" | head -1
curl -fsS "$CRM/index.html" | head -5   # should contain <!doctype html>

# A protected endpoint (should 401 without a token, 200 with one)
curl -sS -o /dev/null -w "unauthed: %{http_code}\n" "$CRM/api/v1/contacts"

A 200 from /health and /api/v1/health means:

  • The container is running and accepting connections.
  • The app can talk to PostgreSQL (the lifespan startup hook runs SELECT 1).

A 200 from /index.html with Content-Type: text/html means:

  • The app/webui/ directory was correctly baked into the image.
  • The static-files mount is active.

A 200 from /api/v1/contacts with a valid Bearer token means:

  • JWT verification works.
  • User has the crm:read permission (or equivalent role).

2. Database migration strategy

Alembic runs automatically on every container start (prestart.sh calls alembic upgrade head before exec'ing uvicorn). To manage migrations manually:

# Connect to the running container (Coolify UI → crm-app → Exec)
# or via Docker on the host if you have SSH access:
docker exec -it <crm-app-container> sh

# Inside the container:
alembic current       # show applied revision
alembic history --verbose | head -20   # show recent migrations
alembic upgrade head   # apply pending (same as prestart)
alembic downgrade -1   # roll back ONE revision (destructive!)

When a migration is risky

  1. Take a database backup first (see § 5).
  2. Deploy the new image in a staging environment if available, run alembic upgrade head there, and smoke-test.
  3. For production: push the new code, redeploy in Coolify. The container will fail to start if the migration is broken — the previous image stays on the old revision (because the upgrade runs in the new container only).

Idempotency rule: every migration must be written to be safely re-runnable for the cases where Coolify restarts the container mid-deploy (set -e in prestart.sh makes a half-applied migration fail loudly).


3. Rollback plan

There are two failure modes: bad code and bad migration.

3.1 Bad code (most common)

Roll back to the previous working commit:

# Locally
cd /path/to/crm-system
git log --oneline -5            # find the last good commit hash, e.g. a1b2c3d
git revert HEAD                  # produce a new commit that undoes HEAD
# or, if you want to force-push the old commit (destructive, only on master if alone):
# git reset --hard a1b2c3d
git push origin master

# In Coolify: crm-app → Deployments → Deploy
# The new build will run alembic upgrade head, but if the schema is unchanged
# it is a no-op.

If the new build itself fails (Dockerfile error etc.): in the Coolify UI you can pick an older image tag under crm-app → Deployments → Deploy → Tag and deploy that. This does not touch git history.

3.2 Bad migration

If a migration corrupted data or ran too long:

  1. Restore from backup (see § 5) into a new Postgres resource.
  2. Update the DATABASE_URL in Coolify to point at the restored DB.
  3. Redeploy the last known good image.
  4. Once stable, fix the migration locally, add a compensating migration, and re-test in staging before re-deploying.

Never edit a migration that has already been deployed to production. Add a new migration that moves the schema forward.


4. Log inspection

Coolify does not expose container logs via the API — you must use the UI or SSH into the server.

4.1 Coolify UI

crm-app → Logs (live tail, last ~5 MB). Best for quick triage.

4.2 Docker on the host

# SSH to the Coolify server (or use the Coolify terminal if enabled)
ssh root@server.media-on.de
docker ps | grep crm-app
docker logs --tail 200 --timestamps crm-app-abc123
docker logs -f crm-app-abc123   # follow live

4.3 Postgres logs

docker logs --tail 200 crm-postgres-xyz789

4.4 What to look for

Symptom in logs Likely cause Fix
alembic.util.exc.CommandError: ... Migration script broken Roll back via § 3.2
asyncpg.exceptions.InvalidPasswordError DATABASE_URL password wrong Update ENV in Coolify, redeploy
asyncpg.exceptions.CannotConnectNowError Postgres still starting up Wait — depends_on: service_healthy should prevent this in compose, but in Coolify the app may start before the DB is reachable. See § 7.
Missing required configuration: AUTH_SECRET AUTH_SECRET ENV empty or < 32 chars Set it in Coolify ENV (see § 6)
pydantic.ValidationError: AUTH_SECRET ... String should have at least 32 characters Same as above Same as above
uvicorn ... ERROR: [Errno 98] Address already in use Port collision — should not happen in a single container Restart container; if persistent, check docker ps for zombies
Repetitive 401s after a deploy AUTH_SECRET was rotated; old tokens invalid Expected — users must log in again (see § 6)

5. Backup strategy

5.1 PostgreSQL backups

Use Coolify's built-in database backup feature for the crm-postgres resource:

  1. Coolify UI → Databases → crm-postgres → Backups → + New.
  2. Configure a daily schedule, e.g. 0 3 * * * (03:00 UTC every day).
  3. Set retention to at least 7 days.
  4. Coolify will run pg_dump and store the file on the host (or your S3/MinIO if configured).

Manual one-off backup:

# From the Coolify server (SSH or terminal)
docker exec crm-postgres-xyz789 pg_dump -U crm_user -d crm_db -Fc -f /tmp/crm.dump
docker cp crm-postgres-xyz789:/tmp/crm.dump ./crm-$(date -u +%Y%m%dT%H%M%SZ).dump

Test the restore quarterly. A backup you never restored from is a backup you don't have. See the Coolify UI's Backups → Restore button.

5.2 Secrets backup

The following secrets must be backed up outside the server (e.g. in a password manager or KMS):

  • AUTH_SECRET
  • POSTGRES_PASSWORD
  • Forgejo deploy credentials (token used to push the repo)

These are not in git. If you lose them you must regenerate them and accept the consequences in § 6 (AUTH_SECRET) or § 3.2 (DB password).

5.3 Pre-upgrade backup ritual

Before any non-trivial code deploy (e.g. a new Alembic migration):

  1. Manual DB backup in addition to the daily schedule (so you have a point-in-time snapshot labelled with the pre-deploy state).
  2. Note the current alembic current revision in the runbook / commit message.
  3. Note the deployed image tag (Coolify → crm-app → Deployments).

6. Secret rotation

6.1 Rotate AUTH_SECRET

Effect: every existing JWT token becomes invalid. All users are silently logged out and must log in again. This is by design.

# 1. Generate a new secret (do NOT use a script output from an old terminal session)
python -c "import secrets; print(secrets.token_urlsafe(48))"

# 2. Update in Coolify: crm-app → Environment Variables → AUTH_SECRET → Save

# 3. Redeploy: crm-app → Deployments → Deploy
#    (Coolify does NOT auto-restart on ENV change.)

There is no global "invalidate all JWTs" button. The secret change is the invalidation — every previously-signed token's signature will no longer verify.

6.2 Rotate POSTGRES_PASSWORD

  1. Update the password on the Postgres resource (Coolify UI → DB → Reset Password, or run ALTER USER crm_user PASSWORD '...' inside the DB).
  2. Update DATABASE_URL in crm-app's ENVs with the new password.
  3. Redeploy crm-app.

6.3 Token / user revocation without rotating AUTH_SECRET

For revoking a single compromised account, change that user's password in the DB (forces logout) and consider adding a per-user token_version column to JWTs (future enhancement; not in v1).


7. Common issues & fixes

7.1 401 Unauthorized after AUTH_SECRET change

Cause: expected. Old JWTs are signed with the old secret and the new secret can't verify them.

Fix:

  • Communicate to users that they must log in again.
  • Optionally, set a longer JWT_EXPIRY_HOURS to reduce how often this happens in normal operation (current default: 24h).

7.2 Container fails to start — DB migration error

Symptoms:

  • Container restarts in a loop in the Coolify UI.
  • Logs show alembic.util.exc.CommandError or sqlalchemy.exc.ProgrammingError.

Fix:

  1. Open the container Exec (Coolify UI → crm-app → Exec).
  2. Run alembic current to see the applied revision.
  3. Run alembic history --verbose | head to see the chain.
  4. If the broken revision was just applied: alembic downgrade -1 to step back. If the new revision hasn't fully run, you may need to restore the DB from the pre-upgrade backup (§ 5.3) instead.
  5. Fix the migration locally, push a new commit, redeploy.

7.3 502 Bad Gateway from the domain

Causes (in order of likelihood):

  1. Domain field has no port. Fix: edit the Domain in Coolify to https://crm.media-on.de:443 (see § 0).
  2. App container is starting or unhealthy. Wait 30s, retry.
  3. App container exited (DB password wrong, missing ENV). Check Logs (§ 4).
  4. Traefik is restarting. Wait 30s, retry.
  5. DNS A record for crm.media-on.de does not point to the Coolify server's public IP. Check with dig +short crm.media-on.de.

7.4 CSP header too strict in production

Cause: the security-headers middleware sets Content-Security-Policy: default-src 'self'; script-src 'self' 'unsafe-inline'; ... to keep unsafe-inline out of production. The 13 frontend pages currently use inline <script> Alpine components, which need unsafe-inline to work.

Fix for v1: keep script-src 'self' 'unsafe-inline' in production. The CSP still blocks external scripts and inline event handlers (onclick=...).

Fix for v2: refactor the 13 pages to load Alpine and inline scripts from external files, then drop unsafe-inline and use a per-request nonce.

7.5 cannot import name 'app' from 'app' during container start

Cause: the prestart.sh exec uvicorn line is correct (app.main:app), but if you see this error, the working directory inside the container is wrong (should be /app).

Fix: the Dockerfile sets WORKDIR /app and the COPY --chown=... . . copies the repo contents (including app/, alembic/, alembic.ini, prestart.sh) to that path. Verify with:

docker exec <crm-app-container> ls /app
# expected: alembic  alembic.ini  app  docker-compose.yml  Dockerfile  ...

7.6 git clone in Coolify fails with Permission denied (publickey)

Cause: the Forgejo repo URL uses a token / SSH key that Coolify doesn't have, or the deploy key was revoked.

Fix:

  • For HTTPS repos: ensure the URL embeds a valid token (https://user:TOKEN@forge.media-on.de/.../crm-system.git).
  • For SSH repos: re-add the public key in the repo's Deploy keys in Forgejo and re-paste the matching private key in Coolify → Private keys.

7.7 docker compose locally fails with bind: address already in use

Cause: another process on the host is already listening on :5432 (local Postgres) or :8000 (the app).

Fix: either stop the conflicting process, or change the host port mapping in docker-compose.yml (e.g. "5433:5432", "8001:8000").


8. Escalation

If a service-impacting incident lasts longer than 30 minutes and you cannot restore it with this runbook:

  1. Roll back to the last known good image (Coolify UI → crm-app → Deployments → Deploy → older tag).
  2. Post a status note to the team's incident channel.
  3. Hand off to the next on-call engineer with: incident start time, last known good commit, what you tried, what you observed.

9. References