Skip to content

Operational Troubleshooting Runbook

This operational runbook provides field diagnostic playbooks for common incidents encountered on the Planovi VPS host.


🚨 Incident Matrix & Quick Navigation

SymptomProbable CauseAction Runbook
HTTP 502 Bad GatewayUpstream container exited or restartingPlaybook 1: 502 Bad Gateway
Docker Disk Full (No space left on device)Dangling build layers, old logs, or unpruned imagesPlaybook 2: Docker Disk Exhaustion
mTLS Telemetry Handshake Fails (400 Bad Request)Client cert expired, untrusted CA, or Cloudflare Orange CloudPlaybook 3: mTLS Telemetry Failures
GitHub Webhook Fails (403 Forbidden)HMAC signature mismatch or secret typoPlaybook 4: Deployer Webhook Rejections
Supabase Studio / Docs Access Denied (403)Client IP not on allowed_ips.conf whitelistPlaybook 5: Whitelist Access Lockout

Playbook 1: HTTP 502 Bad Gateway

Diagnostic Steps

  1. Identify the affected domain (e.g. api.planovi.app, dealflow.planovi.app, cloud.planovi.app).
  2. Inspect Nginx error logs:
    Terminal window
    docker logs --tail 100 vps-proxy | grep error
  3. Check status of upstream container:
    Terminal window
    docker compose ps

Resolution

  • If supabase-kong is down:
    Terminal window
    docker compose restart supabase-kong
  • If php-app is down:
    Terminal window
    docker compose logs --tail 50 php-app
    docker compose restart php-app
  • If converters-api-* is down: Verify database connectivity: docker exec supabase-db pg_isready -U postgres.

Playbook 2: Docker Disk Exhaustion

Diagnostic Steps

Terminal window
df -h /
docker system df

Resolution

Safely reclaim unused build caches, stopped containers, and untagged images:

Terminal window
# Remove stopped containers, unused networks, and dangling images
docker system prune -f
# Clean multi-stage build cache
docker builder prune -a -f
# Truncate runaway container JSON logs
find /var/lib/docker/containers/ -name "*.log" -exec truncate -s 0 {} +
# Clean database backups older than 14 days if cron was disabled
find /opt/vps-stack/backups/ -name "*.sql.gz" -mtime +14 -delete

Playbook 3: mTLS Telemetry Failures

Diagnostic Steps

IoT converters cannot submit payloads to https://telemetry.planovi.app/functions/v1/ingest-telemetry.

  1. Verify Cloudflare DNS Mode: Subdomains telemetry and dev-telemetry MUST be set to DNS Only (Grey Cloud) in Cloudflare DNS. If set to Proxied (Orange Cloud), Cloudflare intercepts the TLS handshake and strips client certificates.
  2. Verify Nginx Client CA Configuration: Check /etc/nginx/conf.d/supabase.conf:
    ssl_client_certificate /etc/ssl/cloudflare/planovi-device-ca.pem;
    ssl_verify_client on;
  3. Test with Curl using Hardware Client Certificate:
    Terminal window
    curl -v -k \
    --cert /opt/vps-stack/cloudflare-ssl/test-device.crt \
    --key /opt/vps-stack/cloudflare-ssl/test-device.key \
    https://telemetry.planovi.app/functions/v1/ingest-telemetry

Playbook 4: Deployer Webhook Rejections

Diagnostic Steps

GitHub reports delivery failure with HTTP 403.

  1. Inspect deployer logs:
    Terminal window
    docker logs --tail 50 vps-deployer
  2. Check if Invalid HMAC signature is logged.

Resolution

  • Verify that WEBHOOK_SECRET in /opt/vps-stack/.env exactly matches the secret entered in GitHub repository Settings ➔ Webhooks.
  • Restart the deployer container if .env was modified:
    Terminal window
    cd /opt/vps-stack && docker compose up -d vps-deployer

Playbook 5: Whitelist Access Lockout

If your ISP changes your IP and you cannot access Supabase Studio or Docs:

Terminal window
# SSH into the VPS
ssh root@191.218.165.149
# Run the update script to whitelist your new IP
/opt/vps-stack/scripts/update-allowed-ip.sh add <YOUR_NEW_IP>

Nginx is automatically tested and reloaded with zero downtime.