Operational Troubleshooting Runbook
This operational runbook provides field diagnostic playbooks for common incidents encountered on the Planovi VPS host.
🚨 Incident Matrix & Quick Navigation
| Symptom | Probable Cause | Action Runbook |
|---|---|---|
| HTTP 502 Bad Gateway | Upstream container exited or restarting | Playbook 1: 502 Bad Gateway |
| Docker Disk Full (No space left on device) | Dangling build layers, old logs, or unpruned images | Playbook 2: Docker Disk Exhaustion |
| mTLS Telemetry Handshake Fails (400 Bad Request) | Client cert expired, untrusted CA, or Cloudflare Orange Cloud | Playbook 3: mTLS Telemetry Failures |
| GitHub Webhook Fails (403 Forbidden) | HMAC signature mismatch or secret typo | Playbook 4: Deployer Webhook Rejections |
| Supabase Studio / Docs Access Denied (403) | Client IP not on allowed_ips.conf whitelist | Playbook 5: Whitelist Access Lockout |
Playbook 1: HTTP 502 Bad Gateway
Diagnostic Steps
- Identify the affected domain (e.g.
api.planovi.app,dealflow.planovi.app,cloud.planovi.app). - Inspect Nginx error logs:
Terminal window docker logs --tail 100 vps-proxy | grep error - Check status of upstream container:
Terminal window docker compose ps
Resolution
- If
supabase-kongis down:Terminal window docker compose restart supabase-kong - If
php-appis down:Terminal window docker compose logs --tail 50 php-appdocker compose restart php-app - If
converters-api-*is down: Verify database connectivity:docker exec supabase-db pg_isready -U postgres.
Playbook 2: Docker Disk Exhaustion
Diagnostic Steps
df -h /docker system dfResolution
Safely reclaim unused build caches, stopped containers, and untagged images:
# Remove stopped containers, unused networks, and dangling imagesdocker system prune -f
# Clean multi-stage build cachedocker builder prune -a -f
# Truncate runaway container JSON logsfind /var/lib/docker/containers/ -name "*.log" -exec truncate -s 0 {} +
# Clean database backups older than 14 days if cron was disabledfind /opt/vps-stack/backups/ -name "*.sql.gz" -mtime +14 -deletePlaybook 3: mTLS Telemetry Failures
Diagnostic Steps
IoT converters cannot submit payloads to https://telemetry.planovi.app/functions/v1/ingest-telemetry.
- Verify Cloudflare DNS Mode:
Subdomains
telemetryanddev-telemetryMUST be set to DNS Only (Grey Cloud) in Cloudflare DNS. If set to Proxied (Orange Cloud), Cloudflare intercepts the TLS handshake and strips client certificates. - Verify Nginx Client CA Configuration:
Check
/etc/nginx/conf.d/supabase.conf:ssl_client_certificate /etc/ssl/cloudflare/planovi-device-ca.pem;ssl_verify_client on; - Test with Curl using Hardware Client Certificate:
Terminal window curl -v -k \--cert /opt/vps-stack/cloudflare-ssl/test-device.crt \--key /opt/vps-stack/cloudflare-ssl/test-device.key \https://telemetry.planovi.app/functions/v1/ingest-telemetry
Playbook 4: Deployer Webhook Rejections
Diagnostic Steps
GitHub reports delivery failure with HTTP 403.
- Inspect deployer logs:
Terminal window docker logs --tail 50 vps-deployer - Check if
Invalid HMAC signatureis logged.
Resolution
- Verify that
WEBHOOK_SECRETin/opt/vps-stack/.envexactly matches the secret entered in GitHub repository Settings ➔ Webhooks. - Restart the deployer container if
.envwas modified:Terminal window cd /opt/vps-stack && docker compose up -d vps-deployer
Playbook 5: Whitelist Access Lockout
If your ISP changes your IP and you cannot access Supabase Studio or Docs:
# SSH into the VPSssh root@191.218.165.149
# Run the update script to whitelist your new IP/opt/vps-stack/scripts/update-allowed-ip.sh add <YOUR_NEW_IP>Nginx is automatically tested and reloaded with zero downtime.