Blog

SSH, Linux, and production operations — page 22

Practical runbooks for servers, containers, databases, cloud hosting, incident response, and safe AI-assisted operations.

Dokku gives you a tiny Heroku; you still own production

Dokku can deploy open-source software quickly, but its runtime and storage model create tradeoffs. Here is where the friction appears and when a VPS is simpler.

The Tryssh team · July 21, 2026

Encrypted backups are useless without recoverable keys

Encrypted backups are useless without recoverable keys. A practical production guide with diagnostic commands, failure interpretation, and a safe decision…

The Tryssh team · July 21, 2026

Ephemeral port exhaustion: the hidden client-side outage

Ephemeral port exhaustion: the hidden client-side outage. A practical production guide with diagnostic commands, failure interpretation, and a safe…

The Tryssh team · July 21, 2026

Fail2ban for SSH: useful guardrail, not your primary lock

Fail2ban for SSH: useful guardrail, not your primary lock. Audit the exposure, understand the threat model, apply changes safely, and keep an independent…

The Tryssh team · July 21, 2026

Feature flag rollout caused an incident: contain and learn

Feature flag rollout caused an incident: contain and learn. A practical production guide with diagnostic commands, failure interpretation, and a safe…

The Tryssh team · July 21, 2026

Fly Volumes and the hidden work behind stateful open source

Fly.io can deploy open-source software quickly, but its runtime and storage model create tradeoffs. Here is where the friction appears and when a VPS is…

The Tryssh team · July 21, 2026

Run a disaster-recovery game day that teaches the truth

Run a disaster-recovery game day that teaches the truth. A practical production guide with diagnostic commands, failure interpretation, and a safe decision…

The Tryssh team · July 21, 2026

GitHub Actions permission denied after token hardening

GitHub Actions permission denied after token hardening. A practical production guide with diagnostic commands, failure interpretation, and a safe decision…

The Tryssh team · July 21, 2026

GitHub Actions self-hosted runner offline: service and network debugging

GitHub Actions self-hosted runner offline: service and network debugging. A practical production guide with diagnostic commands, failure interpretation,…

The Tryssh team · July 21, 2026

Health checks vs synthetic monitoring: green is not enough

Health checks vs synthetic monitoring: green is not enough. A practical production guide with diagnostic commands, failure interpretation, and a safe…

The Tryssh team · July 21, 2026

Heroku's ephemeral filesystem versus file-backed open source apps

Heroku can deploy open-source software quickly, but its runtime and storage model create tradeoffs. Here is where the friction appears and when a VPS is…

The Tryssh team · July 21, 2026

HTTP 502 vs 503 vs 504: use the status as a starting clue

HTTP 502 vs 503 vs 504: use the status as a starting clue. A practical production guide with diagnostic commands, failure interpretation, and a safe…

The Tryssh team · July 21, 2026

Idle servers cost more than compute: patching and attack surface

Idle servers cost more than compute: patching and attack surface. A practical production guide with diagnostic commands, failure interpretation, and a safe…

The Tryssh team · July 21, 2026

Immutable deployment artifacts: stop rebuilding on the server

Immutable deployment artifacts: stop rebuilding on the server. A practical production guide with diagnostic commands, failure interpretation, and a safe…

The Tryssh team · July 21, 2026

API latency spike: a no-panic investigation

API latency spike: a no-panic investigation. A time-boxed incident workflow: verify impact, gather high-signal evidence, restore service safely, and…

The Tryssh team · July 21, 2026

Credential rotation after a server incident: the order matters

Credential rotation after a server incident: the order matters. Audit the exposure, understand the threat model, apply changes safely, and keep an…

The Tryssh team · July 21, 2026

Application cannot reach the database: trace the whole path

Application cannot reach the database: trace the whole path. A time-boxed incident workflow: verify impact, gather high-signal evidence, restore service…

The Tryssh team · July 21, 2026

The deploy broke production: rollback first or debug first?

The deploy broke production: rollback first or debug first?. A time-boxed incident workflow: verify impact, gather high-signal evidence, restore service…

The Tryssh team · July 21, 2026

DNS change went wrong: TTLs, caches, and split answers

DNS change went wrong: TTLs, caches, and split answers. A time-boxed incident workflow: verify impact, gather high-signal evidence, restore service safely,…

The Tryssh team · July 21, 2026

Docker Compose outage: dependency order is not readiness

Docker Compose outage: dependency order is not readiness. A time-boxed incident workflow: verify impact, gather high-signal evidence, restore service…

The Tryssh team · July 21, 2026

High I/O wait: find the disk queue behind the slowdown

High I/O wait: find the disk queue behind the slowdown. A time-boxed incident workflow: verify impact, gather high-signal evidence, restore service safely,…

The Tryssh team · July 21, 2026

Load balancer says unhealthy while the app looks fine

Load balancer says unhealthy while the app looks fine. A time-boxed incident workflow: verify impact, gather high-signal evidence, restore service safely,…

The Tryssh team · July 21, 2026

No space left on device: disk blocks, inodes, and deleted files

No space left on device: disk blocks, inodes, and deleted files. A time-boxed incident workflow: verify impact, gather high-signal evidence, restore…

The Tryssh team · July 21, 2026

The OOM killer terminated your app: what to inspect next

The OOM killer terminated your app: what to inspect next. A time-boxed incident workflow: verify impact, gather high-signal evidence, restore service…

The Tryssh team · July 21, 2026