Incident Room
Something broke. Now investigate.
Guided troubleshooting workflows for Linux, networking, TLS, Kubernetes, processes, files, and the systems behind them. Start with the symptom rather than the command — every step tells you why you are running it, what to look for, and where that answer sends you next.
/
Linux 4
Load average is climbing and everything feels slow Separate real CPU work from processes stuck waiting on disk, then find the thread actually burning the cycles. Free memory is falling, or something just got killed Tell cache from genuine usage, find what is growing, and read the OOM killer’s own account of what it did. Access refused despite the permissions looking correct Walk the whole path, then the layers most people forget: ACLs, SELinux, capabilities, mount options and namespaces. systemctl says failed, or the service keeps restarting Read the Result and the exit status, which name the cause precisely — systemd’s own codes distinguish a missing binary from a bad user from a readiness timeout.
Networking 3
Connection refused, timed out, or hanging Work outwards from the process to the socket to the firewall. The distinction between refused and timed out narrows it enormously. Name or service not known, or the wrong address comes back Separate the resolver from the record, and find out which of the several DNS paths on a modern machine is actually being used. The load balancer returns 502 Bad Gateway or 503 Service Unavailable The two codes mean different things: 502 is a broken conversation with a backend, 503 is having no backend to talk to.
TLS 2
Certificate errors, handshake failures, or a refused client certificate Get the server to tell you what it is presenting, then work through name, chain, expiry, protocol and client certificates in the order they fail. Clients started rejecting the certificate, or it expires shortly Find which certificate in the chain actually expired, why a renewal did not take effect, and every other place the same certificate is deployed.
Kubernetes 3
A pod starts, dies, and restarts forever Read the previous container’s logs and its exit code — those two facts resolve most crash loops before you touch anything else. The pod will not start because the image will not pull The event message names the cause almost every time — it is a question of knowing which of four failures it describes. A container is killed with exit code 137 Find out whether the limit was too low or the application is leaking, and why a JVM or Node process ignores the limit you set.
Nothing matches that. Try a symptom rather than a cause — “timeout”, “killed”, “refused”.
Reference
Command Atlas
The invocations worth remembering for ss, lsof, strace, tcpdump, dig, openssl, find
and the rest — with the mistake each one invites.
File Investigator
Give it a path and it writes the whole forensic sequence for that file: timestamps,
ownership, ACLs, open handles, checksums, recent changes and the processes touching
it.