Not manual pages — the handful of invocations of each tool that are actually worth
remembering, and the mistake each one invites when you are moving quickly.
/
ss
Networking
Socket statistics. The replacement for netstat, and considerably faster on a busy host.
$ss -tlnp
Every TCP socket listening, with the owning process
$ss -tnp state established
Established connections and who owns them
$ss -tn state time-wait | wc -l
Count TIME_WAIT sockets before blaming them
$ss -tnpi dst 10.0.0.5
Connections to one peer, with RTT and congestion window
$ss -s
Summary totals by socket state
Watch out. Recv-Q on a listening socket is the accept backlog, not buffered data. Non-zero there means the application is not calling accept() fast enough.
lsof
Linux
What files, sockets and devices a process has open — and who has a given file open.
$lsof -nP -iTCP:8080 -sTCP:LISTEN
Who is listening on a port
$lsof +L1
Deleted files still held open, which is where missing disk space hides
$lsof -p <PID>
Everything one process has open
$lsof /var/log/app.log
Which processes hold this file
$lsof -u www-data
Everything one user has open
Watch out. -n skips DNS and -P skips port name lookup. Without them lsof can take a very long time on a host with unreachable DNS.
strace
Process
Every syscall a process makes. The definitive answer to "what is it actually doing".
$strace -p <PID> -f -tt
Attach to a running process and all its threads, with timestamps
$strace -c -f -p <PID>
Histogram instead of a firehose — the first thing to run
$strace -e trace=openat,stat -f <cmd>
What files it looks for, including the ones it fails to find
$strace -e trace=network -f -p <PID>
Network syscalls only
$strace -f -o out.txt <cmd>
Capture to a file for something that fails at startup
Watch out. strace stops the process at every syscall. On a busy service the slowdown is order-of-magnitude, so attach briefly and never silently.
tcpdump
Networking
Capture packets. The final arbiter when two sides disagree about what was sent.
$tcpdump -i any -nn port 443 -c 100
First hundred packets on a port, no name resolution
$tcpdump -i eth0 -s 0 -w cap.pcap host 10.0.0.5
Full packets to a file for later analysis
$tcpdump -i any -nn 'tcp[tcpflags] & (tcp-syn|tcp-rst) != 0'
Connection attempts and refusals only
$tcpdump -i any -nn -A port 80
Print payload as ASCII for plaintext protocols
Watch out. Capturing on the box you are SSHed into records your own session, which then generates more traffic. Exclude it: `not port 22`.
dig
Networking
Query DNS directly, without the resolver library getting in the way.
$dig +short example.com
Just the answer
$dig example.com @8.8.8.8
Ask a specific resolver, to compare against your own
$dig +trace example.com
Walk the delegation from the root — finds which step breaks
$dig -x 10.0.0.5
Reverse lookup
$dig example.com AAAA
A and AAAA are separate records and fail separately
Watch out. dig bypasses NSS, so it can succeed while the application fails. `getent hosts` follows the same path the application does.
openssl s_client
TLS
Speak TLS to a server by hand and see exactly what it presents.
Watch out. Without -servername no SNI is sent, so a shared host answers with its default certificate and you debug the wrong one.
ps
Process
A snapshot of processes, with whichever columns you actually need.
$ps -eo pid,ppid,stat,pcpu,pmem,etime,comm --sort=-pcpu | head
Top CPU consumers with parentage and age
$ps -L -o pid,tid,pcpu,comm -p <PID>
Per thread, which is where the CPU actually is
$ps -eo pid,stat,wchan:30,comm | grep " D"
Processes stuck in uninterruptible sleep
$ps --ppid <PID>
Children of a process
Watch out. STAT is the fastest diagnosis available. D cannot be killed, Z is already dead, T has been stopped by a signal.
stat
Linux
Everything the filesystem records about a file, including the three timestamps.
$stat file
Size, inode, permissions, and atime/mtime/ctime
$stat -c "%n %s %U:%G %a %y" file
A format you can put in a script
$stat -f /path
The filesystem rather than the file
Watch out. ctime is inode change time, not creation time. It moves when permissions or ownership change, and Linux mostly does not record a creation time at all.
journalctl
Linux
Query the systemd journal, which is where the evidence usually is.
$journalctl -u nginx -S -1h --no-pager
One unit, last hour
$journalctl -p err -S today
Errors and worse since midnight
$journalctl -k -S -10min
Kernel messages — OOM kills and hardware errors land here
$journalctl -u app -f
Follow, like tail -f
$journalctl --disk-usage
How much space the journal is holding
Watch out. A journal without persistent storage is lost on reboot. If you need history after a crash, `Storage=persistent` must be set before the crash.
find
Linux
Walk a tree with predicates. The most useful and most misused tool on the box.
$find /var -xdev -type f -size +1G
Large files, staying on one filesystem
$find /etc -mmin -60 -type f
Changed in the last hour — what did that deploy touch
$find / -xdev -type f -perm -4000 2>/dev/null
setuid binaries, for an audit
$find . -name "*.log" -mtime +30 -delete
Delete carefully, and run it without -delete first
$find . -newer reference-file
Everything modified after a known point in time
Watch out. -xdev keeps find on one filesystem. Without it you walk /proc, /sys and every network mount, which is slow and produces nonsense.
curl
Networking
Make an HTTP request and see precisely what happened.
A shell with network tools beside a distroless container
Watch out. Events are namespaced and expire, by default after an hour. If an incident is older than that, the events are already gone.
Nothing matches that.
I would like to count visits with Google Analytics, which sets cookies. The tools
themselves never send your data anywhere either way. What is collected.