Reference · 14 commands

Command Atlas

Not manual pages — the handful of invocations of each tool that are actually worth remembering, and the mistake each one invites when you are moving quickly.

ss

Networking

Socket statistics. The replacement for netstat, and considerably faster on a busy host.

$ ss -tlnp

Every TCP socket listening, with the owning process

$ ss -tnp state established

Established connections and who owns them

$ ss -tn state time-wait | wc -l

Count TIME_WAIT sockets before blaming them

$ ss -tnpi dst 10.0.0.5

Connections to one peer, with RTT and congestion window

$ ss -s

Summary totals by socket state

Watch out. Recv-Q on a listening socket is the accept backlog, not buffered data. Non-zero there means the application is not calling accept() fast enough.

lsof

Linux

What files, sockets and devices a process has open — and who has a given file open.

$ lsof -nP -iTCP:8080 -sTCP:LISTEN

Who is listening on a port

$ lsof +L1

Deleted files still held open, which is where missing disk space hides

$ lsof -p <PID>

Everything one process has open

$ lsof /var/log/app.log

Which processes hold this file

$ lsof -u www-data

Everything one user has open

Watch out. -n skips DNS and -P skips port name lookup. Without them lsof can take a very long time on a host with unreachable DNS.

strace

Process

Every syscall a process makes. The definitive answer to "what is it actually doing".

$ strace -p <PID> -f -tt

Attach to a running process and all its threads, with timestamps

$ strace -c -f -p <PID>

Histogram instead of a firehose — the first thing to run

$ strace -e trace=openat,stat -f <cmd>

What files it looks for, including the ones it fails to find

$ strace -e trace=network -f -p <PID>

Network syscalls only

$ strace -f -o out.txt <cmd>

Capture to a file for something that fails at startup

Watch out. strace stops the process at every syscall. On a busy service the slowdown is order-of-magnitude, so attach briefly and never silently.

tcpdump

Networking

Capture packets. The final arbiter when two sides disagree about what was sent.

$ tcpdump -i any -nn port 443 -c 100

First hundred packets on a port, no name resolution

$ tcpdump -i eth0 -s 0 -w cap.pcap host 10.0.0.5

Full packets to a file for later analysis

$ tcpdump -i any -nn 'tcp[tcpflags] & (tcp-syn|tcp-rst) != 0'

Connection attempts and refusals only

$ tcpdump -i any -nn -A port 80

Print payload as ASCII for plaintext protocols

Watch out. Capturing on the box you are SSHed into records your own session, which then generates more traffic. Exclude it: `not port 22`.

dig

Networking

Query DNS directly, without the resolver library getting in the way.

$ dig +short example.com

Just the answer

$ dig example.com @8.8.8.8

Ask a specific resolver, to compare against your own

$ dig +trace example.com

Walk the delegation from the root — finds which step breaks

$ dig -x 10.0.0.5

Reverse lookup

$ dig example.com AAAA

A and AAAA are separate records and fail separately

Watch out. dig bypasses NSS, so it can succeed while the application fails. `getent hosts` follows the same path the application does.

openssl s_client

TLS

Speak TLS to a server by hand and see exactly what it presents.

$ openssl s_client -connect host:443 -servername host </dev/null

The handshake, chain and verify result

$ openssl s_client -connect host:443 -showcerts </dev/null

Every certificate sent — check the chain is complete

$ openssl s_client -connect host:443 -cert c.crt -key c.key

Present a client certificate for mutual TLS

$ openssl s_client -connect host:443 -tls1_2 </dev/null

Force a protocol version to test support

Watch out. Without -servername no SNI is sent, so a shared host answers with its default certificate and you debug the wrong one.

ps

Process

A snapshot of processes, with whichever columns you actually need.

$ ps -eo pid,ppid,stat,pcpu,pmem,etime,comm --sort=-pcpu | head

Top CPU consumers with parentage and age

$ ps -L -o pid,tid,pcpu,comm -p <PID>

Per thread, which is where the CPU actually is

$ ps -eo pid,stat,wchan:30,comm | grep " D"

Processes stuck in uninterruptible sleep

$ ps --ppid <PID>

Children of a process

Watch out. STAT is the fastest diagnosis available. D cannot be killed, Z is already dead, T has been stopped by a signal.

stat

Linux

Everything the filesystem records about a file, including the three timestamps.

$ stat file

Size, inode, permissions, and atime/mtime/ctime

$ stat -c "%n %s %U:%G %a %y" file

A format you can put in a script

$ stat -f /path

The filesystem rather than the file

Watch out. ctime is inode change time, not creation time. It moves when permissions or ownership change, and Linux mostly does not record a creation time at all.

journalctl

Linux

Query the systemd journal, which is where the evidence usually is.

$ journalctl -u nginx -S -1h --no-pager

One unit, last hour

$ journalctl -p err -S today

Errors and worse since midnight

$ journalctl -k -S -10min

Kernel messages — OOM kills and hardware errors land here

$ journalctl -u app -f

Follow, like tail -f

$ journalctl --disk-usage

How much space the journal is holding

Watch out. A journal without persistent storage is lost on reboot. If you need history after a crash, `Storage=persistent` must be set before the crash.

find

Linux

Walk a tree with predicates. The most useful and most misused tool on the box.

$ find /var -xdev -type f -size +1G

Large files, staying on one filesystem

$ find /etc -mmin -60 -type f

Changed in the last hour — what did that deploy touch

$ find / -xdev -type f -perm -4000 2>/dev/null

setuid binaries, for an audit

$ find . -name "*.log" -mtime +30 -delete

Delete carefully, and run it without -delete first

$ find . -newer reference-file

Everything modified after a known point in time

Watch out. -xdev keeps find on one filesystem. Without it you walk /proc, /sys and every network mount, which is slow and produces nonsense.

curl

Networking

Make an HTTP request and see precisely what happened.

$ curl -sv https://host/path

Headers, TLS handshake and response

$ curl -s -o /dev/null -w '%{http_code} %{time_total}s\n' url

Status and timing only, for scripting

$ curl -w '@curl-format.txt' -o /dev/null -s url

DNS, connect, TLS and first-byte timings separately

$ curl --resolve host:443:10.0.0.5 https://host/

Test one backend directly while keeping the right SNI and Host

$ curl --cacert ca.pem --cert c.crt --key c.key https://host/

Mutual TLS from the command line

Watch out. --resolve is the right way to test a single backend behind a load balancer. Editing /etc/hosts changes it for everything on the machine.

nc

Networking

Open a raw TCP connection. The simplest possible reachability test.

$ nc -vz host 443

Is the port open — refused and timeout mean different things

$ nc -vz -w 3 host 443

With a timeout, so a drop fails fast

$ nc -l 9000

Listen, to test connectivity from the other direction

$ nc -u -vz host 53

UDP, where success is much less meaningful

Watch out. Reaching a port proves TCP works, nothing more. A healthy TCP connection to a broken application still fails every request.

vmstat

Linux

System-wide activity over time. The fastest way to characterise what kind of busy a box is.

$ vmstat 1 10

Ten one-second samples — never trust the first line

$ vmstat -s

Totals since boot

$ vmstat -d

Per-disk statistics

Watch out. The first line is an average since boot and is almost always misleading. Read from the second line onwards.

kubectl

Kubernetes

The cluster API, from the command line. These are the invocations that matter during an incident.

$ kubectl logs <pod> --previous

The logs of the container that died, not the fresh one

$ kubectl get events --sort-by=.lastTimestamp | tail -30

Events in the order they happened, which is not the default

$ kubectl get endpoints <svc>

Whether a Service has any backends at all

$ kubectl describe pod <pod>

Probes, limits, last state and exit code in one place

$ kubectl debug -it <pod> --image=nicolaka/netshoot

A shell with network tools beside a distroless container

Watch out. Events are namespaced and expire, by default after an hour. If an incident is older than that, the events are already gone.