Security Notes
Linux

Linux Production Deployment Security — A Practical Primer

How to run a Linux server in production securely: how people log in, what should and shouldn't be running, how you audit and log it, and how you patch and update it. This is the "day-2 operations" view — written in plain language, focused on the decisions that actually matter, with the reasoning behind each so you can defend it in an interview.

13 min read 7 sections 5 model answers verified 2026-06

Last verified2026-06 See also: Linux security fundamentals · Privilege escalation · Linux hardening


The Mental Model: Minimize, Control, Observe, Update

Everything below is one of four ideas:

MINIMIZE  → run as little as possible (fewer packages, fewer ports, fewer privileges)
CONTROL   → tightly govern who gets in and what they can do (access + authz)
OBSERVE   → log and audit everything so you can detect and reconstruct
UPDATE    → keep it patched and reproducible, without snowflakes
Memory hook

"a secure server is a boring, small, watched, fresh one." Boring = does one job. Small = minimal attack surface. Watched = logged and audited. Fresh = patched and rebuilt from code, not hand-tweaked over years. If a production decision doesn't make the box more boring/small/watched/fresh, question it.


1. Logging In — How Humans (and Machines) Access the Box

SSH: the front door

SSH is how you reach a Linux server, and it's the most-attacked service on the internet (bots hammer port 22 constantly). The non-negotiables:

SettingValueWhy
Key auth onlyPasswordAuthentication noPasswords are brute-forceable; keys aren't. This single change stops ~all SSH brute-forcing
No root loginPermitRootLogin noForce login as a named user + sudo → accountability and an audit trail
Modern keysEd25519 (or RSA-4096)See SSH keys in the crypto deep dive
Limit whoAllowUsers/AllowGroupsDefault-deny; only named admins
Idle timeoutClientAliveInterval 300Kill abandoned sessions
2FA where possiblePAM + TOTP, or an SSH CADefense in depth for the human factor
Memory hook

"named users in, root never, keys only, password never." Root login + password auth is the combination every credential-stuffing bot prays for. Logging in as a named user and escalating with sudo is what gives you the "who did what" trail — root login erases identity because everyone is "root."

Better than static keys: short-lived access

The modern pattern for fleets is not copying SSH public keys to every server's authorized_keys (which sprawls and is hard to revoke). Instead:

SSH certificates

an internal SSH CA signs short-lived user certs; servers trust the CA, not individual keys. Rotating/revoking is trivial; no key sprawl. (Same PKI idea as TLS, applied to SSH.)

Bastion / jump host

all SSH funnels through one hardened, heavily-logged host; servers accept SSH only from the bastion. One chokepoint to watch.

No SSH at all (the cloud-native ideal)

use AWS SSM Session Manager or equivalent: shell access brokered by the cloud control plane, fully audited, no open port 22, no keys to manage. Increasingly the best-practice default.

Memory hook

the best SSH port is no SSH port. Every open SSH port is attack surface and a key to manage. Brokered access (SSM Session Manager) or short-lived certs beat static keys, which beat passwords. Aim to remove the standing door, not just lock it.

Privilege: sudo, least privilege, no shared accounts

Individual accounts, never shared logins

admin/deploy shared accounts destroy accountability.

sudo for escalation, logged

grant specific commands where possible, not blanket ALL. Every sudo is logged (/var/log/auth.log / journald).

Service accounts are non-login

daemons run as dedicated users with nologin shells and no home/SSH; they exist to run a service, not to be logged into.


2. What Should (and Shouldn't) Be Running

Run the minimum

Every running service is attack surface. The discipline:

One job per box (or container).

A server that's a web server should not also be a database, a build agent, and a Tor relay. Blast radius shrinks when roles are isolated.

Remove what you don't need.

No compilers, no telnet, no unused interpreters, no leftover dev tools on a production box — these are exactly what attackers use to build a foothold (living off the land).

Audit listening ports.

ss -tulpn shows every listening socket and its process. Anything listening that you can't explain is either a misconfiguration or a compromise. Bind services to 127.0.0.1 unless they genuinely must be reachable.

bash
ss -tulpn          # what's listening, and which process owns it
systemctl list-units --type=service --state=running   # what's running and why
Memory hook

"if you can't name why it's running, turn it off." The single highest-leverage hardening question is "what is listening on the network, and does it need to be?" Most breaches start with a service that was running, exposed, and forgotten. Attack surface you removed can't be exploited.

Run services with least privilege (systemd hardening)

Modern Linux runs services under systemd, which can sandbox them heavily — most teams don't use this and should:

ini
# /etc/systemd/system/myapp.service — a hardened unit
[Service]
User=myapp                    # never root
NoNewPrivileges=true          # the service can't gain privileges (blocks setuid escalation)
ProtectSystem=strict          # filesystem is read-only except explicitly allowed paths
ProtectHome=true              # no access to /home
PrivateTmp=true               # isolated /tmp (can't tamper with others' temp files)
CapabilityBoundingSet=        # drop ALL Linux capabilities (add back only what's needed)
SystemCallFilter=@system-service   # seccomp: only allow normal syscalls
RestrictAddressFamilies=AF_INET AF_INET6   # no weird socket families
Memory hook

assume the service will be popped, then contain it. Running a service as a non-root user with NoNewPrivileges, a read-only filesystem, dropped capabilities, and a seccomp filter means that even a full RCE in the app lands the attacker in a tiny, powerless box. This is the same "container is a process, not a boundary — so sandbox the process" idea, applied with systemd. It turns "app compromised" into "app compromised, and they can't do anything with it."

Containers in production

If you deploy containers, the security rules rhyme:

  • Non-root inside the container (USER directive), read-only root filesystem, drop capabilities, no --privileged, no docker socket mounted.
  • Minimal base images (distroless / Alpine) — fewer packages = fewer CVEs = smaller attack surface.
  • Scan images for CVEs (Trivy/Grype) in CI and at registry.
  • Remember: a container is not a security boundary — it shares the host kernel. For hostile multi-tenant workloads, use stronger isolation (gVisor, Firecracker/microVMs).

3. Auditing & Logging — So You Can See What Happened

You cannot respond to what you can't see. Production logging has three jobs: operational (is it healthy?), security (was it attacked?), and forensic (what exactly happened?).

What to log

SourceCapturesWhy it matters
journald / syslogSystem & service logsBaseline; service crashes, errors
auth.log / secureLogins, sudo, SSH, PAMWho got in, who escalated — the access trail
auditdKernel-level syscall/file auditingThe forensic gold: file access, execve, config changes, with who/when
App logsApplication eventsBusiness-logic abuse, app attacks
NetworkFirewall logs, NetFlow/Flow LogsConnections in/out (C2, exfil, scans)

auditd — the forensic backbone

auditd watches the kernel and records security-relevant events with the responsible user and timestamp. A few high-value rules:

bash
# Watch sensitive files for any change (who modified /etc/passwd?)
-w /etc/passwd -p wa -k identity     # -p wa = alert on Write + Attribute-change; -k = searchable tag
-w /etc/sudoers -p wa -k privilege   # someone editing sudoers = granting themselves root → high signal
-w /etc/ssh/sshd_config -p wa -k sshd

# Record every command execution (execve) — powerful for IR
-a always,exit -F arch=b64 -S execve -k exec

# Watch for loading kernel modules (rootkit indicator)
-w /sbin/insmod -p x -k modules
Memory hook

auth.log tells you who got in; auditd tells you what they did. SSH/sudo logs answer "who logged in and escalated"; auditd answers "what files did they touch, what did they execute, what did they change." Together they reconstruct the timeline. If you run one extra thing on a production box for security, it's auditd with rules on identity files and execve.

Ship logs OFF the box (the rule attackers hate)

Local logs are worthless if the attacker can edit them. The first thing a competent intruder does is clear /var/log. So:

  • Forward logs in real time to a central, append-only store (SIEM, a logging account, an immutable bucket) the production box can write to but not modify or delete.
  • This preserves evidence even if the host is wiped, and enables cross-host correlation and detection.
Memory hook

"logs that live only on the box die with the box." Centralized, immutable logging is what separates "we got owned and have no idea what happened" from "we have the full timeline." An attacker can rm -rf /var/log locally, but they can't reach back into your logging account. Ship logs off-host, make them append-only. (Ties to order of volatility and analyst OPSEC.)

File integrity monitoring

Tools like AIDE take a cryptographic baseline of important files and alert on changes — catching a backdoored binary, a modified sshd, or a new SUID file. Tripwire/AIDE answers "did anything important change that shouldn't have?"


4. Updating & Patching — Staying Fresh Without Snowflakes

Patch, and patch fast

Most breaches exploit known vulnerabilities that had a patch available. So:

Automate security updates

unattended-upgrades (Debian/Ubuntu) or dnf-automatic (RHEL) for security patches at minimum.

Track your exposure

know what's installed (an SBOM helps) and scan for CVEs so you can prioritize the ones that are exploitable and exposed.

Have a fast lane for criticals

when something like a remote-exploitable RCE drops (Log4Shell-class), you need to patch in hours, not the monthly cycle. That requires knowing where the vulnerable component runs.

Immutable infrastructure — the modern way to update

The old way: long-lived servers, patched in place for years, each one a unique "snowflake" nobody fully understands. The modern way is immutable infrastructure:

DON'T patch the running server.
DO  rebuild a fresh, patched image (AMI/container) from code,
    deploy it, and DESTROY the old one.
No SSH-and-fix

in production — changes go through code (IaC + image builds), reviewed and version-controlled.

Reproducible

every server is built identically from the same definition, so there are no mystery configurations.

Easy rollback

bad deploy? redeploy the previous image.

Security bonus

frequent rebuilds wipe any attacker persistence; a box that's recreated every week is hostile ground for a foothold.

Memory hook

"cattle, not pets." Pets are named, hand-fed, nursed back to health when sick (patched in place forever). Cattle are numbered and replaced when they break (rebuild from image, destroy the old). Treating servers as cattle makes them reproducible, patchable, and — as a side effect — resistant to persistence, because an attacker's foothold evaporates at the next deploy. "Don't fix the server, replace it."

Configuration as code

How the box is configured — packages, users, services, firewall — should live in code (Ansible, or baked into the image via Packer), not in someone's memory or ad-hoc commands. Benefits: reviewable, auditable, reproducible, and you can prove what should be there (so drift = suspicious).


The Production Hardening Checklist

ACCESS
[ ] SSH: key-only, no root login, named users + sudo, AllowUsers/Groups
[ ] Prefer brokered access (SSM Session Manager) or SSH certs over static keys
[ ] No shared accounts; service accounts are non-login (nologin shell)
[ ] MFA on human access where feasible

MINIMIZE
[ ] One role per box/container; remove unused packages, tools, interpreters
[ ] Audit listening ports (ss -tulpn); bind to localhost unless must be public
[ ] Host firewall default-deny inbound (nftables/ufw); only required ports open
[ ] Services run non-root with systemd hardening (NoNewPrivileges, ProtectSystem, seccomp)
[ ] Containers: non-root, read-only FS, dropped caps, minimal base, scanned

OBSERVE
[ ] auditd with rules on identity files + execve
[ ] Logs shipped OFF-box in real time to append-only/central store
[ ] File integrity monitoring (AIDE)
[ ] Alerts on: root/sudo use, new SUID, new listening ports, config changes

UPDATE
[ ] Automated security patching (unattended-upgrades / dnf-automatic)
[ ] Fast lane for critical CVEs; know where vulnerable components run
[ ] Immutable infra: rebuild+replace, don't patch in place ("cattle not pets")
[ ] Config as code (Ansible/Packer/IaC), reviewed and version-controlled

Interview Questions

Q
How would you secure SSH access to a fleet of production Linux servers?
Model answer

Start with the basics on every host: key-only authentication with passwords disabled, no direct root login so people log in as named users and escalate with sudo for an audit trail, modern Ed25519 keys, AllowUsers/Groups to default-deny, and idle timeouts. But at fleet scale, static keys in authorized_keys sprawl and are hard to revoke, so I'd move to short-lived access — an internal SSH certificate authority that signs brief user certs the servers trust, or funnel everything through a hardened, logged bastion. The ideal in cloud is to remove standing SSH entirely and use brokered access like SSM Session Manager: no open port 22, no keys to manage, and every session is audited by the control plane. The principle is the best SSH port is no SSH port.

Q
A production server should run as little as possible — how do you put that into practice and verify it?
Model answer

One role per box, minimal packages — no compilers, telnet, or leftover dev tools that attackers use to live off the land — and a host firewall that default-denies inbound, opening only the required ports. I verify with ss -tulpn to enumerate every listening socket and the process behind it; anything I can't explain is a misconfiguration or a compromise, and services that don't need network exposure get bound to localhost. I also run each service non-root under a hardened systemd unit — NoNewPrivileges, a read-only filesystem, dropped capabilities, and a seccomp syscall filter — so even if the app is exploited, the attacker lands in a powerless sandbox. The rule of thumb is if you can't name why something is running, turn it off.

Q
What's your logging strategy for production Linux, and what's the single most important rule?
Model answer

Three layers: system and service logs via journald, the access trail in auth.log for logins/sudo/SSH, and auditd for kernel-level forensic detail — rules watching identity files like /etc/passwd and /etc/sudoers and recording every execve. auth.log tells me who got in and escalated; auditd tells me what they actually did. But the single most important rule is to ship logs off the box in real time to a central, append-only store the server can write to but not modify or delete. Local logs are worthless because the first thing a competent attacker does is clear /var/log — centralized immutable logging is what preserves the evidence and enables cross-host detection. Logs that live only on the box die with the box.

Q
How do you keep production servers patched, and what is immutable infrastructure?
Model answer

At minimum, automate security updates with unattended-upgrades or dnf-automatic, track what's installed so I can prioritize exploitable, exposed CVEs, and keep a fast lane to patch critical RCEs in hours rather than waiting for a monthly window. But the better model is immutable infrastructure: instead of SSHing in to patch a long-lived server — which produces unique snowflakes nobody fully understands — I rebuild a fresh, patched image from code, deploy it, and destroy the old one. That makes every server reproducible and identical, makes rollback trivial, and as a security bonus wipes any attacker persistence on each deploy, since a box rebuilt frequently is hostile ground for a foothold. It's cattle, not pets: don't nurse the server back to health, replace it.

Q
Why run a service as non-root with systemd hardening if the app is "trusted"?
Model answer

Because you assume the app will eventually be compromised and you want to contain the blast radius when it is. No application is immune to a vulnerability, and if it runs as root with full capabilities and a writable filesystem, a single RCE gives the attacker the whole box. Running it as a dedicated non-root user with NoNewPrivileges so it can't escalate via setuid, ProtectSystem making the filesystem read-only, PrivateTmp, dropped Linux capabilities, and a seccomp syscall filter means the same RCE lands the attacker in a tiny, powerless sandbox — they can't write system files, can't gain privileges, can't make unexpected syscalls. It's defense in depth: trust isn't a control, containment is, and systemd gives you that containment essentially for free.