Security Notes
Cloud

AWS IAM — The Deep Dive

IAM (Identity and Access Management) is the single most important topic in any AWS security interview. It's the authorization brain of AWS — every API call, from reading an S3 object to launching a server, is evaluated by IAM. This file explains IAM in plain language first, then goes deep on the parts interviewers always probe: roles vs users, the policy types and how they're evaluated, permission boundaries, how compute gets credentials, how to provision users properly, and exactly what to do when an access key leaks.

36 min read 15 sections 12 model answers verified 2026-06

Last verified2026-06 See also: AWS fundamentals (product primer) · AWS incident response


What IAM Actually Is

Every request to AWS — clicking in the console, an aws s3 cp, a Lambda reading a secret — is an API call. Before AWS performs it, IAM answers one question: "is this principal allowed to perform this action on this resource, under these conditions?"

That's the whole job. IAM is a giant policy evaluation engine sitting in front of every AWS service. Master "who is asking, what are they asking to do, to which resource, and which policies apply," and you've understood IAM.

Memory hook

every AWS action is Principal + Action + Resource + Condition. Those four words are the shape of every IAM decision and every IAM policy statement. When debugging "why is this denied?", walk those four: who is the principal, what action, which resource, and do the conditions (IP, MFA, region, tags) match?


Identity vs Policy — the two things people conflate

The most common IAM confusion, and worth nailing before anything else: an identity and a policy are completely different objects.

  • An identity is who you are — a badge. IAM users, groups, and roles are identities. AWS authenticates them, and they can hold or assume credentials.
  • A policy is the rules — a JSON document full of Allow/Deny statements. It has no credentials and authenticates nobody; it does nothing on its own. It's inert paper until you attach it to an identity (or to a resource).

The badge opens no doors by itself; the rulebook is just paper until it's pinned to a badge. Access = an identity with a policy attached. Hand the same person a different rulebook and they reach different rooms — the badge never changed, the rules did.

   IDENTITY — the "who" (a badge)          POLICY — the "rules" (a rulebook)
   ┌──────────────────────────┐           ┌───────────────────────────────────┐
   │  IAM user   "alice"        │  attach   │ {                                  │
   │  IAM group  "devs"         │◄─────────►│   "Effect":   "Allow",             │
   │  IAM role   "deploy"       │           │   "Action":   "s3:GetObject",      │
   │                            │           │   "Resource": "arn:…:bucket/*"     │
   │  • AWS authenticates it    │           │ }                                  │
   │  • holds / assumes creds   │           │  · pure JSON — no creds, no powers │
   └──────────────────────────┘           │  · does NOTHING until attached     │
            ▲                               └───────────────────────────────────┘
            └─ one policy can attach to many identities (and an identity to many policies)

   an identity with NO policy attached can do NOTHING  (default deny)

What a policy actually looks like — JSON with one or more statements, each one answering IAM's four questions:

   {
     "Effect":    "Allow",                  ← Allow or Deny
     "Action":    "s3:GetObject",           ← WHAT operation
     "Resource":  "arn:aws:s3:::bucket/*",  ← on WHICH resource
     "Condition": { "Bool": { "aws:MultiFactorAuthPresent": "true" } }   ← WHEN (optional)
   }

Notice there's no Principal field here — that's the giveaway of an identity-based policy: the identity it's attached to is the "who". A Principal field only appears in resource-based policies (a bucket/key/role-trust policy), which name who may touch the resource.

Principal vs identity (another easy mix-up): principal is the broad term — anything that can make a request: an identity, an AWS service like ec2.amazonaws.com, or even anonymous ("*"). Identity is the narrower kind of principal you create and attach policies to — users, groups, roles. (A group isn't something you log in as; it's an identity you attach policies to, and its users inherit them.)

Memory hook

identity ≠ policy"Who you are" and "what's allowed" are separate objects you wire together by attaching. Identities authenticate but carry no permissions of their own; policies carry permissions but authenticate nobody. A user with no policy can do nothing; a policy attached to nothing affects nothing.

The next two sections take each in turn — Principals (the identities) and, further down, Policies (the rules).


Principals — the "who"

A principal is anything that can make an AWS request. The types, and how to think about each:

PrincipalWhat it isLifespan of credsUse it for
Root userThe original account owner, all-powerful, tied to the account's emailPermanentAlmost nothing — lock it away (see below)
IAM userA named identity with long-lived credentials (password and/or access keys)Permanent until rotatedLegacy humans/services — increasingly discouraged
IAM roleAn identity with no permanent credentials that principals assume to get temporary credentialsTemporary (minutes–hours)Almost everything — humans via SSO, EC2, Lambda, cross-account
Federated identityA user from an external IdP (Okta, Entra, Google) mapped to a roleTemporaryHuman workforce access (the modern way)
AWS service principalAn AWS service itself (e.g. ec2.amazonaws.com) acting on your behalfn/aLetting a service assume a role
Memory hook

the whole modern IAM philosophy: roles over users, temporary over permanent. IAM users have long-lived access keys that sit in files waiting to be stolen. IAM roles hand out credentials that expire in an hour. The single biggest IAM improvement most orgs can make is "stop creating IAM users with access keys; use roles and federation." If you remember one design principle, it's this one.


The Root User — handle with extreme care

The root user is the account's god-mode identity, created with the account and tied to its email address. It can do everything, including things no IAM policy can restrict (close the account, change billing, certain S3/KMS recovery actions). An SCP can't even fully restrict the root of the management account.

Correct handling of root

Don't use it for daily work.

Create IAM roles/identities for everything operational.

Enable a hardware MFA

on it and store the credentials in a safe/break-glass vault.

Delete root access keys

root should never have programmatic access keys. (AWS now blocks creating them.)

Alarm on any root usage

a CloudTrail/GuardDuty alert on root login or root API calls is a top-priority signal, because legitimate root use is rare and attacker root use is catastrophic.

Memory hook

Fun fact — root is the one identity guardrails can't fully cage. Service Control Policies, which can restrict every other identity in an org, do not restrict the management account's root. That's why "secure the root, MFA it, and alarm on its use" is AWS security 101 — it's the one key that opens every lock.


Roles & AssumeRole — the heart of IAM

A role is the most important and most misunderstood IAM concept. A role is an identity with permissions but no permanent credentials. To use it, a principal assumes it via sts:AssumeRole, receiving temporary credentials (an access key + secret + session token) that expire.

Every role has two policies, and confusing them is the classic mistake:

            ┌─────────────────────── ROLE ───────────────────────┐
            │                                                     │
   TRUST POLICY (who can assume me?)        PERMISSION POLICY (what can I do?)
   = a resource-based policy on the role    = identity-based policy
   "Principal X is allowed to AssumeRole"   "Allow s3:GetObject on bucket Y"
Trust policy

who is allowed to become this role. (e.g. "the EC2 service" or "account 123's CI role" or "users from our Okta").

Permission policy

what the role can do once assumed.

json
// Trust policy: allow EC2 instances to assume this role
{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Principal": { "Service": "ec2.amazonaws.com" },
    "Action": "sts:AssumeRole"
  }]
}
Memory hook

a role is a hat, not a person. Nobody is the role; principals put on the hat (assume it) temporarily, do the work with its powers, and take it off (the creds expire). The trust policy is the bouncer deciding who's allowed to wear the hat; the permission policy is what the hat lets you do. "AssumeRole = put on the hat and get temporary keys."

AssumeRole, step by step — "alice" assuming the "deploy" role. Note the two gates that must both pass:

   alice (user)              STS                       IAM (checks 2 gates)
     │  AssumeRole(deploy)    │                              │
     ├───────────────────────►│  gate 1: does alice's policy │
     │                        │          allow sts:AssumeRole?├─┐
     │                        │  gate 2: does deploy's TRUST  │ │
     │                        │          policy allow alice?  │◄┘
     │  temp creds            │  both yes → mint creds for     │
     │  (ASIA… + session      │  the deploy role               │
     │   token + expiry)      │                                │
     │◄───────────────────────┤                                │
     │
     ▼  every later call runs as  assumed-role/deploy/alice ,
        and the DEPLOY role's PERMISSION policy decides what's allowed.
        (CloudTrail logs the AssumeRole and each subsequent call.)

Miss either gate — alice lacks sts:AssumeRole, or the role's trust policy doesn't name her — and you get AccessDenied. The trust policy lives on the role (it's a resource-based policy); alice's sts:AssumeRole lives on her identity.

Reading a credential — AKIA vs ASIA

Every set of AWS credentials starts with an access key ID, and its four-letter prefix tells you at a glance what kind it is — useful in CloudTrail and when triaging a leaked credential:

   AKIA EXAMPLE...                    ASIA EXAMPLE...
   + secret access key                + secret access key
                                       + SESSION TOKEN   ◄── the extra piece
   = 2 parts, NEVER expires           = 3 parts, EXPIRES (minutes–hours)
   a long-lived IAM USER key          a temporary STS credential
   (the kind that leaks to GitHub)    (from AssumeRole / IRSA / federation)
AKIA…

a long-lived key belonging to an IAM user. No expiry; valid until someone rotates or deletes it. This is the kind bots scan GitHub for.

ASIA…

a temporary credential minted by STS (every AssumeRole, IRSA, or federated login). It requires a session token alongside the key+secret, and it expires on its own. Seeing ASIA means "a role session — find which role and who assumed it; the creds will die by themselves."

Memory hook

Two more prefixes show up in the userIdentity block of CloudTrail (these are unique-ID prefixes, not access keys): AROA… = a role object, AIDA… = a user object. So assumed-role/deploy/... with an AROA… principalId is a role session; an AIDA… is a plain IAM user. The quick read: AKIA = permanent user key (rotate it), ASIA = temporary role session (it expires).


How Compute Gets Credentials (no hardcoded keys!)

A core interview theme: how does an EC2 instance / Lambda / EKS pod call AWS without anyone hardcoding an access key? Answer: they assume a role and AWS injects temporary credentials.

ComputeMechanismHow creds arrive
EC2Instance profile (a wrapper around a role)The instance fetches temp creds from the instance metadata service (IMDS) at 169.254.169.254
LambdaExecution roleAWS injects temp creds into the function's environment
ECS/FargateTask roleCreds served from a task-local metadata endpoint
EKS podIRSA (IAM Roles for Service Accounts) or EKS Pod IdentityThe pod's Kubernetes service account is mapped to an IAM role; pod gets a projected token → assumes the role via OIDC
Memory hook

the metadata endpoint is the magic and the danger. EC2 gets its credentials from http://169.254.169.254. That's elegant (no stored keys) but it's exactly why SSRF against EC2 is so dangerous: trick the app into fetching that URL and you steal the instance's role credentials. The Capital One breach (2019) was precisely this — SSRF → IMDS → role creds → S3 exfiltration of 100M+ records. The fix is IMDSv2, which requires a session token obtained via a PUT that SSRF typically can't perform. Always: "EC2 creds come from IMDS; protect IMDS with v2."

Memory hook

EKS: IRSA vs Pod IdentityBoth solve "give a specific pod (not the whole node) its own IAM role." IRSA maps a Kubernetes service account to an IAM role via an OIDC trust relationship (older, more setup). EKS Pod Identity (newer, 2023) does the same with a simpler agent-based association. The security point either way: scope IAM to the pod's service account, never to the node's instance role — otherwise every pod on the node inherits the node's permissions (a lateral-movement goldmine).


Policies — the rules

A policy is a JSON document of Allow/Deny statements. There are several types, distinguished by what they attach to and whether they can grant or only restrict.

The policy types (readable version)

TypeAttaches toGrants or only restricts?Think of it as
Identity-baseduser, group, roleGrants"What this identity can do"
Resource-basedS3 bucket, KMS key, SQS, Lambda, role trust policyGrants (incl. cross-account)"Who can touch this resource"
SCP (Service Control Policy)Org OU/accountOnly restricts (a ceiling)"Org-wide guardrail"
Permission boundaryuser or roleOnly restricts (a ceiling)"Max this identity could ever have"
Session policypassed at AssumeRole timeOnly restricts"Shrink this one session"
Memory hook

grant vs. ceilingOnly identity and resource policies grant access. SCPs, permission boundaries, and session policies are ceilings — they can only take away, never add. So if something's denied, check the ceilings; if something needs to be allowed, it must be granted by an identity or resource policy and permitted by every ceiling above it.

How a request is evaluated (the logic interviewers love)

A request is ALLOWED only if ALL of these hold:
  1. No EXPLICIT DENY anywhere            (an explicit Deny ALWAYS wins, full stop)
  2. Permitted by every SCP               (org guardrail)
  3. Permitted by the permission boundary (if one is set)
  4. Permitted by the session policy       (if assumed with one)
  5. GRANTED by an identity policy OR a resource policy

Same-account:  identity policy OR resource policy granting is enough.
Cross-account: you need BOTH — the identity policy in the caller's account
               AND the resource policy in the target account.
Memory hook

"explicit deny always wins; default is deny." Two rules cover 90% of IAM evaluation questions: (1) the default is deny — if nothing explicitly allows it, it's denied; (2) an explicit Deny beats any Allow — no matter how many Allows exist. So SCPs and boundaries enforce limits by denying, and you can't accidentally out-allow a Deny. Say it back in interviews verbatim: "default deny, explicit deny overrides allow."


Permission Boundaries — safe delegation

A permission boundary is a ceiling on what an IAM user or role can do, regardless of how permissive its attached policies are. Its killer use case is delegating IAM safely.

The problem it solvesyou want developers to create their own IAM roles (for their Lambdas, etc.), but if you give them iam:*, a developer could create an admin role and assume it — instant privilege escalation. The fix: require that any role they create has a permission boundary attached, capping it at, say, "S3 + DynamoDB only." Now even a maliciously-crafted role can't exceed the boundary.

Developer's effective permissions = (their identity policy) ∩ (their permission boundary)
                                     = the INTERSECTION, never more than the boundary
Memory hook

boundary = "you can grant, but not beyond this line." A permission boundary lets you hand someone the power to create permissions while guaranteeing they can never create more than you allowed. It's the "intersection, not union" idea: effective access is the overlap of what's granted and what the boundary permits. This is the answer to "how do you let developers self-serve IAM without letting them escalate?"


Service Control Policies (SCPs) — org-wide guardrails

SCPs apply at the AWS Organizations level (an OU or account) and set the maximum permissions for everything in that account — they're a ceiling that even account admins can't exceed.

Common SCP uses (preventive controls):

Deny disabling CloudTrail/GuardDuty/Config

so an attacker who gets admin still can't blind your logging.

Region lock

deny all actions outside approved regions (limits where an attacker can spin up crypto-mining).

Deny root usage

in member accounts.

Deny leaving the organization

.

json
// SCP: prevent anyone (even account admins) from disabling CloudTrail
{
  "Effect": "Deny",
  "Action": ["cloudtrail:StopLogging", "cloudtrail:DeleteTrail"],
  "Resource": "*"
}
Memory hook

SCPs are the guardrail, not the road. They never grant anything — an empty SCP allows nothing extra. They define the outer fence inside which account-level IAM operates. The classic SCP win: even if an attacker becomes account admin, an SCP that denies cloudtrail:StopLogging means they can't turn off the cameras — your evidence keeps flowing.


AWS Organizations, Multi-Account & Blast-Radius Guardrails (SCPs + RCPs)

Last verified2026-06

This is the section interviewers use to separate "I clicked around one AWS account" from "I've designed a production landing zone." It ties together three things: how you set an organisation up from scratch, how you split it into accounts, and which guardrails (SCPs + RCPs) you pin where to contain a breach. We'll end with a worked example for a real workload — an account that runs only S3, EKS, and EC2 — and a diagram of exactly what goes where.

What is an AWS Organization?

An AWS Organization is a way to manage many AWS accounts as one tree. At the top is a single management account (the one that creates the org — historically called the "payer" account because all the bills roll up to it). Under it you build a tree of Organizational Units (OUs) — folders — and drop member accounts into those folders.

        AWS ORGANIZATION  (one tree, billing rolls up to the top)
        ┌─────────────────────────────────────────────────────┐
        │  MANAGEMENT ACCOUNT  (root of the org)                │
        │   • owns the org, billing, SCPs/RCPs                   │
        │   • runs almost NO workloads (keep it nearly empty)    │
        └───────────────────────────┬──────────────────────────┘
                                     │
        ┌──────────┬─────────────────┼──────────────┬───────────┐
       OU         OU                OU             OU          OU
    Security   Infrastructure    Workloads      Sandbox    Suspended
       │            │            ┌───┴────┐         │           │
   ┌───┴───┐    ┌───┴───┐      Prod-OU  Nonprod-OU  …       (quarantine)
 Log-     Audit  Network  Shared  │        │
 Archive  /Sec  (VPCs)   Services prod    dev/stage
 acct     acct                    acct    accts

Why split into many accounts at all? An AWS account is the strongest isolation boundary AWS gives you — stronger than a VPC, stronger than an IAM role. A blast radius is naturally capped at the account: a compromised role in dev can't touch prod resources because they live in a different account with different credentials. The pattern is "one account per workload per environment" — e.g. payments-prod, payments-dev, analytics-prod — plus a few shared accounts.

Memory hook

the account is the blast-radius boundary"Why not just run everything in one account with IAM separating teams?" Because IAM is one misconfigured Resource: "*" away from leaking across teams, whereas crossing an account boundary requires an explicit cross-account trust you can see and audit. Separate accounts make isolation the default and sharing the deliberate exception. This is the single most important multi-account talking point.

Setting it up from scratch — the landing zone

You almost never hand-build this. The from-scratch sequence, in order:

1. Start from a clean account → it becomes the MANAGEMENT account.
   Enable AWS Organizations (creates the org root).

2. Turn on AWS CONTROL TOWER (or Landing Zone Accelerator).
   This automates steps 3–6 below into a "landing zone".

3. It creates the foundational accounts:
   • Log Archive  → central, write-once CloudTrail/Config log bucket
   • Audit/Security → GuardDuty, Security Hub, cross-account read roles

4. It creates a baseline OU layout (Security OU, Sandbox OU, …)
   and enrolls accounts into it.

5. It wires IAM IDENTITY CENTER (SSO) → humans log in centrally and
   assume roles into each account. NO IAM users anywhere.

6. It applies a baseline set of SCPs (and you add RCPs) as guardrails.

7. You vend NEW accounts on demand via Account Factory — each one is
   born already enrolled, logged, and guard-railed.
Memory hook

"Organizations is the plumbing, Control Tower is the house." Organizations gives you the raw tree, accounts, and policy attachment points. Control Tower is the opinionated layer on top that stamps out a secure-by-default landing zone — the log-archive account, the SSO wiring, the baseline guardrails — so every new account starts compliant instead of being hardened by hand. In an interview: "I wouldn't hand-roll the org; I'd use Control Tower / Landing Zone Accelerator so logging, SSO, and guardrails are baked in from account #1."

Two kinds of guardrail: SCP (the principals) vs RCP (the resources)

Until late 2024 the org only had one guardrail type — the SCP. SCPs cap what the identities in your accounts can do. But they say nothing about who, from outside, can reach your resources. Resource Control Policies (RCPs) — launched November 2024 — fill that gap: they cap what can be done to the resources in your accounts, no matter who is asking. Together they form a data perimeter — the two halves of one fence.

        ┌──────────── YOUR ACCOUNT ────────────┐
        │                                        │
  SCP ──┤  the IDENTITIES here                   │── RCP
 caps   │  (roles, users)  ───► can act on ───►  │   caps
 what   │                          RESOURCES     │   who/how
 THEY   │                       (S3, KMS, STS…)  │   can touch
 can do │                                        │   the RESOURCE
        └────────────────────────────────────────┘

  SCP answers:  "what may the principals in my accounts do?"   (identity side)
  RCP answers:  "who may touch the resources in my accounts?"  (resource side)
SCP (Service Control Policy)RCP (Resource Control Policy)
Caps the permissions ofPrincipals (identities) in the accountResources in the account
Classic question"Can my roles call X?""Can anyone — even an external account — read my bucket?"
Attaches toRoot / OU / accountRoot / OU / account
Grants anything?No — ceiling onlyNo — ceiling only
Applies to root user?No (mgmt acct root exempt)Yes — RCPs also bound the resource owner's root
Services coveredAll servicesS3, STS, KMS, SQS, Secrets Manager (Nov 2024); + ECR, OpenSearch Serverless (Jun 2025)
Default attached policyFullAWSAccessRCPFullAWSAccess
Memory hook

SCP = "what my people can do", RCP = "who can touch my things". The reason RCPs matter: an SCP can't stop a misconfigured resource policy from sharing your S3 bucket with the whole internet, because the SCP only governs your principals, not the external one reading the bucket. An RCP sits on the resource side and can enforce "no principal outside my org may access this bucket — full stop", overriding any too-generous bucket policy. RCPs are how you enforce a data perimeter centrally instead of auditing thousands of bucket policies.

Resource policies — the best-practice recommendation

A resource-based policy (a bucket policy, KMS key policy, role trust policy, SQS/Secrets policy) is the Principal-bearing policy that lives on the resource and says who may touch it. The recommendations interviewers want to hear:

Default to identity-based policies; use resource policies for the cases only they can do

namely cross-account access and letting an AWS service in (e.g. CloudTrail writing to your log bucket). Don't scatter access logic across both sides if one will do; it gets unauditable fast.

Never use "Principal": "*" without a Condition that scopes it.

A wildcard principal with no condition is how buckets end up public. If you need broad access, fence it with aws:PrincipalOrgID, aws:SourceArn, aws:SourceAccount, or a VPC endpoint condition.

Pin cross-account trust to your org

, not to a bare account: aws:PrincipalOrgID is more durable than listing account IDs.

Always require TLS

(aws:SecureTransport: true) and, on S3, enforce Block Public Access at the account level so no bucket policy can accidentally go public.

Then enforce all of the above centrally with an RCP

, so a single fat-fingered bucket policy in one of 200 accounts can't punch a hole — the RCP is the backstop the per-resource policy can't override.

Memory hook

resource policy is the lock on the door; the RCP is the building's master rule. Per-resource policies are easy to get subtly wrong at scale. The modern answer to "how do you keep 500 buckets from leaking?" isn't "review every bucket policy" — it's "enforce a data-perimeter RCP at the org so external principals are denied by default, and let bucket policies grant only inside that fence."

Worked example — an account running only S3, EKS, EC2

Now the concrete ask: a production workload account that runs S3 in one region (say eu-west-1), EKS, and EC2 — nothing else. Here is exactly which guardrails to apply and where, to keep the blast radius tiny if a role in that account is compromised.

Where the policies attach (the graph)

 ORG ROOT
   │  RCP:  ✦ data-perimeter      (deny any principal NOT in my org from
   │         (org-wide backstop)    touching S3/STS/KMS/SQS/Secrets/ECR)
   │  RCP:  ✦ enforce-TLS          (deny non-HTTPS on those resources)
   │  SCP:  ✦ protect-security     (deny disabling CloudTrail/Config/
   │         (org-wide baseline)    GuardDuty/SecurityHub; deny leaving org;
   │                                deny deleting the org CloudTrail/log bucket;
   │                                deny tampering with org-managed IAM roles)
   │
   ├── Security OU      ── (inherits root guardrails; log-archive + audit accts)
   │
   └── Workloads OU
         │  SCP:  ✦ no-IAM-users   (deny iam:CreateUser / CreateAccessKey —
         │                          this org is SSO-only)
         │  SCP:  ✦ require-IMDSv2  (deny ec2:RunInstances unless
         │                          MetadataHttpTokens = required)
         │  SCP:  ✦ s3-public-block (deny s3:PutAccountPublicAccessBlock off,
         │                          deny disabling bucket public-access block)
         │
         └── Prod OU
               │  SCP:  ✦ REGION-LOCK      (deny everything outside eu-west-1,
               │                            except global services: IAM/STS/
               │                            CloudFront/Route53/Organizations)
               │  SCP:  ✦ SERVICE-ALLOWLIST(deny every service EXCEPT the ones
               │                            this workload needs — see below)
               │
               └── ● payments-prod  (the account: S3 + EKS + EC2 only)
                       │
                       ├─ S3 bucket  ◄── bucket policy: grant only the EKS/EC2
                       │                 roles in THIS account + TLS-only.
                       │                 RCP from the root is the backstop.
                       ├─ EKS cluster ── pods get IAM via IRSA / Pod Identity
                       │                 (role scoped to the pod's SA, NOT node)
                       └─ EC2 (nodes) ── instance role least-priv; IMDSv2 forced

The two heavy hitters for blast radius are the Prod-OU SCPs:

(a) Region-lock — if a role is stolen, the attacker can't spin up GPU miners in 30 other regions:

json
// SCP on Prod OU: confine all action to eu-west-1, but DON'T break global services
{
  "Effect": "Deny",
  "NotAction": [
    "iam:*", "sts:*", "organizations:*",
    "cloudfront:*", "route53:*", "support:*", "waf:*"
  ],
  "Resource": "*",
  "Condition": { "StringNotEquals": { "aws:RequestedRegion": "eu-west-1" } }
}

(b) Service-allowlist — deny every service except the handful this workload genuinely uses. This is the biggest blast-radius win: a compromised admin in this account still can't reach SageMaker, Bedrock, or 200 other services:

json
// SCP on Prod OU: only the services an S3 + EKS + EC2 workload actually needs
{
  "Effect": "Deny",
  "NotAction": [
    "s3:*",                                      // the data store
    "ec2:*", "autoscaling:*",                    // nodes + scaling
    "eks:*", "ecr:*",                            // cluster + image pulls
    "elasticloadbalancing:*",                    // ingress
    "kms:*", "logs:*", "cloudwatch:*",           // encryption + telemetry
    "iam:*", "sts:*",                            // roles (still capped by other SCPs)
    "cloudtrail:*"                               // its own trail
  ],
  "Resource": "*"
}

The org-wide RCP is the resource-side backstop — even if someone writes a bucket policy with Principal: "*", this denies any caller outside your org:

json
// RCP on Org Root: data perimeter — no external principal may touch these resources
{
  "Version": "2012-10-17",
  "Statement": [{
    "Sid": "DenyExternalPrincipals",
    "Effect": "Deny",
    "Principal": "*",
    "Action": ["s3:*", "sts:*", "kms:*", "sqs:*", "secretsmanager:*", "ecr:*"],
    "Resource": "*",
    "Condition": {
      "StringNotEqualsIfExists": { "aws:PrincipalOrgID": "o-myorg123" },
      "BoolIfExists": { "aws:PrincipalIsAWSService": "false" }
    }
  }]
}

StringNotEqualsIfExists + aws:PrincipalIsAWSService together mean "deny unless the caller is in my org or is a legitimate AWS service (like CloudFront reading the bucket)" — so you don't accidentally break AWS's own integrations.

Memory hook

blast radius = "where can a stolen role go?" and you shrink it on four axes. Account (isolation boundary), region (deny everywhere but one), service (allowlist only what you run), and external reach (RCP data-perimeter so the data can't leave the org). For the S3+EKS+EC2 account the punchline is: region-lock + service-allowlist SCPs on the Prod OU, a data-perimeter + TLS RCP at the root, IMDSv2 forced, and IRSA so pods don't inherit the node role. That sentence is a complete, senior-level blast-radius answer.


iam:PassRole, Confused Deputy & ExternalId

Two related, frequently-asked concepts:

iam:PassRole — handing a role to a service. When you launch an EC2 with an instance profile, you're "passing" that role to EC2. The danger: if a user can iam:PassRole any role plus ec2:RunInstances, they can launch an instance with an admin role and read its credentials → privilege escalation. So PassRole must be scoped to only the specific roles a user may pass.

   Privesc:  iam:PassRole (on a powerful role)  +  ec2:RunInstances

   low-priv user (holds PassRole + RunInstances, but not admin)
        │ RunInstances( IamInstanceProfile = AdminRole )   ← "passes" AdminRole to EC2
        ▼
   EC2 boots WITH AdminRole attached
        │ attacker reaches the box, then curls the metadata service
        ▼
   169.254.169.254  →  AdminRole temp creds  →  attacker is now admin
   ──────────────────────────────────────────────────────────────────────
   Fix: scope PassRole's "Resource" to ONLY the specific low-priv roles a user
        may pass — never  "Resource": "*"  — and pair it with a PassedToService.

The confused deputy & ExternalId — when you let a third party (e.g. a SaaS vendor) assume a role in your account, an attacker who is another customer of that vendor could trick the vendor into assuming your role. The fix is ExternalId: a shared secret the vendor must present in the AssumeRole call, proving the request is really for your account.

json
// Trust policy requiring an ExternalId (anti-confused-deputy)
{
  "Effect": "Allow",
  "Principal": { "AWS": "arn:aws:iam::VENDOR-ACCOUNT:root" },
  "Action": "sts:AssumeRole",
  "Condition": { "StringEquals": { "sts:ExternalId": "unique-secret-per-customer" } }
}
Memory hook

the "confused deputy" is a deputy tricked into misusing its authority. The vendor (deputy) has the power to assume your role; an attacker tricks it into doing so on the attacker's behalf. ExternalId is a password that ties each assume to a specific customer, so the deputy can't be confused about whose role to assume.


IAM Privilege Escalation Paths

These are the misconfigurations that let a low-privileged identity become admin. Interviewers love "name some IAM privesc paths":

iam:CreatePolicyVersion + iam:SetDefaultPolicyVersion  → rewrite a policy to admin
iam:PassRole + ec2:RunInstances                        → launch EC2 w/ admin role, read IMDS creds
iam:PassRole + lambda:CreateFunction + Invoke          → run code as an admin role
iam:CreateAccessKey (on another user)                  → mint creds for a privileged user
iam:AddUserToGroup                                     → add self to an admin group
iam:AttachUserPolicy / PutUserPolicy                   → attach AdministratorAccess to self
iam:UpdateAssumeRolePolicy                             → rewrite a role's trust policy to trust you
sts:AssumeRole on overly-broad trust                   → assume a more powerful role

ToolsPMapper (builds an IAM privilege-escalation graph — "who can become admin"), Pacu (AWS exploitation framework), Cloudsplaining (flags risky policies). The defensive equivalent of BloodHound for AWS.

Memory hook

most AWS privesc is "permission to grant permissions." Almost every path above is an identity that can modify IAM itself — create keys, attach policies, pass roles, rewrite trust. The lesson: IAM-write permissions (iam:*, PassRole) are as sensitive as admin, because they're a path to admin. Treat the ability to change permissions as the crown jewel it is.


Provisioning Users — the right way

How you should set up human and machine access today:

Humans

  • Use AWS IAM Identity Center (formerly AWS SSO) federated to your IdP (Okta/Entra). Humans log in through SSO and assume roles to get temporary credentials — no IAM users, no long-lived keys.
  • Assign access via permission sets mapped to groups, following least privilege. Start minimal, add as needed.
  • Enforce MFA at the IdP; require it via policy conditions for sensitive actions.

Machines/services

  • Use roles: instance profiles (EC2), execution roles (Lambda), task roles (ECS), IRSA/Pod Identity (EKS). Never bake access keys into code or AMIs.
  • For external CI/CD (e.g. GitHub Actions), use OIDC federation so the pipeline assumes a role with short-lived creds instead of stored keys.

General hygiene

  • Prefer groups over per-user policies; prefer managed over inline policies for auditability.
  • Use permission boundaries to delegate safely.
  • Run IAM Access Analyzer to find external sharing and unused permissions; right-size with last-accessed data.
Memory hook

the modern AWS identity stack is "SSO for humans, roles for machines, keys for nobody." If an interviewer asks "how would you provision access for a new team," that one line plus "least privilege via permission sets/groups, MFA, and IAM Access Analyzer to keep it tidy" is a complete, modern answer.


Incident Response: a Compromised AWS Access Key

This is one of the most common AWS IR interview scenarios. A long-lived access key (AKIA...) leaks — committed to GitHub, found in a breached laptop, exposed via SSRF. What do you do, in order?

1. CONTAIN (fast, reversible) — don't delete yet, you need it for scoping:
   - Attach an explicit DENY-ALL policy to the user, OR deactivate the key
     (aws iam update-access-key --status Inactive). Deactivating is instantly
     reversible and stops the key while preserving it as evidence.
   - If a session/role is involved, also REVOKE active sessions (see below).

2. SCOPE — what did the key do?
   - Pull CloudTrail for that access key ID: every API call, source IP, time,
     user-agent. Look for the tell-tale attacker pattern: a burst of recon
     (List*/Describe*/GetCallerIdentity), then privesc, then resource creation.
   - Did it create new IAM users/keys/roles? Launch EC2 (crypto-mining)? Touch
     S3? Modify CloudTrail/GuardDuty? Assume other roles (follow the chain)?

3. ERADICATE:
   - Delete the compromised key (after scoping). Delete any IAM users, keys,
     or roles the attacker CREATED for persistence — this is the step people
     forget; rotating the one key leaves the attacker's backdoor identities.
   - Roll any secrets the key could read (Secrets Manager, SSM, env).

4. RECOVER & HARDEN:
   - Rotate the legitimate key properly (or better: replace the IAM user with a
     role/SSO so there's no long-lived key to leak again).
   - Add detections: GitHub secret scanning / push protection, GuardDuty
     credential-exfil findings, alerts on new IAM users/keys and on AssumeRole
     from unusual IPs.
Memory hook

"rotate the key" is NOT enoughThe #1 mistake in AWS key-compromise response is rotating the leaked key and closing the ticket — while the attacker's newly created IAM users, access keys, and roles quietly persist. Always scope what the key did and eradicate the persistence it established. And remember AWS access keys don't expire on their own, so a leaked AKIA key is good forever until you act — which is exactly why bots scan GitHub for them within seconds of a commit.

Memory hook

Fun fact — leaked keys are found in seconds. Researchers have shown that an AWS access key committed to a public GitHub repo is often discovered and used by automated bots in under a minute — typically to spin up expensive GPU instances for crypto mining. This is why AWS partnered with GitHub for automatic secret scanning that auto-quarantines exposed keys. The takeaway: there is no "I'll fix it tomorrow" with a leaked key.


Interview Questions

Q
Explain the difference between an IAM user and an IAM role, and why roles are preferred.
Model answer

An IAM user is a permanent identity with long-lived credentials — a password and/or access keys that don't expire until rotated. A role is an identity with permissions but no permanent credentials; principals assume it via STS and receive temporary credentials that expire in minutes to hours. Roles are preferred because temporary credentials dramatically reduce risk — there's no long-lived secret sitting in a file or AMI to be stolen, and a leaked temporary credential expires quickly. Roles also enable clean patterns: EC2/Lambda/EKS get credentials by assuming a role rather than embedding keys, and humans federate through SSO and assume roles. The modern principle is roles over users, temporary over permanent, keys for nobody.

Q
Walk me through IAM policy evaluation. What wins, an allow or a deny?
Model answer

The default is deny — if nothing explicitly allows an action, it's denied. For something to be allowed, it must be granted by an identity-based or resource-based policy and permitted by every ceiling that applies: SCPs, the permission boundary, and any session policy. And the overriding rule is that an explicit Deny anywhere always wins — no number of Allows can override it. For same-account access, either an identity policy or the resource policy granting is enough; for cross-account, you need both the identity policy in the caller's account and the resource policy in the target account. So the two sentences are: default deny, and explicit deny beats allow.

Q
What's a permission boundary and what problem does it solve?
Model answer

A permission boundary is a ceiling on the maximum permissions an IAM user or role can have, regardless of how permissive its attached policies are — effective access is the intersection of the granted policies and the boundary. It solves safe delegation: say you want developers to create their own IAM roles for their services, but giving them iam:* would let them create an admin role and escalate. By requiring that any role they create carries a permission boundary capping it at, say, S3 and DynamoDB, you let them self-serve IAM while guaranteeing they can never create something more powerful than the boundary allows. It's the standard answer to "how do you delegate IAM without enabling privilege escalation."

Q
How does an EC2 instance call AWS APIs without stored credentials, and why is that a security concern?
Model answer

You attach an instance profile, which wraps an IAM role, to the EC2 instance. The instance retrieves temporary credentials for that role from the instance metadata service at 169.254.169.254, and the SDK uses them automatically — no keys are stored. The security concern is that anything able to make the instance fetch that metadata URL can steal those role credentials, which is why SSRF against EC2 is so dangerous — it was the core of the Capital One breach, where SSRF reached the metadata service, grabbed role credentials, and exfiltrated S3 data. The mitigation is IMDSv2, which requires a session token obtained via a PUT request with a hop limit, something SSRF generally can't perform; enforce IMDSv2 and scope the instance role tightly.

Q
A developer's AWS access key was committed to a public GitHub repo. Walk me through your response.
Model answer

First contain reversibly — deactivate the key (set it inactive) rather than deleting it, so it stops working immediately but I preserve it for scoping. Then scope using CloudTrail filtered to that access key ID: every API call, source IP, and timestamp, looking for the classic pattern of recon calls, then privilege escalation, then resource creation. The critical question is what the attacker created for persistence — new IAM users, access keys, roles, or trust-policy changes — plus any EC2 they launched for mining, any S3 they read, and whether they touched CloudTrail or GuardDuty. Then eradicate: delete the leaked key and, crucially, remove the backdoor identities they created, because rotating the one key while leaving attacker-created users is the most common mistake. Roll any secrets the key could read. Finally harden — replace the IAM user with SSO or a role so there's no long-lived key to leak again, and add push-protection and GuardDuty detections. And I'd note these keys are found by bots in under a minute, so speed matters.

Q
What is iam:PassRole and how is it abused?
Model answer

PassRole is the permission to hand an IAM role to an AWS service — for example, specifying an instance profile when launching EC2 passes that role to the EC2 service. It's abused for privilege escalation: if a low-privileged user can pass any role and also run a compute service, they can launch an EC2 instance, Lambda, or similar with a powerful admin role attached and then read that role's credentials or run code as it. So PassRole effectively lets you borrow the permissions of any role you can pass. The mitigation is to scope PassRole tightly with a Resource condition listing only the specific roles a principal may pass, and to treat PassRole as a sensitive, admin-adjacent permission.

Q
What are SCPs and how do they differ from IAM policies?
Model answer

Service Control Policies are organization-level guardrails applied to an OU or account that set the maximum permissions for everything in that account — they're a ceiling, not a grant. The key difference from identity policies is that SCPs never grant anything; an action is only allowed if both an IAM policy grants it and no SCP denies it. They're used for preventive controls that even account admins can't override — like denying the ability to disable CloudTrail or GuardDuty, locking actions to approved regions, or blocking leaving the org. A powerful property is that an SCP denying cloudtrail:StopLogging means even a fully-compromised account admin can't turn off logging. The caveat is they don't apply to the management account's root.

Q
How would you provision AWS access for a new engineering team?
Model answer

For the humans, AWS IAM Identity Center federated to our IdP — they log in through SSO and assume roles for temporary credentials, with no IAM users and no long-lived keys. Access is granted via permission sets mapped to groups, scoped least-privilege, starting minimal and expanding as needed, with MFA enforced at the IdP. For their workloads, roles everywhere — instance profiles, Lambda execution roles, IRSA or Pod Identity for EKS — and OIDC federation for their CI/CD so pipelines assume short-lived roles instead of storing keys. I'd delegate IAM self-service safely with permission boundaries, prefer groups and managed policies for auditability, and run IAM Access Analyzer plus last-accessed data to catch external sharing and prune unused permissions. The summary: SSO for humans, roles for machines, keys for nobody.

Q
Why split an AWS environment into many accounts instead of separating teams with IAM in one account?
Model answer

Because the account is the strongest isolation boundary AWS gives you — stronger than a VPC or an IAM role — so it's the natural cap on a blast radius. In one big account, isolation depends on every IAM policy being perfect, and a single misconfigured Resource: "*" can leak access across teams; across an account boundary, access requires an explicit, auditable cross-account trust. The pattern is one account per workload per environment — payments-prod, payments-dev, and so on — plus shared accounts for logging and security. So a compromised role in dev simply has no path to prod resources, because they're different accounts with different credentials. Isolation becomes the default and sharing the deliberate exception.

Q
What's the difference between an SCP and an RCP, and why did AWS add RCPs?
Model answer

Both are organization-level ceilings that only restrict, never grant, but they govern opposite sides of a request: a Service Control Policy caps what the principals — the identities — inside your accounts can do, while a Resource Control Policy, launched in November 2024, caps who and how anyone can access the resources in your accounts. AWS added RCPs because SCPs only govern your own principals, so they can't stop a misconfigured bucket or key policy from sharing a resource with an external account or the public internet. An RCP sits on the resource side and lets you enforce centrally — for example, deny any principal that isn't in my org from touching S3, STS, KMS, SQS, or Secrets Manager — as a backstop that a too-generous resource policy can't override. Together they form a data perimeter: SCP for "what my people can do," RCP for "who can touch my things."

Q
You have a production account running only S3, EKS, and EC2 in one region. How would you set up guardrails to limit the blast radius if a role is compromised?
Model answer

I'd shrink the blast radius on four axes. The account itself is the isolation boundary, so this workload is already its own account under a Prod OU. On that OU I'd put two SCPs: a region-lock that denies every action outside the one region — with a carve-out for global services like IAM, STS, CloudFront, and Route 53 — so a stolen role can't spin up miners in thirty other regions, and a service-allowlist that denies every service except the handful this workload needs, like S3, EC2, EKS, ECR, ELB, KMS, and logging. At the org root I'd attach a data-perimeter RCP that denies any principal outside my org from touching S3, STS, KMS, and the other supported resources, plus a TLS-only RCP, so even a fat-fingered bucket policy can't leak data externally. Then IMDSv2 forced on the instances and IRSA or Pod Identity so pods get a scoped role instead of inheriting the node's. The one-liner: region-lock plus service-allowlist SCPs, a data-perimeter RCP, IMDSv2, and pod-scoped IAM.

Q
How should you decide between an identity-based policy and a resource-based policy, and how do RCPs change the recommendation?
Model answer

Default to identity-based policies and reserve resource-based policies for the things only they can do — namely cross-account access and letting an AWS service in, like CloudTrail writing to your log bucket — because splitting access logic across both sides gets unauditable fast. When you do write a resource policy, never use a wildcard principal without a condition that scopes it, pin cross-account trust to aws:PrincipalOrgID rather than bare account IDs, and require TLS. What RCPs change is the backstop: instead of hoping every one of hundreds of bucket policies is perfect, you enforce the data perimeter once with an org-level RCP that denies external principals by default, and let individual resource policies grant access only inside that fence. The resource policy is the lock on the door; the RCP is the building's master rule that a single bad door can't override.