AWS IAM — The Deep Dive
IAM (Identity and Access Management) is the single most important topic in any AWS security interview. It's the authorization brain of AWS — every API call, from reading an S3 object to launching a server, is evaluated by IAM. This file explains IAM in plain language first, then goes deep on the parts interviewers always probe: roles vs users, the policy types and how they're evaluated, permission boundaries, how compute gets credentials, how to provision users properly, and exactly what to do when an access key leaks.
Last verified2026-06 See also: AWS fundamentals (product primer) · AWS incident response
What IAM Actually Is
Every request to AWS — clicking in the console, an aws s3 cp, a Lambda reading a secret — is an API call. Before AWS performs it, IAM answers one question: "is this principal allowed to perform this action on this resource, under these conditions?"
That's the whole job. IAM is a giant policy evaluation engine sitting in front of every AWS service. Master "who is asking, what are they asking to do, to which resource, and which policies apply," and you've understood IAM.
Memory hookevery AWS action is
Principal+Action+Resource+Condition. Those four words are the shape of every IAM decision and every IAM policy statement. When debugging "why is this denied?", walk those four: who is the principal, what action, which resource, and do the conditions (IP, MFA, region, tags) match?
Identity vs Policy — the two things people conflate
The most common IAM confusion, and worth nailing before anything else: an identity and a policy are completely different objects.
- An identity is who you are — a badge. IAM users, groups, and roles are identities. AWS authenticates them, and they can hold or assume credentials.
- A policy is the rules — a JSON document full of
Allow/Denystatements. It has no credentials and authenticates nobody; it does nothing on its own. It's inert paper until you attach it to an identity (or to a resource).
The badge opens no doors by itself; the rulebook is just paper until it's pinned to a badge. Access = an identity with a policy attached. Hand the same person a different rulebook and they reach different rooms — the badge never changed, the rules did.
IDENTITY — the "who" (a badge) POLICY — the "rules" (a rulebook)
┌──────────────────────────┐ ┌───────────────────────────────────┐
│ IAM user "alice" │ attach │ { │
│ IAM group "devs" │◄─────────►│ "Effect": "Allow", │
│ IAM role "deploy" │ │ "Action": "s3:GetObject", │
│ │ │ "Resource": "arn:…:bucket/*" │
│ • AWS authenticates it │ │ } │
│ • holds / assumes creds │ │ · pure JSON — no creds, no powers │
└──────────────────────────┘ │ · does NOTHING until attached │
▲ └───────────────────────────────────┘
└─ one policy can attach to many identities (and an identity to many policies)
an identity with NO policy attached can do NOTHING (default deny)What a policy actually looks like — JSON with one or more statements, each one answering IAM's four questions:
{
"Effect": "Allow", ← Allow or Deny
"Action": "s3:GetObject", ← WHAT operation
"Resource": "arn:aws:s3:::bucket/*", ← on WHICH resource
"Condition": { "Bool": { "aws:MultiFactorAuthPresent": "true" } } ← WHEN (optional)
}Notice there's no Principal field here — that's the giveaway of an identity-based
policy: the identity it's attached to is the "who". A Principal field only appears
in resource-based policies (a bucket/key/role-trust policy), which name who may
touch the resource.
Principal vs identity (another easy mix-up): principal is the broad term — anything that can make a request: an identity, an AWS service like
ec2.amazonaws.com, or even anonymous ("*"). Identity is the narrower kind of principal you create and attach policies to — users, groups, roles. (A group isn't something you log in as; it's an identity you attach policies to, and its users inherit them.)
Memory hookidentity ≠ policy"Who you are" and "what's allowed" are separate objects you wire together by attaching. Identities authenticate but carry no permissions of their own; policies carry permissions but authenticate nobody. A user with no policy can do nothing; a policy attached to nothing affects nothing.
The next two sections take each in turn — Principals (the identities) and, further down, Policies (the rules).
Principals — the "who"
A principal is anything that can make an AWS request. The types, and how to think about each:
| Principal | What it is | Lifespan of creds | Use it for |
|---|---|---|---|
| Root user | The original account owner, all-powerful, tied to the account's email | Permanent | Almost nothing — lock it away (see below) |
| IAM user | A named identity with long-lived credentials (password and/or access keys) | Permanent until rotated | Legacy humans/services — increasingly discouraged |
| IAM role | An identity with no permanent credentials that principals assume to get temporary credentials | Temporary (minutes–hours) | Almost everything — humans via SSO, EC2, Lambda, cross-account |
| Federated identity | A user from an external IdP (Okta, Entra, Google) mapped to a role | Temporary | Human workforce access (the modern way) |
| AWS service principal | An AWS service itself (e.g. ec2.amazonaws.com) acting on your behalf | n/a | Letting a service assume a role |
Memory hookthe whole modern IAM philosophy: roles over users, temporary over permanent. IAM users have long-lived access keys that sit in files waiting to be stolen. IAM roles hand out credentials that expire in an hour. The single biggest IAM improvement most orgs can make is "stop creating IAM users with access keys; use roles and federation." If you remember one design principle, it's this one.
The Root User — handle with extreme care
The root user is the account's god-mode identity, created with the account and tied to its email address. It can do everything, including things no IAM policy can restrict (close the account, change billing, certain S3/KMS recovery actions). An SCP can't even fully restrict the root of the management account.
Correct handling of root
Create IAM roles/identities for everything operational.
on it and store the credentials in a safe/break-glass vault.
root should never have programmatic access keys. (AWS now blocks creating them.)
a CloudTrail/GuardDuty alert on root login or root API calls is a top-priority signal, because legitimate root use is rare and attacker root use is catastrophic.
Memory hookFun fact — root is the one identity guardrails can't fully cage. Service Control Policies, which can restrict every other identity in an org, do not restrict the management account's root. That's why "secure the root, MFA it, and alarm on its use" is AWS security 101 — it's the one key that opens every lock.
Roles & AssumeRole — the heart of IAM
A role is the most important and most misunderstood IAM concept. A role is an identity with permissions but no permanent credentials. To use it, a principal assumes it via sts:AssumeRole, receiving temporary credentials (an access key + secret + session token) that expire.
Every role has two policies, and confusing them is the classic mistake:
┌─────────────────────── ROLE ───────────────────────┐
│ │
TRUST POLICY (who can assume me?) PERMISSION POLICY (what can I do?)
= a resource-based policy on the role = identity-based policy
"Principal X is allowed to AssumeRole" "Allow s3:GetObject on bucket Y"who is allowed to become this role. (e.g. "the EC2 service" or "account 123's CI role" or "users from our Okta").
what the role can do once assumed.
// Trust policy: allow EC2 instances to assume this role
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Principal": { "Service": "ec2.amazonaws.com" },
"Action": "sts:AssumeRole"
}]
}Memory hooka role is a hat, not a person. Nobody is the role; principals put on the hat (assume it) temporarily, do the work with its powers, and take it off (the creds expire). The trust policy is the bouncer deciding who's allowed to wear the hat; the permission policy is what the hat lets you do. "AssumeRole = put on the hat and get temporary keys."
AssumeRole, step by step — "alice" assuming the "deploy" role. Note the two
gates that must both pass:
alice (user) STS IAM (checks 2 gates)
│ AssumeRole(deploy) │ │
├───────────────────────►│ gate 1: does alice's policy │
│ │ allow sts:AssumeRole?├─┐
│ │ gate 2: does deploy's TRUST │ │
│ │ policy allow alice? │◄┘
│ temp creds │ both yes → mint creds for │
│ (ASIA… + session │ the deploy role │
│ token + expiry) │ │
│◄───────────────────────┤ │
│
▼ every later call runs as assumed-role/deploy/alice ,
and the DEPLOY role's PERMISSION policy decides what's allowed.
(CloudTrail logs the AssumeRole and each subsequent call.)Miss either gate — alice lacks sts:AssumeRole, or the role's trust policy doesn't
name her — and you get AccessDenied. The trust policy lives on the role (it's a
resource-based policy); alice's sts:AssumeRole lives on her identity.
Reading a credential — AKIA vs ASIA
Every set of AWS credentials starts with an access key ID, and its four-letter prefix tells you at a glance what kind it is — useful in CloudTrail and when triaging a leaked credential:
AKIA EXAMPLE... ASIA EXAMPLE...
+ secret access key + secret access key
+ SESSION TOKEN ◄── the extra piece
= 2 parts, NEVER expires = 3 parts, EXPIRES (minutes–hours)
a long-lived IAM USER key a temporary STS credential
(the kind that leaks to GitHub) (from AssumeRole / IRSA / federation)AKIA…a long-lived key belonging to an IAM user. No expiry; valid until someone rotates or deletes it. This is the kind bots scan GitHub for.
ASIA…a temporary credential minted by STS (every AssumeRole, IRSA,
or federated login). It requires a session token alongside the key+secret, and it
expires on its own. Seeing ASIA means "a role session — find which role and who
assumed it; the creds will die by themselves."
Memory hookTwo more prefixes show up in the
userIdentityblock of CloudTrail (these are unique-ID prefixes, not access keys):AROA…= a role object,AIDA…= a user object. Soassumed-role/deploy/...with anAROA…principalId is a role session; anAIDA…is a plain IAM user. The quick read:AKIA= permanent user key (rotate it),ASIA= temporary role session (it expires).
How Compute Gets Credentials (no hardcoded keys!)
A core interview theme: how does an EC2 instance / Lambda / EKS pod call AWS without anyone hardcoding an access key? Answer: they assume a role and AWS injects temporary credentials.
| Compute | Mechanism | How creds arrive |
|---|---|---|
| EC2 | Instance profile (a wrapper around a role) | The instance fetches temp creds from the instance metadata service (IMDS) at 169.254.169.254 |
| Lambda | Execution role | AWS injects temp creds into the function's environment |
| ECS/Fargate | Task role | Creds served from a task-local metadata endpoint |
| EKS pod | IRSA (IAM Roles for Service Accounts) or EKS Pod Identity | The pod's Kubernetes service account is mapped to an IAM role; pod gets a projected token → assumes the role via OIDC |
Memory hookthe metadata endpoint is the magic and the danger. EC2 gets its credentials from
http://169.254.169.254. That's elegant (no stored keys) but it's exactly why SSRF against EC2 is so dangerous: trick the app into fetching that URL and you steal the instance's role credentials. The Capital One breach (2019) was precisely this — SSRF → IMDS → role creds → S3 exfiltration of 100M+ records. The fix is IMDSv2, which requires a session token obtained via aPUTthat SSRF typically can't perform. Always: "EC2 creds come from IMDS; protect IMDS with v2."
Memory hookEKS: IRSA vs Pod IdentityBoth solve "give a specific pod (not the whole node) its own IAM role." IRSA maps a Kubernetes service account to an IAM role via an OIDC trust relationship (older, more setup). EKS Pod Identity (newer, 2023) does the same with a simpler agent-based association. The security point either way: scope IAM to the pod's service account, never to the node's instance role — otherwise every pod on the node inherits the node's permissions (a lateral-movement goldmine).
Policies — the rules
A policy is a JSON document of Allow/Deny statements. There are several types, distinguished by what they attach to and whether they can grant or only restrict.
The policy types (readable version)
| Type | Attaches to | Grants or only restricts? | Think of it as |
|---|---|---|---|
| Identity-based | user, group, role | Grants | "What this identity can do" |
| Resource-based | S3 bucket, KMS key, SQS, Lambda, role trust policy | Grants (incl. cross-account) | "Who can touch this resource" |
| SCP (Service Control Policy) | Org OU/account | Only restricts (a ceiling) | "Org-wide guardrail" |
| Permission boundary | user or role | Only restricts (a ceiling) | "Max this identity could ever have" |
| Session policy | passed at AssumeRole time | Only restricts | "Shrink this one session" |
Memory hookgrant vs. ceilingOnly identity and resource policies grant access. SCPs, permission boundaries, and session policies are ceilings — they can only take away, never add. So if something's denied, check the ceilings; if something needs to be allowed, it must be granted by an identity or resource policy and permitted by every ceiling above it.
How a request is evaluated (the logic interviewers love)
A request is ALLOWED only if ALL of these hold:
1. No EXPLICIT DENY anywhere (an explicit Deny ALWAYS wins, full stop)
2. Permitted by every SCP (org guardrail)
3. Permitted by the permission boundary (if one is set)
4. Permitted by the session policy (if assumed with one)
5. GRANTED by an identity policy OR a resource policy
Same-account: identity policy OR resource policy granting is enough.
Cross-account: you need BOTH — the identity policy in the caller's account
AND the resource policy in the target account.Memory hook"explicit deny always wins; default is deny." Two rules cover 90% of IAM evaluation questions: (1) the default is deny — if nothing explicitly allows it, it's denied; (2) an explicit
Denybeats anyAllow— no matter how many Allows exist. So SCPs and boundaries enforce limits by denying, and you can't accidentally out-allow a Deny. Say it back in interviews verbatim: "default deny, explicit deny overrides allow."
Permission Boundaries — safe delegation
A permission boundary is a ceiling on what an IAM user or role can do, regardless of how permissive its attached policies are. Its killer use case is delegating IAM safely.
The problem it solvesyou want developers to create their own IAM roles (for their Lambdas, etc.), but if you give them iam:*, a developer could create an admin role and assume it — instant privilege escalation. The fix: require that any role they create has a permission boundary attached, capping it at, say, "S3 + DynamoDB only." Now even a maliciously-crafted role can't exceed the boundary.
Developer's effective permissions = (their identity policy) ∩ (their permission boundary)
= the INTERSECTION, never more than the boundaryMemory hookboundary = "you can grant, but not beyond this line." A permission boundary lets you hand someone the power to create permissions while guaranteeing they can never create more than you allowed. It's the "intersection, not union" idea: effective access is the overlap of what's granted and what the boundary permits. This is the answer to "how do you let developers self-serve IAM without letting them escalate?"
Service Control Policies (SCPs) — org-wide guardrails
SCPs apply at the AWS Organizations level (an OU or account) and set the maximum permissions for everything in that account — they're a ceiling that even account admins can't exceed.
Common SCP uses (preventive controls):
so an attacker who gets admin still can't blind your logging.
deny all actions outside approved regions (limits where an attacker can spin up crypto-mining).
in member accounts.
.
// SCP: prevent anyone (even account admins) from disabling CloudTrail
{
"Effect": "Deny",
"Action": ["cloudtrail:StopLogging", "cloudtrail:DeleteTrail"],
"Resource": "*"
}Memory hookSCPs are the guardrail, not the road. They never grant anything — an empty SCP allows nothing extra. They define the outer fence inside which account-level IAM operates. The classic SCP win: even if an attacker becomes account admin, an SCP that denies
cloudtrail:StopLoggingmeans they can't turn off the cameras — your evidence keeps flowing.
AWS Organizations, Multi-Account & Blast-Radius Guardrails (SCPs + RCPs)
Last verified2026-06
This is the section interviewers use to separate "I clicked around one AWS account" from "I've designed a production landing zone." It ties together three things: how you set an organisation up from scratch, how you split it into accounts, and which guardrails (SCPs + RCPs) you pin where to contain a breach. We'll end with a worked example for a real workload — an account that runs only S3, EKS, and EC2 — and a diagram of exactly what goes where.
What is an AWS Organization?
An AWS Organization is a way to manage many AWS accounts as one tree. At the top is a single management account (the one that creates the org — historically called the "payer" account because all the bills roll up to it). Under it you build a tree of Organizational Units (OUs) — folders — and drop member accounts into those folders.
AWS ORGANIZATION (one tree, billing rolls up to the top)
┌─────────────────────────────────────────────────────┐
│ MANAGEMENT ACCOUNT (root of the org) │
│ • owns the org, billing, SCPs/RCPs │
│ • runs almost NO workloads (keep it nearly empty) │
└───────────────────────────┬──────────────────────────┘
│
┌──────────┬─────────────────┼──────────────┬───────────┐
OU OU OU OU OU
Security Infrastructure Workloads Sandbox Suspended
│ │ ┌───┴────┐ │ │
┌───┴───┐ ┌───┴───┐ Prod-OU Nonprod-OU … (quarantine)
Log- Audit Network Shared │ │
Archive /Sec (VPCs) Services prod dev/stage
acct acct acct acctsWhy split into many accounts at all? An AWS account is the strongest isolation boundary AWS gives you — stronger than a VPC, stronger than an IAM role. A blast radius is naturally capped at the account: a compromised role in dev can't touch prod resources because they live in a different account with different credentials. The pattern is "one account per workload per environment" — e.g. payments-prod, payments-dev, analytics-prod — plus a few shared accounts.
Memory hookthe account is the blast-radius boundary"Why not just run everything in one account with IAM separating teams?" Because IAM is one misconfigured
Resource: "*"away from leaking across teams, whereas crossing an account boundary requires an explicit cross-account trust you can see and audit. Separate accounts make isolation the default and sharing the deliberate exception. This is the single most important multi-account talking point.
Setting it up from scratch — the landing zone
You almost never hand-build this. The from-scratch sequence, in order:
1. Start from a clean account → it becomes the MANAGEMENT account.
Enable AWS Organizations (creates the org root).
2. Turn on AWS CONTROL TOWER (or Landing Zone Accelerator).
This automates steps 3–6 below into a "landing zone".
3. It creates the foundational accounts:
• Log Archive → central, write-once CloudTrail/Config log bucket
• Audit/Security → GuardDuty, Security Hub, cross-account read roles
4. It creates a baseline OU layout (Security OU, Sandbox OU, …)
and enrolls accounts into it.
5. It wires IAM IDENTITY CENTER (SSO) → humans log in centrally and
assume roles into each account. NO IAM users anywhere.
6. It applies a baseline set of SCPs (and you add RCPs) as guardrails.
7. You vend NEW accounts on demand via Account Factory — each one is
born already enrolled, logged, and guard-railed.Memory hook"Organizations is the plumbing, Control Tower is the house." Organizations gives you the raw tree, accounts, and policy attachment points. Control Tower is the opinionated layer on top that stamps out a secure-by-default landing zone — the log-archive account, the SSO wiring, the baseline guardrails — so every new account starts compliant instead of being hardened by hand. In an interview: "I wouldn't hand-roll the org; I'd use Control Tower / Landing Zone Accelerator so logging, SSO, and guardrails are baked in from account #1."
Two kinds of guardrail: SCP (the principals) vs RCP (the resources)
Until late 2024 the org only had one guardrail type — the SCP. SCPs cap what the identities in your accounts can do. But they say nothing about who, from outside, can reach your resources. Resource Control Policies (RCPs) — launched November 2024 — fill that gap: they cap what can be done to the resources in your accounts, no matter who is asking. Together they form a data perimeter — the two halves of one fence.
┌──────────── YOUR ACCOUNT ────────────┐
│ │
SCP ──┤ the IDENTITIES here │── RCP
caps │ (roles, users) ───► can act on ───► │ caps
what │ RESOURCES │ who/how
THEY │ (S3, KMS, STS…) │ can touch
can do │ │ the RESOURCE
└────────────────────────────────────────┘
SCP answers: "what may the principals in my accounts do?" (identity side)
RCP answers: "who may touch the resources in my accounts?" (resource side)| SCP (Service Control Policy) | RCP (Resource Control Policy) | |
|---|---|---|
| Caps the permissions of | Principals (identities) in the account | Resources in the account |
| Classic question | "Can my roles call X?" | "Can anyone — even an external account — read my bucket?" |
| Attaches to | Root / OU / account | Root / OU / account |
| Grants anything? | No — ceiling only | No — ceiling only |
| Applies to root user? | No (mgmt acct root exempt) | Yes — RCPs also bound the resource owner's root |
| Services covered | All services | S3, STS, KMS, SQS, Secrets Manager (Nov 2024); + ECR, OpenSearch Serverless (Jun 2025) |
| Default attached policy | FullAWSAccess | RCPFullAWSAccess |
Memory hookSCP = "what my people can do", RCP = "who can touch my things". The reason RCPs matter: an SCP can't stop a misconfigured resource policy from sharing your S3 bucket with the whole internet, because the SCP only governs your principals, not the external one reading the bucket. An RCP sits on the resource side and can enforce "no principal outside my org may access this bucket — full stop", overriding any too-generous bucket policy. RCPs are how you enforce a data perimeter centrally instead of auditing thousands of bucket policies.
Resource policies — the best-practice recommendation
A resource-based policy (a bucket policy, KMS key policy, role trust policy, SQS/Secrets policy) is the Principal-bearing policy that lives on the resource and says who may touch it. The recommendations interviewers want to hear:
namely cross-account access and letting an AWS service in (e.g. CloudTrail writing to your log bucket). Don't scatter access logic across both sides if one will do; it gets unauditable fast.
"Principal": "*" without a Condition that scopes it.A wildcard principal with no condition is how buckets end up public. If you need broad access, fence it with aws:PrincipalOrgID, aws:SourceArn, aws:SourceAccount, or a VPC endpoint condition.
, not to a bare account: aws:PrincipalOrgID is more durable than listing account IDs.
(aws:SecureTransport: true) and, on S3, enforce Block Public Access at the account level so no bucket policy can accidentally go public.
, so a single fat-fingered bucket policy in one of 200 accounts can't punch a hole — the RCP is the backstop the per-resource policy can't override.
Memory hookresource policy is the lock on the door; the RCP is the building's master rule. Per-resource policies are easy to get subtly wrong at scale. The modern answer to "how do you keep 500 buckets from leaking?" isn't "review every bucket policy" — it's "enforce a data-perimeter RCP at the org so external principals are denied by default, and let bucket policies grant only inside that fence."
Worked example — an account running only S3, EKS, EC2
Now the concrete ask: a production workload account that runs S3 in one region (say eu-west-1), EKS, and EC2 — nothing else. Here is exactly which guardrails to apply and where, to keep the blast radius tiny if a role in that account is compromised.
Where the policies attach (the graph)
ORG ROOT
│ RCP: ✦ data-perimeter (deny any principal NOT in my org from
│ (org-wide backstop) touching S3/STS/KMS/SQS/Secrets/ECR)
│ RCP: ✦ enforce-TLS (deny non-HTTPS on those resources)
│ SCP: ✦ protect-security (deny disabling CloudTrail/Config/
│ (org-wide baseline) GuardDuty/SecurityHub; deny leaving org;
│ deny deleting the org CloudTrail/log bucket;
│ deny tampering with org-managed IAM roles)
│
├── Security OU ── (inherits root guardrails; log-archive + audit accts)
│
└── Workloads OU
│ SCP: ✦ no-IAM-users (deny iam:CreateUser / CreateAccessKey —
│ this org is SSO-only)
│ SCP: ✦ require-IMDSv2 (deny ec2:RunInstances unless
│ MetadataHttpTokens = required)
│ SCP: ✦ s3-public-block (deny s3:PutAccountPublicAccessBlock off,
│ deny disabling bucket public-access block)
│
└── Prod OU
│ SCP: ✦ REGION-LOCK (deny everything outside eu-west-1,
│ except global services: IAM/STS/
│ CloudFront/Route53/Organizations)
│ SCP: ✦ SERVICE-ALLOWLIST(deny every service EXCEPT the ones
│ this workload needs — see below)
│
└── ● payments-prod (the account: S3 + EKS + EC2 only)
│
├─ S3 bucket ◄── bucket policy: grant only the EKS/EC2
│ roles in THIS account + TLS-only.
│ RCP from the root is the backstop.
├─ EKS cluster ── pods get IAM via IRSA / Pod Identity
│ (role scoped to the pod's SA, NOT node)
└─ EC2 (nodes) ── instance role least-priv; IMDSv2 forcedThe two heavy hitters for blast radius are the Prod-OU SCPs:
(a) Region-lock — if a role is stolen, the attacker can't spin up GPU miners in 30 other regions:
// SCP on Prod OU: confine all action to eu-west-1, but DON'T break global services
{
"Effect": "Deny",
"NotAction": [
"iam:*", "sts:*", "organizations:*",
"cloudfront:*", "route53:*", "support:*", "waf:*"
],
"Resource": "*",
"Condition": { "StringNotEquals": { "aws:RequestedRegion": "eu-west-1" } }
}(b) Service-allowlist — deny every service except the handful this workload genuinely uses. This is the biggest blast-radius win: a compromised admin in this account still can't reach SageMaker, Bedrock, or 200 other services:
// SCP on Prod OU: only the services an S3 + EKS + EC2 workload actually needs
{
"Effect": "Deny",
"NotAction": [
"s3:*", // the data store
"ec2:*", "autoscaling:*", // nodes + scaling
"eks:*", "ecr:*", // cluster + image pulls
"elasticloadbalancing:*", // ingress
"kms:*", "logs:*", "cloudwatch:*", // encryption + telemetry
"iam:*", "sts:*", // roles (still capped by other SCPs)
"cloudtrail:*" // its own trail
],
"Resource": "*"
}The org-wide RCP is the resource-side backstop — even if someone writes a bucket policy with Principal: "*", this denies any caller outside your org:
// RCP on Org Root: data perimeter — no external principal may touch these resources
{
"Version": "2012-10-17",
"Statement": [{
"Sid": "DenyExternalPrincipals",
"Effect": "Deny",
"Principal": "*",
"Action": ["s3:*", "sts:*", "kms:*", "sqs:*", "secretsmanager:*", "ecr:*"],
"Resource": "*",
"Condition": {
"StringNotEqualsIfExists": { "aws:PrincipalOrgID": "o-myorg123" },
"BoolIfExists": { "aws:PrincipalIsAWSService": "false" }
}
}]
}StringNotEqualsIfExists + aws:PrincipalIsAWSService together mean "deny unless the caller is in my org or is a legitimate AWS service (like CloudFront reading the bucket)" — so you don't accidentally break AWS's own integrations.
Memory hookblast radius = "where can a stolen role go?" and you shrink it on four axes. Account (isolation boundary), region (deny everywhere but one), service (allowlist only what you run), and external reach (RCP data-perimeter so the data can't leave the org). For the S3+EKS+EC2 account the punchline is: region-lock + service-allowlist SCPs on the Prod OU, a data-perimeter + TLS RCP at the root, IMDSv2 forced, and IRSA so pods don't inherit the node role. That sentence is a complete, senior-level blast-radius answer.
iam:PassRole, Confused Deputy & ExternalId
Two related, frequently-asked concepts:
iam:PassRole — handing a role to a service. When you launch an EC2 with an instance profile, you're "passing" that role to EC2. The danger: if a user can iam:PassRole any role plus ec2:RunInstances, they can launch an instance with an admin role and read its credentials → privilege escalation. So PassRole must be scoped to only the specific roles a user may pass.
Privesc: iam:PassRole (on a powerful role) + ec2:RunInstances
low-priv user (holds PassRole + RunInstances, but not admin)
│ RunInstances( IamInstanceProfile = AdminRole ) ← "passes" AdminRole to EC2
▼
EC2 boots WITH AdminRole attached
│ attacker reaches the box, then curls the metadata service
▼
169.254.169.254 → AdminRole temp creds → attacker is now admin
──────────────────────────────────────────────────────────────────────
Fix: scope PassRole's "Resource" to ONLY the specific low-priv roles a user
may pass — never "Resource": "*" — and pair it with a PassedToService.The confused deputy & ExternalId — when you let a third party (e.g. a SaaS vendor) assume a role in your account, an attacker who is another customer of that vendor could trick the vendor into assuming your role. The fix is ExternalId: a shared secret the vendor must present in the AssumeRole call, proving the request is really for your account.
// Trust policy requiring an ExternalId (anti-confused-deputy)
{
"Effect": "Allow",
"Principal": { "AWS": "arn:aws:iam::VENDOR-ACCOUNT:root" },
"Action": "sts:AssumeRole",
"Condition": { "StringEquals": { "sts:ExternalId": "unique-secret-per-customer" } }
}Memory hookthe "confused deputy" is a deputy tricked into misusing its authority. The vendor (deputy) has the power to assume your role; an attacker tricks it into doing so on the attacker's behalf.
ExternalIdis a password that ties each assume to a specific customer, so the deputy can't be confused about whose role to assume.
IAM Privilege Escalation Paths
These are the misconfigurations that let a low-privileged identity become admin. Interviewers love "name some IAM privesc paths":
iam:CreatePolicyVersion + iam:SetDefaultPolicyVersion → rewrite a policy to admin
iam:PassRole + ec2:RunInstances → launch EC2 w/ admin role, read IMDS creds
iam:PassRole + lambda:CreateFunction + Invoke → run code as an admin role
iam:CreateAccessKey (on another user) → mint creds for a privileged user
iam:AddUserToGroup → add self to an admin group
iam:AttachUserPolicy / PutUserPolicy → attach AdministratorAccess to self
iam:UpdateAssumeRolePolicy → rewrite a role's trust policy to trust you
sts:AssumeRole on overly-broad trust → assume a more powerful roleToolsPMapper (builds an IAM privilege-escalation graph — "who can become admin"), Pacu (AWS exploitation framework), Cloudsplaining (flags risky policies). The defensive equivalent of BloodHound for AWS.
Memory hookmost AWS privesc is "permission to grant permissions." Almost every path above is an identity that can modify IAM itself — create keys, attach policies, pass roles, rewrite trust. The lesson: IAM-write permissions (
iam:*,PassRole) are as sensitive as admin, because they're a path to admin. Treat the ability to change permissions as the crown jewel it is.
Provisioning Users — the right way
How you should set up human and machine access today:
Humans
- Use AWS IAM Identity Center (formerly AWS SSO) federated to your IdP (Okta/Entra). Humans log in through SSO and assume roles to get temporary credentials — no IAM users, no long-lived keys.
- Assign access via permission sets mapped to groups, following least privilege. Start minimal, add as needed.
- Enforce MFA at the IdP; require it via policy conditions for sensitive actions.
Machines/services
- Use roles: instance profiles (EC2), execution roles (Lambda), task roles (ECS), IRSA/Pod Identity (EKS). Never bake access keys into code or AMIs.
- For external CI/CD (e.g. GitHub Actions), use OIDC federation so the pipeline assumes a role with short-lived creds instead of stored keys.
General hygiene
- Prefer groups over per-user policies; prefer managed over inline policies for auditability.
- Use permission boundaries to delegate safely.
- Run IAM Access Analyzer to find external sharing and unused permissions; right-size with last-accessed data.
Memory hookthe modern AWS identity stack is "SSO for humans, roles for machines, keys for nobody." If an interviewer asks "how would you provision access for a new team," that one line plus "least privilege via permission sets/groups, MFA, and IAM Access Analyzer to keep it tidy" is a complete, modern answer.
Incident Response: a Compromised AWS Access Key
This is one of the most common AWS IR interview scenarios. A long-lived access key (AKIA...) leaks — committed to GitHub, found in a breached laptop, exposed via SSRF. What do you do, in order?
1. CONTAIN (fast, reversible) — don't delete yet, you need it for scoping:
- Attach an explicit DENY-ALL policy to the user, OR deactivate the key
(aws iam update-access-key --status Inactive). Deactivating is instantly
reversible and stops the key while preserving it as evidence.
- If a session/role is involved, also REVOKE active sessions (see below).
2. SCOPE — what did the key do?
- Pull CloudTrail for that access key ID: every API call, source IP, time,
user-agent. Look for the tell-tale attacker pattern: a burst of recon
(List*/Describe*/GetCallerIdentity), then privesc, then resource creation.
- Did it create new IAM users/keys/roles? Launch EC2 (crypto-mining)? Touch
S3? Modify CloudTrail/GuardDuty? Assume other roles (follow the chain)?
3. ERADICATE:
- Delete the compromised key (after scoping). Delete any IAM users, keys,
or roles the attacker CREATED for persistence — this is the step people
forget; rotating the one key leaves the attacker's backdoor identities.
- Roll any secrets the key could read (Secrets Manager, SSM, env).
4. RECOVER & HARDEN:
- Rotate the legitimate key properly (or better: replace the IAM user with a
role/SSO so there's no long-lived key to leak again).
- Add detections: GitHub secret scanning / push protection, GuardDuty
credential-exfil findings, alerts on new IAM users/keys and on AssumeRole
from unusual IPs.Memory hook"rotate the key" is NOT enoughThe #1 mistake in AWS key-compromise response is rotating the leaked key and closing the ticket — while the attacker's newly created IAM users, access keys, and roles quietly persist. Always scope what the key did and eradicate the persistence it established. And remember AWS access keys don't expire on their own, so a leaked
AKIAkey is good forever until you act — which is exactly why bots scan GitHub for them within seconds of a commit.
Memory hookFun fact — leaked keys are found in seconds. Researchers have shown that an AWS access key committed to a public GitHub repo is often discovered and used by automated bots in under a minute — typically to spin up expensive GPU instances for crypto mining. This is why AWS partnered with GitHub for automatic secret scanning that auto-quarantines exposed keys. The takeaway: there is no "I'll fix it tomorrow" with a leaked key.
Interview Questions
An IAM user is a permanent identity with long-lived credentials — a password and/or access keys that don't expire until rotated. A role is an identity with permissions but no permanent credentials; principals assume it via STS and receive temporary credentials that expire in minutes to hours. Roles are preferred because temporary credentials dramatically reduce risk — there's no long-lived secret sitting in a file or AMI to be stolen, and a leaked temporary credential expires quickly. Roles also enable clean patterns: EC2/Lambda/EKS get credentials by assuming a role rather than embedding keys, and humans federate through SSO and assume roles. The modern principle is roles over users, temporary over permanent, keys for nobody.
The default is deny — if nothing explicitly allows an action, it's denied. For something to be allowed, it must be granted by an identity-based or resource-based policy and permitted by every ceiling that applies: SCPs, the permission boundary, and any session policy. And the overriding rule is that an explicit Deny anywhere always wins — no number of Allows can override it. For same-account access, either an identity policy or the resource policy granting is enough; for cross-account, you need both the identity policy in the caller's account and the resource policy in the target account. So the two sentences are: default deny, and explicit deny beats allow.
A permission boundary is a ceiling on the maximum permissions an IAM user or role can have, regardless of how permissive its attached policies are — effective access is the intersection of the granted policies and the boundary. It solves safe delegation: say you want developers to create their own IAM roles for their services, but giving them iam:* would let them create an admin role and escalate. By requiring that any role they create carries a permission boundary capping it at, say, S3 and DynamoDB, you let them self-serve IAM while guaranteeing they can never create something more powerful than the boundary allows. It's the standard answer to "how do you delegate IAM without enabling privilege escalation."
You attach an instance profile, which wraps an IAM role, to the EC2 instance. The instance retrieves temporary credentials for that role from the instance metadata service at 169.254.169.254, and the SDK uses them automatically — no keys are stored. The security concern is that anything able to make the instance fetch that metadata URL can steal those role credentials, which is why SSRF against EC2 is so dangerous — it was the core of the Capital One breach, where SSRF reached the metadata service, grabbed role credentials, and exfiltrated S3 data. The mitigation is IMDSv2, which requires a session token obtained via a PUT request with a hop limit, something SSRF generally can't perform; enforce IMDSv2 and scope the instance role tightly.
First contain reversibly — deactivate the key (set it inactive) rather than deleting it, so it stops working immediately but I preserve it for scoping. Then scope using CloudTrail filtered to that access key ID: every API call, source IP, and timestamp, looking for the classic pattern of recon calls, then privilege escalation, then resource creation. The critical question is what the attacker created for persistence — new IAM users, access keys, roles, or trust-policy changes — plus any EC2 they launched for mining, any S3 they read, and whether they touched CloudTrail or GuardDuty. Then eradicate: delete the leaked key and, crucially, remove the backdoor identities they created, because rotating the one key while leaving attacker-created users is the most common mistake. Roll any secrets the key could read. Finally harden — replace the IAM user with SSO or a role so there's no long-lived key to leak again, and add push-protection and GuardDuty detections. And I'd note these keys are found by bots in under a minute, so speed matters.
PassRole is the permission to hand an IAM role to an AWS service — for example, specifying an instance profile when launching EC2 passes that role to the EC2 service. It's abused for privilege escalation: if a low-privileged user can pass any role and also run a compute service, they can launch an EC2 instance, Lambda, or similar with a powerful admin role attached and then read that role's credentials or run code as it. So PassRole effectively lets you borrow the permissions of any role you can pass. The mitigation is to scope PassRole tightly with a Resource condition listing only the specific roles a principal may pass, and to treat PassRole as a sensitive, admin-adjacent permission.
Service Control Policies are organization-level guardrails applied to an OU or account that set the maximum permissions for everything in that account — they're a ceiling, not a grant. The key difference from identity policies is that SCPs never grant anything; an action is only allowed if both an IAM policy grants it and no SCP denies it. They're used for preventive controls that even account admins can't override — like denying the ability to disable CloudTrail or GuardDuty, locking actions to approved regions, or blocking leaving the org. A powerful property is that an SCP denying cloudtrail:StopLogging means even a fully-compromised account admin can't turn off logging. The caveat is they don't apply to the management account's root.
For the humans, AWS IAM Identity Center federated to our IdP — they log in through SSO and assume roles for temporary credentials, with no IAM users and no long-lived keys. Access is granted via permission sets mapped to groups, scoped least-privilege, starting minimal and expanding as needed, with MFA enforced at the IdP. For their workloads, roles everywhere — instance profiles, Lambda execution roles, IRSA or Pod Identity for EKS — and OIDC federation for their CI/CD so pipelines assume short-lived roles instead of storing keys. I'd delegate IAM self-service safely with permission boundaries, prefer groups and managed policies for auditability, and run IAM Access Analyzer plus last-accessed data to catch external sharing and prune unused permissions. The summary: SSO for humans, roles for machines, keys for nobody.
Because the account is the strongest isolation boundary AWS gives you — stronger than a VPC or an IAM role — so it's the natural cap on a blast radius. In one big account, isolation depends on every IAM policy being perfect, and a single misconfigured Resource: "*" can leak access across teams; across an account boundary, access requires an explicit, auditable cross-account trust. The pattern is one account per workload per environment — payments-prod, payments-dev, and so on — plus shared accounts for logging and security. So a compromised role in dev simply has no path to prod resources, because they're different accounts with different credentials. Isolation becomes the default and sharing the deliberate exception.
Both are organization-level ceilings that only restrict, never grant, but they govern opposite sides of a request: a Service Control Policy caps what the principals — the identities — inside your accounts can do, while a Resource Control Policy, launched in November 2024, caps who and how anyone can access the resources in your accounts. AWS added RCPs because SCPs only govern your own principals, so they can't stop a misconfigured bucket or key policy from sharing a resource with an external account or the public internet. An RCP sits on the resource side and lets you enforce centrally — for example, deny any principal that isn't in my org from touching S3, STS, KMS, SQS, or Secrets Manager — as a backstop that a too-generous resource policy can't override. Together they form a data perimeter: SCP for "what my people can do," RCP for "who can touch my things."
I'd shrink the blast radius on four axes. The account itself is the isolation boundary, so this workload is already its own account under a Prod OU. On that OU I'd put two SCPs: a region-lock that denies every action outside the one region — with a carve-out for global services like IAM, STS, CloudFront, and Route 53 — so a stolen role can't spin up miners in thirty other regions, and a service-allowlist that denies every service except the handful this workload needs, like S3, EC2, EKS, ECR, ELB, KMS, and logging. At the org root I'd attach a data-perimeter RCP that denies any principal outside my org from touching S3, STS, KMS, and the other supported resources, plus a TLS-only RCP, so even a fat-fingered bucket policy can't leak data externally. Then IMDSv2 forced on the instances and IRSA or Pod Identity so pods get a scoped role instead of inheriting the node's. The one-liner: region-lock plus service-allowlist SCPs, a data-perimeter RCP, IMDSv2, and pod-scoped IAM.
Default to identity-based policies and reserve resource-based policies for the things only they can do — namely cross-account access and letting an AWS service in, like CloudTrail writing to your log bucket — because splitting access logic across both sides gets unauditable fast. When you do write a resource policy, never use a wildcard principal without a condition that scopes it, pin cross-account trust to aws:PrincipalOrgID rather than bare account IDs, and require TLS. What RCPs change is the backstop: instead of hoping every one of hundreds of bucket policies is perfect, you enforce the data perimeter once with an org-level RCP that denies external principals by default, and let individual resource policies grant access only inside that fence. The resource policy is the lock on the door; the RCP is the building's master rule that a single bad door can't override.