Cloud Security Concepts & Architecture (Vendor-Neutral)
The provider-independent ideas behind cloud security: what makes something "cloud," who is responsible for what, how cloud platforms are built and where they break, how to judge a provider, and the security components cloud applications are built from. This is the core of CCSP (Certified Cloud Security Professional) domains 1, 3 and 4. AWS- and GCP-specific material lives in AWS and GCP.
Cloud Computing Concepts and Reference Architecture
What is this?
Cloud computing is renting computing resources over a network, on demand, from a shared pool, paying for what you use. NIST SP 800-145 (2011) gives the definition everyone uses: five essential characteristics, three service models and four deployment models.
Why it matters
These definitions are the shared vocabulary for every cloud contract, audit and exam question. "Which model is this?" decides who is responsible for patching, logging and encryption.
How it works
Five essential characteristics
you provision resources yourself, without a ticket to a human.
available over the network from many device types.
the provider serves many customers (tenants) from shared hardware; you don't know or control exactly where.
capacity scales out and in quickly, appearing unlimited.
usage is metered and billed.
Service models
virtual machines, networks, storage. You manage the OS upward. Example: EC2, Compute Engine.
a managed runtime or database; you manage code, data and configuration. Example: managed databases, App Engine.
a finished application; you manage users, data and settings. Example: email, CRM.
Newer variants follow the same logic: serverless / FaaS (function as a service), CaaS (containers as a service).
Deployment models
shared by anyone who pays.
for one organisation, on-prem or hosted.
shared by organisations with common needs (government agencies, a research consortium).
two or more models connected, such as on-prem private plus public cloud. Multi-cloud means using several public providers.
Roles (ISO/IEC 17789)
the organisation using the service.
runs the service.
supports either side, including auditors, integrators and cloud service brokers (who aggregate, integrate or customise services from several providers).
Cross-cutting properties to recognise
Interoperability (works with other systems), portability (can move workloads and data elsewhere), reversibility (can get your data back and leave cleanly at contract end), availability, resiliency, auditability, governance, regulatory compliance, security, privacy, performance, maintenance, versioning, and service levels.
Related technologies
Virtualisation, containers, serverless, edge computing, confidential computing (processing data inside hardware-isolated enclaves so even the host can't read it), DevSecOps, quantum computing (a threat to today's public-key cryptography; see the quantum question) and AI.
In practiceEach characteristic shows up as something you can do in a minute and get billed for by the second:
$ aws ec2 run-instances --image-id ami-0abc --instance-type c7g.large --count 50 ← self-service + elasticity
$ aws ce get-cost-and-usage --time-period Start=2026-10-01,End=2026-10-11 \
--granularity DAILY --metrics UnblendedCost ← measured serviceThe security consequence of "anyone with the right API key can create 50 machines in a minute" is that the API key is now the most valuable thing you own — and a cost spike is often the first sign it was stolen.
🎯On the job. In any architecture review, name the service model of each component (IaaS VM, PaaS database, SaaS CRM): it immediately tells you which controls are yours to check.
Security angle
Resource pooling is the root of most cloud-specific risk: your workload shares hardware, hypervisors and management planes with strangers, and you rely on the provider's isolation.
Shared Responsibility
What is this?
Who secures which layer. The provider always owns the physical data centre; the customer always owns their data, their users and how they configure access. Everything in between shifts with the service model.
Why it matters
Most cloud breaches are on the customer side of the line: misconfigured storage, leaked keys, excessive permissions. Knowing the line is the first step in any cloud risk assessment or contract review.
How it works
| Layer | On-prem | IaaS | PaaS | SaaS |
|---|---|---|---|---|
| Data, classification, accountability | You | You | You | You |
| Identities and access configuration | You | You | You | You |
| Application | You | You | You | Provider |
| Runtime, middleware | You | You | Provider | Provider |
| Operating system | You | You | Provider | Provider |
| Virtualisation, hypervisor | You | Provider | Provider | Provider |
| Physical servers, network, facility | You | Provider | Provider | Provider |
You can delegate responsibility but never accountability. If customer data leaks from a SaaS provider, regulators still hold the customer (the data controller) accountable.
📰Real incident — Snowflake customers (2024). Around 165 organisations using Snowflake had data stolen after attackers logged in with passwords harvested by infostealer malware, on accounts with no MFA. The provider's platform was intact; identity configuration sat on the customer's side of the line. See cloud contracts.
🎯On the job. For each SaaS or PaaS service, write down the customer-side controls explicitly — MFA enforcement, network allow-lists, key ownership, log export, backup — and check them in the product. The provider's marketing diagram is not your control list.
Security angle
Read the provider's actual responsibility documentation and contract, not the marketing diagram: logging, backups and encryption key ownership are often less covered than customers assume.
Secure Cloud Design Principles
What is this?
The design rules that apply specifically to cloud workloads, on top of general secure design principles.
Why it matters
Cloud changes the economics and the failure modes: infrastructure is code, everything is reachable by API, and capacity is cheap but misconfiguration is instant and global.
How it works
apply controls at every phase: create → store → use → share → archive → destroy (the Cloud Security Alliance lifecycle). See cloud data security.
multiple availability zones and regions, automated recovery, backups in a separate account. See business continuity and DR.
every API call is authenticated and authorised; network controls are a second layer.
rebuild rather than patch in place; review changes as code.
preventive policies, detective rules and automatic remediation instead of manual review.
the cheapest compliant design that meets the risk appetite; elasticity also means attackers can run up your bill.
portable data formats and an exit plan, weighed against the security benefits of native services.
IaaS demands OS hardening and network design; PaaS demands configuration and IAM rigour; SaaS demands identity, data and vendor management.
Security angle
The API is the new perimeter. Credentials that can call the management API are worth more than any single server.
Evaluating Cloud Service Providers
What is this?
How a customer decides whether a provider's security is good enough, without being able to inspect the data centre.
Why it matters
You inherit your provider's controls. Since you can't audit hyperscalers yourself, you rely on independent certifications and attestations, and you need to know what each one does and doesn't cover.
How it works
the provider runs a certified information security management system.
additional security controls specific to cloud services, for both providers and customers.
protecting personal data (PII) in public clouds acting as processors.
(Cloud Security Alliance Security, Trust, Assurance and Risk) — Level 1 self-assessment published in a public registry; Level 2 third-party audit or certification. Built on the CSA Cloud Controls Matrix (CCM) and its questionnaire (CAIQ).
an auditor's report on controls for security, availability, processing integrity, confidentiality and privacy. Type II covers operating effectiveness over a period. See SOC 2.
the provider's environment is assessed for card-data workloads.
US federal authorisation at low, moderate or high impact. See FedRAMP.
and FIPS 140-3 — product evaluation and validation of cryptographic modules.
Always check the scope: which services, regions and time period the certificate or report covers, and the complementary user entity controls the customer must implement.
In practiceA fast vendor triage, before the long questionnaire:
1. Certifications in scope? ISO 27001 + 27017/27018 certificate covers "Platform X, EU regions"? ✔
2. SOC 2 Type II current? Period ends < 12 months ago, bridge letter for the gap? ✔
3. Exceptions? Any in access control, change management, logging? ⚠ 2 in CC6.1
4. Our obligations (CUECs)? MFA, key management, user reviews — can we actually do them? ✔
5. Sub-processors? Who else touches our data, and where? list attached
6. Breach history? Public incidents, how they communicated one in 2024, handled well🎯On the job. The CSA STAR registry is a quick public check of what a provider has self-assessed or had audited against the Cloud Controls Matrix. Absence from STAR isn't disqualifying; inconsistency between STAR answers and the SOC 2 report is a red flag.
Security angle
A certificate proves a control framework was met at audit time; it doesn't prove your configuration is safe. Pair provider assurance with your own posture management.
AI and Machine Learning in the Cloud
Last verified2026-10 — added to the CCSP outline effective 1 August 2026.
What is this?
Most AI workloads run on cloud platforms: training on rented GPUs, models served as managed endpoints, and generative AI consumed as an API. This section covers the security concepts specific to them; the data side is in AI and ML data protection.
Why it matters
AI adds new assets (models, training data, prompts) and new attack paths (prompt injection, poisoning, model theft) on top of the usual cloud risks.
How it works
collect data → prepare → train → evaluate → deploy → monitor → retrain. Each stage has its own risks and owners.
with a hosted model API, the provider secures the model and infrastructure; you secure your prompts, data sent to it, outputs, the tools your agents can call, and who can invoke the model.
prompt injection, insecure output handling, training-data poisoning, model theft or extraction, excessive agency in AI agents, and sensitive data in prompts or training sets. See the OWASP LLM Top 10.
least-privilege access to models and data, guardrails that filter inputs and outputs, logging of model invocations, isolation of training environments, provenance of models and datasets, and human approval for high-impact agent actions.
📰Real incident — EchoLeak (2025, CVE-2025-32711). Researchers showed that a single crafted email could make Microsoft 365 Copilot leak internal data with no click from the victim: when the user later asked Copilot an ordinary question, it read the attacker's email as context, followed its hidden instructions, gathered data from OneDrive, SharePoint and Teams, and sent it out through a link that passed a trusted Microsoft domain. Microsoft fixed it server-side in June 2025. It is the textbook case of prompt injection plus excessive access.
🎯On the job. For an AI assistant connected to company data, ask: what can it read (and is that scoped to the user's own permissions), what can it do (send, write, call APIs), can untrusted content reach its context (email, web pages, uploaded files), and is every action logged?
Security angle
Treat an AI agent with tool access like a new employee with system access who can be talked into anything by any document it reads. Scope its permissions accordingly.
Cloud Platform and Infrastructure Security
What is this?
How the provider's underlying platform is built: physical data centres, networks, compute, virtualisation, storage and the management plane, and the risks each layer creates.
Why it matters
Even though customers don't run this layer, CCSP expects you to understand it to assess provider risk and to secure private clouds you may operate yourself.
How it works
Components
- Physical environment — data centres, power, cooling.
- Network and communications — physical network plus software-defined overlays that give each tenant an isolated virtual network.
- Compute — physical hosts running many tenants' workloads.
- Virtualisation — the hypervisor divides a host into virtual machines:Type 1 (bare metal)
runs directly on hardware (KVM, Xen, ESXi, Hyper-V). Used by cloud providers; smaller attack surface.
Type 2 (hosted)runs on top of an ordinary OS (VirtualBox, VMware Workstation). Inherits the host OS's attack surface.
- Storage — shared storage arrays and object stores, logically separated per tenant.
- Management plane — the APIs and consoles that control everything. Compromising it means compromising every tenant.
Secure data-centre design
tenant partitioning, access control to management systems, redundancy.
location away from flood plains and flight paths, hardened buildings, layered physical access (see physical security).
power, HVAC, fire suppression, and multiple independent network carriers entering by different routes.
Tier I basic capacity, Tier II redundant components, Tier III concurrently maintainable (any component can be serviced without downtime), Tier IV fault tolerant (survives any single failure without impact).
Risks specific to cloud infrastructure
code in a guest breaks out into the hypervisor and reaches other tenants.
co-resident tenants infer secrets through shared CPU caches (the Spectre and Meltdown class, CVE-2017-5753 and CVE-2017-5754).
stolen admin or API credentials.
data left on reallocated storage; mitigated by encryption and crypto-shredding.
, misconfiguration, insufficient isolation between tenants, and supply-chain risk in provider hardware and software.
Countermeasures: patched and minimal hypervisors, hardware-assisted isolation, encryption everywhere, strong management-plane IAM with MFA, dedicated hosts for sensitive workloads, and continuous monitoring.
📰Real incident — ChaosDB (2021). Researchers at Wiz found that a notebook feature in Azure Cosmos DB let them obtain the primary access keys of other customers' databases — a cross-tenant break of the isolation every customer relies on. Microsoft disabled the feature and told thousands of customers to rotate keys. Tenant isolation failures are rare, but they do happen, and key rotation was the customer's only lever.
📰Real incident — OMIGOD (2021, CVE-2021-38647). Azure silently installed an agent called OMI on many Linux VMs when certain services were enabled. A bug allowed unauthenticated remote code execution as root on any VM where its port was exposed. Customers were exposed by software they didn't know they were running — part of the platform, but running inside their VMs.
📰Real incident — VENOM (2015, CVE-2015-3456). A bug in the virtual floppy-disk controller code used by QEMU-based hypervisors (Xen, KVM) could let a guest VM escape to the host. Cloud providers patched and rebooted fleets urgently. Even legacy virtual hardware nobody used was attack surface.
Security angle
For customers, the realistic infrastructure risk is not VM escape but the management plane: your own cloud credentials. Protect them like domain admin.
Cloud Application Architecture
What is this?
The extra security components and patterns cloud applications are assembled from, beyond the application code itself.
Why it matters
Cloud apps are built from managed services glued together by APIs. Their security depends as much on the components around the code as on the code.
How it works
filters HTTP attacks at the edge.
watches queries to databases for abuse and exfiltration.
validate, authenticate, rate-limit and log API calls before they reach the service.
sits between users and SaaS apps to enforce policy: discover shadow IT, apply DLP, control sharing, detect risky sign-ins.
TLS everywhere, encryption at rest with managed keys, secrets in a secrets manager.
run untrusted code in isolated environments (separate accounts, gVisor, Firecracker micro-VMs).
containers and microservices orchestrated by Kubernetes; see Kubernetes security.
federation, SSO and MFA for users; workload identity for services; see authentication.
threat modelling (STRIDE, PASTA), SAST, DAST, IAST, SCA and abuse-case testing; see secure software development and AppSec.
In practiceAn API gateway turns several application-architecture controls into configuration:
# API gateway route policy (generic)
route: POST /v1/payments
auth: { type: jwt, issuer: https://login.acme.com, audience: payments-api } # authenticate
authorize: { require_scope: payments:write } # authorise
limits: { rate: 20/min per client_id, body_max: 64KB } # abuse and DoS
validate: { schema: openapi/payments.yaml } # reject malformed input
log: { fields: [client_id, route, status, latency], destination: siem } # traceability🎯On the job. In a design review, ask which of these controls live at the gateway and which inside the service. "The gateway authenticates, so the service doesn't check" is a common gap when internal traffic can bypass the gateway.
Security angle
Every managed service you add is another configuration and another identity to get right. Fewer, well-understood components beat many lightly understood ones.
Interview Questions
The provider always owns physical security and the customer always owns their data, identities and access configuration. In IaaS you also own the operating system, runtime and application; in PaaS the provider takes the OS and runtime and you own code and configuration; in SaaS you're left with users, data and settings. The catch is that responsibility can be delegated but accountability can't — if a SaaS vendor leaks your customers' data, regulators still come to you.
NIST SP 800-145 lists on-demand self-service, broad network access, resource pooling across tenants, rapid elasticity and measured service. Resource pooling is the security-relevant one: your workload shares hardware, hypervisors and management systems with other customers, so isolation becomes something you rely on the provider for.
Through independent assurance: their ISO 27001 certificate with the 27017 cloud and 27018 privacy extensions, a SOC 2 Type II report, their CSA STAR entry built on the Cloud Controls Matrix, and FedRAMP or PCI where relevant. I always check the scope — which services, regions and period — and the complementary controls I'm expected to run. A certificate shows their controls worked at audit time; it says nothing about my configuration.
A Type 1 hypervisor runs directly on the hardware, like KVM, Xen or ESXi, while a Type 2 runs as an application on a normal operating system, like VirtualBox. Type 1 has a much smaller attack surface, which is why cloud providers use it. A Type 2 inherits every weakness of its host OS, so compromising the host compromises every guest.
Not VM escape, which is rare and the provider's problem, but the management plane — your own credentials that can call the cloud API. With those, an attacker can read data, disable logging and create persistence without exploiting anything. So protecting cloud admin credentials with MFA, short-lived roles and guardrails matters more than worrying about hypervisor bugs.
A cloud access security broker sits between users and SaaS applications, via proxy or API, to give visibility and control you otherwise lose. It discovers shadow IT, applies DLP to uploads and sharing, enforces access policy such as blocking unmanaged devices, and flags risky behaviour like mass downloads. It's essentially the SaaS-era replacement for the visibility the corporate perimeter used to give you.
It adds an actor that takes instructions from content it reads, so any document, web page or email it processes can try to steer it — that's prompt injection. If the agent has tools or cloud permissions, injection turns into real actions like data exfiltration. So I scope its permissions to the minimum, filter inputs and outputs with guardrails, log every invocation and require human approval for high-impact actions.