Security Notes
Cloud

Cloud Data Security (Vendor-Neutral)

How to protect data in the cloud across its whole life: where it is stored, how it is obscured or encrypted, how you find and classify it, how rights travel with it, how it is retained and destroyed, and how every access is traced. This is CCSP domain 2, the most heavily weighted domain (20%). Provider-specific key management is in AWS data protection.

17 min read 9 sections 7 model answers verified 2026-10

Data Concepts and the Cloud Data Lifecycle

What is this?

The Cloud Security Alliance describes data as moving through six phases. Security controls are chosen per phase, because what can go wrong when data is created differs from what can go wrong when it is shared or destroyed.

Why it matters

"Where are the controls for each phase?" is a structured way to review any data flow, and it is how CCSP frames most data questions.

How it works

CREATE ──► STORE ──► USE ──► SHARE ──► ARCHIVE ──► DESTROY
classify   encrypt,   access   DLP, IRM,  retention,  crypto-shred,
at birth   back up    control, encrypt in lock,       sanitise,
           (at once)  monitor  transit    legal hold   certify
Create

classify at creation (or immediately after); generated, imported or modified data all count.

Store

usually happens at the same time as create; apply encryption, access control and backup.

Use

data is decrypted in memory to be processed; controls are access control, monitoring and confidential computing.

Share

data leaves its home: DLP, information rights management, encryption in transit, contracts.

Archive

long-term storage: retention rules, immutability, readable formats and keys still available years later.

Destroy

remove permanently; in the cloud, usually by crypto-shredding.

Data dispersion (also called bit splitting or erasure coding) splits data into fragments spread across locations, so the loss of one location, or theft of one fragment, doesn't expose or lose the whole.

Data flows should be diagrammed: where data enters, which services and regions it passes through, and where it leaves. Every arrow is a place to ask "encrypted? logged? allowed by law?"

In practiceClassifying at creation usually means a tag or label that later controls can read:

console
$ aws s3api put-object-tagging --bucket acme-exports --key 2026/q3/customers.csv \
    --tagging 'TagSet=[{Key=classification,Value=restricted},{Key=owner,Value=crm-team}]'

A bucket policy, DLP rule or lifecycle rule can then act on classification=restricted — deny sharing, require KMS, delete after 90 days.

📰

Real incident — Toyota (disclosed 2023). A cloud database at a Toyota subsidiary was set to public by mistake and exposed vehicle location data for about 2.15 million customers in Japan for nearly ten years (2013–2023) before anyone noticed. A store-phase control (a public-access block) or a discovery scan would have caught it in days.

🎯

On the job. Draw the data flow for any new feature and ask at every arrow: encrypted in transit, logged, allowed to cross this border, and who owns the copy at the other end?

Security angle

The use phase is the weakest: data must be decrypted to be processed, so access control and monitoring carry the load there.


Storage Architectures and Their Threats

What is this?

The kinds of storage each cloud service model provides, and what typically goes wrong with each.

Why it matters

Controls differ by storage type: an object store's main risk is public exposure, a virtual disk's main risk is snapshot sharing, and SaaS storage's main risk is over-sharing by users.

How it works

IaaS
Volume (block) storage

virtual disks attached to VMs. Risks: unencrypted volumes, public or cross-account snapshot sharing.

Object storage

files in buckets reached by API or URL. Risks: public buckets, overly broad policies, leaked signed URLs.

Ephemeral storage

temporary disks lost when the instance stops; don't keep the only copy of anything there.

PaaS
Structured

relational and managed databases. Risks: public endpoints, weak credentials, injection.

Unstructured

big data, data lakes and blob stores. Risks: broad analyst access to raw sensitive data.

SaaS
Information storage and management

data entered into the application.

Content and file storage

documents shared through the app. Risk: anyone-with-the-link sharing.

Threats across all types: unauthorised access, misconfiguration, accidental deletion, ransomware, loss of keys, data remanence and jurisdiction (data stored in a country whose laws allow access).

In practiceTwo checks that find many real exposures:

console
$ aws s3api get-public-access-block --bucket acme-exports
{ "PublicAccessBlockConfiguration": { "BlockPublicAcls": true, "IgnorePublicAcls": true,
  "BlockPublicPolicy": true, "RestrictPublicBuckets": true } }           ← all four true = good

$ aws ec2 describe-snapshot-attribute --snapshot-id snap-0abc --attribute createVolumePermission
{ "CreateVolumePermissions": [ { "Group": "all" } ] }                     ← PUBLIC disk snapshot: anyone can copy it
📰

Real incident — Microsoft AI research SAS token (2023). Researchers shared training data on GitHub using an Azure storage link (a SAS token) that was scoped to the whole storage account, granted write access and didn't expire until 2051. It exposed 38 TB of internal data, including backups of employees' workstations and Teams messages. Signed links are bearer credentials; scope and expiry are the security.

🎯

On the job. Posture tools (CSPM) flag public buckets, public snapshots and over-broad signed links; the real work is fixing the pipeline or template that keeps creating them.

Security angle

The majority of public cloud data exposures are configuration, not exploitation. Block public access at the account or organisation level, then grant exceptions.


Data Security Technologies

What is this?

The techniques for making data useless to anyone who shouldn't have it: encryption, hashing, masking, tokenisation, anonymisation and data loss prevention.

Why it matters

Each technique fits a different need. Choosing tokenisation instead of encryption can take a whole system out of PCI DSS scope; choosing masking instead of anonymisation can leave you still subject to GDPR.

How it works

  • Encryption — reversible with the key. The question is who holds the key:
    Provider-managed keys

    simplest; the provider can technically decrypt.

    Customer-managed keys in the provider's KMS

    you control policy and rotation.

    BYOK (bring your own key)

    you generate the key and import it.

    HYOK (hold your own key)

    the key never leaves your own HSM; the provider can't decrypt without you.

  • Hashing — one-way fingerprint for integrity checks and lookups; not for confidentiality of guessable values (a hashed phone number is easily brute-forced). See hashing.
  • Masking — hiding part of a value: **** **** **** 4242.
    Static masking

    a masked copy of a database for test environments.

    Dynamic masking

    the original is stored; values are masked on the fly based on who is asking.

  • Tokenisation — replace a sensitive value with a random token; the real value lives only in a separate, secured token vault. Systems that only see tokens can drop out of compliance scope.
  • Pseudonymisation — replace identifiers with pseudonyms that can be re-linked using separately held information. Under GDPR, pseudonymised data is still personal data.
  • Anonymisation — irreversibly remove the link to a person, including indirect identifiers (postcode + birth date + gender can identify most people). Truly anonymous data falls outside GDPR, but true anonymisation is hard.
  • DLP (data loss prevention) — discovers sensitive data and monitors or blocks it leaving approved locations, in storage, in transit and on endpoints.
  • Keys, secrets and certificates management — central managers, rotation, least privilege, and logging of every use.

In practiceThe same customer record, protected four different ways:

Original        : name=Jane Smith  card=4111 1111 1111 1234  dob=1984-03-07  postcode=SW1A 1AA
Masked (display): name=J*** S****  card=**** **** **** 1234  dob=****-**-**  postcode=SW1A ***
Tokenised       : name=Jane Smith  card=tok_9f2c81d7e4       ← real PAN only in the token vault
Pseudonymised   : id=user_58213    card=…                    ← re-linkable with a separate key table
Anonymised      : age_band=40-44   region=London             ← no route back to Jane (if done properly)
📰

Real incident — the Netflix Prize de-anonymisation (2008). Netflix released "anonymised" movie ratings of half a million subscribers for a competition. Researchers matched a few ratings and dates against public IMDb reviews and re-identified users, revealing their full viewing histories. Removing names isn't anonymisation; a few quasi-identifiers can single people out.

🎯

On the job. When a team says "the data is anonymised," ask how: which fields were removed or generalised, and has anyone tried to re-identify it? If it can be re-linked, it's pseudonymised and still personal data.

Security angle

Tokenisation and HYOK both concentrate risk: the token vault and the key server become the crown jewels. Protect and monitor them accordingly.


Data Discovery and Classification

What is this?

Discovery finds where sensitive data actually lives. Classification labels it by sensitivity so the right controls apply.

Why it matters

Organisations consistently underestimate where their sensitive data is: copies in test databases, exports in object storage, attachments in SaaS.

How it works

  • Data types
    Structured

    rows and columns in databases; easiest to scan by schema.

    Semi-structured

    JSON, XML, logs; scan by keys and patterns.

    Unstructured

    documents, images, chat; needs content inspection (regular expressions, machine-learning classifiers, OCR).

  • Discovery methods — metadata (column names, file names), labels already applied, and content analysis. Cloud services such as Macie (AWS) or Sensitive Data Protection (GCP) automate it.
  • Data location — record where data physically resides (region, country) because law follows location.
  • Classification — the data owner assigns a level (public, internal, confidential, restricted) based on impact. Sensitive categories include PII (personally identifiable information), PHI (protected health information) and cardholder data.
  • Mapping and labelling — map data elements to their classification, then apply labels (tags, metadata, document labels) that tools can enforce on.

In practiceWhat an automated discovery finding looks like (Amazon Macie, abridged):

json
{ "type": "SensitiveData:S3Object/Personal",
  "resourcesAffected": { "s3Bucket": { "name": "acme-tmp-exports", "publicAccess": { "effectivePermission": "NOT_PUBLIC" } },
                         "s3Object": { "key": "dumps/crm_full_2024.csv" } },
  "classificationDetails": { "result": { "sensitiveData": [
      { "category": "PERSONAL_INFORMATION", "detections": [ { "type": "EMAIL_ADDRESS", "count": 182344 } ] },
      { "category": "FINANCIAL_INFORMATION", "detections": [ { "type": "CREDIT_CARD_NUMBER", "count": 1210 } ] } ] } } }

The classic pattern: a full CRM export in a "tmp" bucket nobody owns.

🎯

On the job. Run discovery against the places data leaks to — temp and export buckets, analytics sandboxes, test databases, shared drives — not just production systems that are already controlled.

Security angle

Classification only helps if it drives enforcement: labels should automatically change encryption, sharing permissions and DLP rules.


Information Rights Management (IRM)

What is this?

IRM (also called digital rights management for enterprise documents) attaches protection to the document itself, so rules such as "view only, no print, expires Friday" apply wherever the file goes.

Why it matters

Once a file leaves your storage — emailed to a partner, synced to a laptop — perimeter controls no longer apply. IRM keeps control after sharing.

How it works

  1. The document is encrypted and wrapped with a policy at creation or sharing.
  2. To open it, the user's application contacts the rights server, which checks identity and policy.
  3. The server issues a key and the allowed actions (view, edit, print, copy, forward).
  4. The owner can change or revoke rights later, and every access is logged.

Properties CCSP expects you to name: persistence (protection travels with the data), dynamic policy control (rights can change after distribution), expiration, continuous audit trail, replication restrictions (no copy, print or screenshot where enforceable), and support for remote rights revocation. Access is based on certificates or identities issued and revoked by the IRM system.

Limits: users need compatible software, and nothing stops someone photographing the screen.

🎯

On the job. In Microsoft 365 or Google Workspace, IRM shows up as sensitivity labels ("Confidential — Board only: do not forward, no print, expires in 30 days"). The common failure is labels that only mark documents without encrypting them; check that the high-sensitivity labels actually apply protection, and that external sharing of those labels is blocked or logged.

Security angle

IRM is strongest for a small set of very sensitive documents shared externally (board papers, M&A data rooms), not as a blanket control.


What is this?

Policies for how long data is kept, how it is stored long-term, how it is destroyed, and how deletion is suspended when litigation is expected.

Why it matters

Keeping data too long is a breach and privacy liability; deleting it too early can break the law. Both are compliance failures.

How it works

Retention policy

per data class: retention period, storage format, legal and regulatory basis, and who approves exceptions.

Archiving

cheaper, slower storage; must remain readable for the whole period, which means keeping formats, software and encryption keys available.

Deletion

in the cloud you can't physically wipe the provider's disks, so the practical method is crypto-shredding: encrypt data with keys you control and destroy the keys. Keep proof of destruction.

Legal hold

when litigation or an investigation is reasonably expected, deletion stops for the relevant data regardless of the retention schedule. Cloud storage supports this directly (for example object legal holds).

Immutability

write-once storage for records that must not change. See integrity and immutability.

In practiceLegal hold and deletion are explicit API actions you can audit:

console
$ aws s3api put-object-legal-hold --bucket acme-records --key hr/case-4471.pdf --legal-hold Status=ON
$ aws kms schedule-key-deletion --key-id 1234abcd-... --pending-window-in-days 30   ← crypto-shred: all data under this key becomes unreadable
🎯

On the job. When legal issues a hold, someone must actually suspend lifecycle deletion, backup expiry and mailbox retention for the affected data, and prove it. Make that a runbook with named owners, not an email.

Security angle

Backups and replicas must follow the same retention and deletion rules, or "deleted" data survives in a forgotten snapshot.


Auditability, Traceability and Accountability of Data Events

What is this?

Being able to answer, for any data event, who did what, when, where and from where, and prove the record is intact.

Why it matters

Without attribution you can't investigate, can't prove compliance and can't hold anyone accountable. Logging is also required by most regulations.

How it works

Event sources

differ by service model: in IaaS you can collect OS, network and API logs; in PaaS mostly service and API logs; in SaaS only what the vendor exposes. Check what the provider offers before you sign.

Identity attribution

every event tied to a specific user or workload identity, not a shared account.

Centralised, protected log storage

logs sent to a separate account or system the logged parties can't modify.

SIEM (security information and event management)

correlation and alerting. See the SIEM pipeline.

Chain of custody and non-repudiation

integrity protection (hashing, signed digests, write-once storage) so logs stand up as evidence, and so a user can't credibly deny an action.

📰

Real incident — Storm-0558 and the missing logs (2023). When attackers forged tokens to read government mailboxes in Microsoft's cloud, the US State Department spotted it through an alert on the MailItemsAccessed audit event. Many other customers couldn't have, because that event was only available in a premium logging tier. After public pressure, Microsoft expanded default logging for all customers. Lesson: in SaaS, the logs you didn't buy are the investigation you can't run.

🎯

On the job. For each critical SaaS application, list which events you get (sign-in, admin change, data read, sharing), how long they're kept, and whether they reach your SIEM. Fill the gaps before the incident.

Security angle

SaaS logging is often a paid tier or limited to 30–90 days. Negotiate log access and retention in the contract, not after the incident.


Protecting AI and ML Data

Last verified2026-10 — added to the CCSP outline effective 1 August 2026.

What is this?

Protecting the data used to train, tune and run AI models, and the models themselves, which are a form of data derived from it.

Why it matters

Training data is often the largest concentration of sensitive data in an organisation, and models can leak what they were trained on.

How it works

Data provenance and lineage

know where every training dataset came from, under what consent or licence, and what transformations were applied.

Minimisation and de-identification

remove or pseudonymise personal data before training where possible.

Integrity

protect training data and pipelines from data poisoning (malicious samples that change model behaviour). Version datasets and restrict write access.

Confidentiality of models

model weights are valuable intellectual property; protect them like source code and watch for model extraction through excessive API queries.

Privacy attacks

membership inference (was this person in the training data?) and model inversion (reconstructing training examples from outputs). Mitigations include differential privacy, output filtering and rate limiting.

Prompts and outputs

prompts can contain sensitive data and outputs can leak it; apply DLP and logging to both, and check provider terms on whether your data is used for training.

Deletion requests

removing a person's data from a trained model may require retraining; plan for it when using personal data.

📰

Real incident — source code pasted into a chatbot (2023). Samsung engineers pasted confidential source code and meeting notes into a public AI chatbot to get help, sending them to a service whose terms then allowed using inputs for training. Samsung restricted generative-AI use on company devices. The leak wasn't a hack; it was the default data flow of a convenient tool.

📰

Real incident — training-data extraction (2023). Researchers showed that asking a production chatbot to repeat a single word forever eventually made it emit memorised training data verbatim, including personal contact details. Models can leak what they were trained on.

🎯

On the job. Before approving an AI tool, check: does the provider train on our inputs, where are prompts stored and for how long, can we enforce enterprise settings and DLP on prompts, and is there an audit log of who sent what.

Security angle

The training pipeline is a supply chain: datasets, pre-trained models and libraries pulled from public sources can all be poisoned. See the OWASP LLM Top 10.


Interview Questions

Q
Walk through the cloud data lifecycle and a control for each phase.
Model answer

Create, store, use, share, archive and destroy. You classify at creation, encrypt and back up on storage, enforce access control and monitoring while in use, apply DLP, rights management and TLS when sharing, enforce retention and immutability in the archive, and crypto-shred on destruction. The use phase is the weak spot because data has to be decrypted to be processed, so access control and monitoring carry the load there.

Q
Tokenisation versus encryption — when would you choose tokenisation?
Model answer

Encryption transforms the value mathematically and anyone with the key can reverse it, whereas tokenisation swaps it for a random token with no mathematical link and keeps the real value in a separate vault. I'd choose tokenisation for things like card numbers, because systems that only ever see tokens can fall out of PCI DSS scope entirely. The trade-off is that the token vault becomes a crown jewel you must protect and monitor.

Q
Is pseudonymised data still personal data under GDPR?
Model answer

Yes. Pseudonymisation replaces identifiers but the data can be re-linked with separately held information, so it's still personal data — it just reduces risk and is recognised as a good security measure. Only truly anonymised data, where re-identification isn't reasonably possible even through indirect identifiers like postcode and birth date, falls outside GDPR, and that's much harder to achieve than people assume.

Q
How do you securely delete data in a public cloud where you can't touch the disks?
Model answer

Crypto-shredding: encrypt the data under keys you control from day one, then destroy the keys so every copy — including replicas and backups encrypted under them — becomes unreadable. It only works if you planned the key hierarchy up front and know which keys cover which data. Keep a record of the key destruction as proof for auditors.

Q
What is information rights management and when is it worth it?
Model answer

IRM wraps a document with encryption and a policy, so rules like view-only, no printing and expiry travel with the file, and the owner can revoke access after it's been shared, with every open logged. It's worth it for a small set of highly sensitive documents shared outside your control, like board packs or deal rooms. It's not a blanket control: it needs compatible software and can't stop someone photographing the screen.

Q
How can a trained model leak its training data, and what do you do about it?
Model answer

Through membership inference, where an attacker tests whether a specific record was in the training set, and model inversion or regurgitation, where outputs reconstruct training examples. Defences start with minimising and de-identifying personal data before training, then differential privacy, output filtering and rate limiting against extraction. And because removing a person later may mean retraining, you decide what personal data goes in very deliberately.

Q
What should you check about logging before signing a SaaS contract?
Model answer

Which events the vendor logs — sign-ins, admin changes, data access, sharing — whether you can export them to your SIEM by API, how long they're retained, and whether that's an extra-cost tier. In SaaS you only get what the vendor exposes, so if data-access logs aren't available, you won't be able to scope a breach. That needs to be in the contract, because after an incident it's too late.