Cloud Data Security (Vendor-Neutral)
How to protect data in the cloud across its whole life: where it is stored, how it is obscured or encrypted, how you find and classify it, how rights travel with it, how it is retained and destroyed, and how every access is traced. This is CCSP domain 2, the most heavily weighted domain (20%). Provider-specific key management is in AWS data protection.
Data Concepts and the Cloud Data Lifecycle
What is this?
The Cloud Security Alliance describes data as moving through six phases. Security controls are chosen per phase, because what can go wrong when data is created differs from what can go wrong when it is shared or destroyed.
Why it matters
"Where are the controls for each phase?" is a structured way to review any data flow, and it is how CCSP frames most data questions.
How it works
CREATE ──► STORE ──► USE ──► SHARE ──► ARCHIVE ──► DESTROY
classify encrypt, access DLP, IRM, retention, crypto-shred,
at birth back up control, encrypt in lock, sanitise,
(at once) monitor transit legal hold certifyclassify at creation (or immediately after); generated, imported or modified data all count.
usually happens at the same time as create; apply encryption, access control and backup.
data is decrypted in memory to be processed; controls are access control, monitoring and confidential computing.
data leaves its home: DLP, information rights management, encryption in transit, contracts.
long-term storage: retention rules, immutability, readable formats and keys still available years later.
remove permanently; in the cloud, usually by crypto-shredding.
Data dispersion (also called bit splitting or erasure coding) splits data into fragments spread across locations, so the loss of one location, or theft of one fragment, doesn't expose or lose the whole.
Data flows should be diagrammed: where data enters, which services and regions it passes through, and where it leaves. Every arrow is a place to ask "encrypted? logged? allowed by law?"
In practiceClassifying at creation usually means a tag or label that later controls can read:
$ aws s3api put-object-tagging --bucket acme-exports --key 2026/q3/customers.csv \
--tagging 'TagSet=[{Key=classification,Value=restricted},{Key=owner,Value=crm-team}]'A bucket policy, DLP rule or lifecycle rule can then act on classification=restricted — deny sharing, require KMS, delete after 90 days.
📰Real incident — Toyota (disclosed 2023). A cloud database at a Toyota subsidiary was set to public by mistake and exposed vehicle location data for about 2.15 million customers in Japan for nearly ten years (2013–2023) before anyone noticed. A store-phase control (a public-access block) or a discovery scan would have caught it in days.
🎯On the job. Draw the data flow for any new feature and ask at every arrow: encrypted in transit, logged, allowed to cross this border, and who owns the copy at the other end?
Security angle
The use phase is the weakest: data must be decrypted to be processed, so access control and monitoring carry the load there.
Storage Architectures and Their Threats
What is this?
The kinds of storage each cloud service model provides, and what typically goes wrong with each.
Why it matters
Controls differ by storage type: an object store's main risk is public exposure, a virtual disk's main risk is snapshot sharing, and SaaS storage's main risk is over-sharing by users.
How it works
virtual disks attached to VMs. Risks: unencrypted volumes, public or cross-account snapshot sharing.
files in buckets reached by API or URL. Risks: public buckets, overly broad policies, leaked signed URLs.
temporary disks lost when the instance stops; don't keep the only copy of anything there.
relational and managed databases. Risks: public endpoints, weak credentials, injection.
big data, data lakes and blob stores. Risks: broad analyst access to raw sensitive data.
data entered into the application.
documents shared through the app. Risk: anyone-with-the-link sharing.
Threats across all types: unauthorised access, misconfiguration, accidental deletion, ransomware, loss of keys, data remanence and jurisdiction (data stored in a country whose laws allow access).
In practiceTwo checks that find many real exposures:
$ aws s3api get-public-access-block --bucket acme-exports
{ "PublicAccessBlockConfiguration": { "BlockPublicAcls": true, "IgnorePublicAcls": true,
"BlockPublicPolicy": true, "RestrictPublicBuckets": true } } ← all four true = good
$ aws ec2 describe-snapshot-attribute --snapshot-id snap-0abc --attribute createVolumePermission
{ "CreateVolumePermissions": [ { "Group": "all" } ] } ← PUBLIC disk snapshot: anyone can copy it📰Real incident — Microsoft AI research SAS token (2023). Researchers shared training data on GitHub using an Azure storage link (a SAS token) that was scoped to the whole storage account, granted write access and didn't expire until 2051. It exposed 38 TB of internal data, including backups of employees' workstations and Teams messages. Signed links are bearer credentials; scope and expiry are the security.
🎯On the job. Posture tools (CSPM) flag public buckets, public snapshots and over-broad signed links; the real work is fixing the pipeline or template that keeps creating them.
Security angle
The majority of public cloud data exposures are configuration, not exploitation. Block public access at the account or organisation level, then grant exceptions.
Data Security Technologies
What is this?
The techniques for making data useless to anyone who shouldn't have it: encryption, hashing, masking, tokenisation, anonymisation and data loss prevention.
Why it matters
Each technique fits a different need. Choosing tokenisation instead of encryption can take a whole system out of PCI DSS scope; choosing masking instead of anonymisation can leave you still subject to GDPR.
How it works
- Encryption — reversible with the key. The question is who holds the key:Provider-managed keys
simplest; the provider can technically decrypt.
Customer-managed keys in the provider's KMSyou control policy and rotation.
BYOK (bring your own key)you generate the key and import it.
HYOK (hold your own key)the key never leaves your own HSM; the provider can't decrypt without you.
- Hashing — one-way fingerprint for integrity checks and lookups; not for confidentiality of guessable values (a hashed phone number is easily brute-forced). See hashing.
- Masking — hiding part of a value:
**** **** **** 4242.Static maskinga masked copy of a database for test environments.
Dynamic maskingthe original is stored; values are masked on the fly based on who is asking.
- Tokenisation — replace a sensitive value with a random token; the real value lives only in a separate, secured token vault. Systems that only see tokens can drop out of compliance scope.
- Pseudonymisation — replace identifiers with pseudonyms that can be re-linked using separately held information. Under GDPR, pseudonymised data is still personal data.
- Anonymisation — irreversibly remove the link to a person, including indirect identifiers (postcode + birth date + gender can identify most people). Truly anonymous data falls outside GDPR, but true anonymisation is hard.
- DLP (data loss prevention) — discovers sensitive data and monitors or blocks it leaving approved locations, in storage, in transit and on endpoints.
- Keys, secrets and certificates management — central managers, rotation, least privilege, and logging of every use.
In practiceThe same customer record, protected four different ways:
Original : name=Jane Smith card=4111 1111 1111 1234 dob=1984-03-07 postcode=SW1A 1AA
Masked (display): name=J*** S**** card=**** **** **** 1234 dob=****-**-** postcode=SW1A ***
Tokenised : name=Jane Smith card=tok_9f2c81d7e4 ← real PAN only in the token vault
Pseudonymised : id=user_58213 card=… ← re-linkable with a separate key table
Anonymised : age_band=40-44 region=London ← no route back to Jane (if done properly)📰Real incident — the Netflix Prize de-anonymisation (2008). Netflix released "anonymised" movie ratings of half a million subscribers for a competition. Researchers matched a few ratings and dates against public IMDb reviews and re-identified users, revealing their full viewing histories. Removing names isn't anonymisation; a few quasi-identifiers can single people out.
🎯On the job. When a team says "the data is anonymised," ask how: which fields were removed or generalised, and has anyone tried to re-identify it? If it can be re-linked, it's pseudonymised and still personal data.
Security angle
Tokenisation and HYOK both concentrate risk: the token vault and the key server become the crown jewels. Protect and monitor them accordingly.
Data Discovery and Classification
What is this?
Discovery finds where sensitive data actually lives. Classification labels it by sensitivity so the right controls apply.
Why it matters
Organisations consistently underestimate where their sensitive data is: copies in test databases, exports in object storage, attachments in SaaS.
How it works
- Data typesStructured
rows and columns in databases; easiest to scan by schema.
Semi-structuredJSON, XML, logs; scan by keys and patterns.
Unstructureddocuments, images, chat; needs content inspection (regular expressions, machine-learning classifiers, OCR).
- Discovery methods — metadata (column names, file names), labels already applied, and content analysis. Cloud services such as Macie (AWS) or Sensitive Data Protection (GCP) automate it.
- Data location — record where data physically resides (region, country) because law follows location.
- Classification — the data owner assigns a level (public, internal, confidential, restricted) based on impact. Sensitive categories include PII (personally identifiable information), PHI (protected health information) and cardholder data.
- Mapping and labelling — map data elements to their classification, then apply labels (tags, metadata, document labels) that tools can enforce on.
In practiceWhat an automated discovery finding looks like (Amazon Macie, abridged):
{ "type": "SensitiveData:S3Object/Personal",
"resourcesAffected": { "s3Bucket": { "name": "acme-tmp-exports", "publicAccess": { "effectivePermission": "NOT_PUBLIC" } },
"s3Object": { "key": "dumps/crm_full_2024.csv" } },
"classificationDetails": { "result": { "sensitiveData": [
{ "category": "PERSONAL_INFORMATION", "detections": [ { "type": "EMAIL_ADDRESS", "count": 182344 } ] },
{ "category": "FINANCIAL_INFORMATION", "detections": [ { "type": "CREDIT_CARD_NUMBER", "count": 1210 } ] } ] } } }The classic pattern: a full CRM export in a "tmp" bucket nobody owns.
🎯On the job. Run discovery against the places data leaks to — temp and export buckets, analytics sandboxes, test databases, shared drives — not just production systems that are already controlled.
Security angle
Classification only helps if it drives enforcement: labels should automatically change encryption, sharing permissions and DLP rules.
Information Rights Management (IRM)
What is this?
IRM (also called digital rights management for enterprise documents) attaches protection to the document itself, so rules such as "view only, no print, expires Friday" apply wherever the file goes.
Why it matters
Once a file leaves your storage — emailed to a partner, synced to a laptop — perimeter controls no longer apply. IRM keeps control after sharing.
How it works
- The document is encrypted and wrapped with a policy at creation or sharing.
- To open it, the user's application contacts the rights server, which checks identity and policy.
- The server issues a key and the allowed actions (view, edit, print, copy, forward).
- The owner can change or revoke rights later, and every access is logged.
Properties CCSP expects you to name: persistence (protection travels with the data), dynamic policy control (rights can change after distribution), expiration, continuous audit trail, replication restrictions (no copy, print or screenshot where enforceable), and support for remote rights revocation. Access is based on certificates or identities issued and revoked by the IRM system.
Limits: users need compatible software, and nothing stops someone photographing the screen.
🎯On the job. In Microsoft 365 or Google Workspace, IRM shows up as sensitivity labels ("Confidential — Board only: do not forward, no print, expires in 30 days"). The common failure is labels that only mark documents without encrypting them; check that the high-sensitivity labels actually apply protection, and that external sharing of those labels is blocked or logged.
Security angle
IRM is strongest for a small set of very sensitive documents shared externally (board papers, M&A data rooms), not as a blanket control.
Retention, Deletion, Archiving and Legal Hold
What is this?
Policies for how long data is kept, how it is stored long-term, how it is destroyed, and how deletion is suspended when litigation is expected.
Why it matters
Keeping data too long is a breach and privacy liability; deleting it too early can break the law. Both are compliance failures.
How it works
per data class: retention period, storage format, legal and regulatory basis, and who approves exceptions.
cheaper, slower storage; must remain readable for the whole period, which means keeping formats, software and encryption keys available.
in the cloud you can't physically wipe the provider's disks, so the practical method is crypto-shredding: encrypt data with keys you control and destroy the keys. Keep proof of destruction.
when litigation or an investigation is reasonably expected, deletion stops for the relevant data regardless of the retention schedule. Cloud storage supports this directly (for example object legal holds).
write-once storage for records that must not change. See integrity and immutability.
In practiceLegal hold and deletion are explicit API actions you can audit:
$ aws s3api put-object-legal-hold --bucket acme-records --key hr/case-4471.pdf --legal-hold Status=ON
$ aws kms schedule-key-deletion --key-id 1234abcd-... --pending-window-in-days 30 ← crypto-shred: all data under this key becomes unreadable🎯On the job. When legal issues a hold, someone must actually suspend lifecycle deletion, backup expiry and mailbox retention for the affected data, and prove it. Make that a runbook with named owners, not an email.
Security angle
Backups and replicas must follow the same retention and deletion rules, or "deleted" data survives in a forgotten snapshot.
Auditability, Traceability and Accountability of Data Events
What is this?
Being able to answer, for any data event, who did what, when, where and from where, and prove the record is intact.
Why it matters
Without attribution you can't investigate, can't prove compliance and can't hold anyone accountable. Logging is also required by most regulations.
How it works
differ by service model: in IaaS you can collect OS, network and API logs; in PaaS mostly service and API logs; in SaaS only what the vendor exposes. Check what the provider offers before you sign.
every event tied to a specific user or workload identity, not a shared account.
logs sent to a separate account or system the logged parties can't modify.
correlation and alerting. See the SIEM pipeline.
integrity protection (hashing, signed digests, write-once storage) so logs stand up as evidence, and so a user can't credibly deny an action.
📰Real incident — Storm-0558 and the missing logs (2023). When attackers forged tokens to read government mailboxes in Microsoft's cloud, the US State Department spotted it through an alert on the
MailItemsAccessedaudit event. Many other customers couldn't have, because that event was only available in a premium logging tier. After public pressure, Microsoft expanded default logging for all customers. Lesson: in SaaS, the logs you didn't buy are the investigation you can't run.
🎯On the job. For each critical SaaS application, list which events you get (sign-in, admin change, data read, sharing), how long they're kept, and whether they reach your SIEM. Fill the gaps before the incident.
Security angle
SaaS logging is often a paid tier or limited to 30–90 days. Negotiate log access and retention in the contract, not after the incident.
Protecting AI and ML Data
Last verified2026-10 — added to the CCSP outline effective 1 August 2026.
What is this?
Protecting the data used to train, tune and run AI models, and the models themselves, which are a form of data derived from it.
Why it matters
Training data is often the largest concentration of sensitive data in an organisation, and models can leak what they were trained on.
How it works
know where every training dataset came from, under what consent or licence, and what transformations were applied.
remove or pseudonymise personal data before training where possible.
protect training data and pipelines from data poisoning (malicious samples that change model behaviour). Version datasets and restrict write access.
model weights are valuable intellectual property; protect them like source code and watch for model extraction through excessive API queries.
membership inference (was this person in the training data?) and model inversion (reconstructing training examples from outputs). Mitigations include differential privacy, output filtering and rate limiting.
prompts can contain sensitive data and outputs can leak it; apply DLP and logging to both, and check provider terms on whether your data is used for training.
removing a person's data from a trained model may require retraining; plan for it when using personal data.
📰Real incident — source code pasted into a chatbot (2023). Samsung engineers pasted confidential source code and meeting notes into a public AI chatbot to get help, sending them to a service whose terms then allowed using inputs for training. Samsung restricted generative-AI use on company devices. The leak wasn't a hack; it was the default data flow of a convenient tool.
📰Real incident — training-data extraction (2023). Researchers showed that asking a production chatbot to repeat a single word forever eventually made it emit memorised training data verbatim, including personal contact details. Models can leak what they were trained on.
🎯On the job. Before approving an AI tool, check: does the provider train on our inputs, where are prompts stored and for how long, can we enforce enterprise settings and DLP on prompts, and is there an audit log of who sent what.
Security angle
The training pipeline is a supply chain: datasets, pre-trained models and libraries pulled from public sources can all be poisoned. See the OWASP LLM Top 10.
Interview Questions
Create, store, use, share, archive and destroy. You classify at creation, encrypt and back up on storage, enforce access control and monitoring while in use, apply DLP, rights management and TLS when sharing, enforce retention and immutability in the archive, and crypto-shred on destruction. The use phase is the weak spot because data has to be decrypted to be processed, so access control and monitoring carry the load there.
Encryption transforms the value mathematically and anyone with the key can reverse it, whereas tokenisation swaps it for a random token with no mathematical link and keeps the real value in a separate vault. I'd choose tokenisation for things like card numbers, because systems that only ever see tokens can fall out of PCI DSS scope entirely. The trade-off is that the token vault becomes a crown jewel you must protect and monitor.
Yes. Pseudonymisation replaces identifiers but the data can be re-linked with separately held information, so it's still personal data — it just reduces risk and is recognised as a good security measure. Only truly anonymised data, where re-identification isn't reasonably possible even through indirect identifiers like postcode and birth date, falls outside GDPR, and that's much harder to achieve than people assume.
Crypto-shredding: encrypt the data under keys you control from day one, then destroy the keys so every copy — including replicas and backups encrypted under them — becomes unreadable. It only works if you planned the key hierarchy up front and know which keys cover which data. Keep a record of the key destruction as proof for auditors.
IRM wraps a document with encryption and a policy, so rules like view-only, no printing and expiry travel with the file, and the owner can revoke access after it's been shared, with every open logged. It's worth it for a small set of highly sensitive documents shared outside your control, like board packs or deal rooms. It's not a blanket control: it needs compatible software and can't stop someone photographing the screen.
Through membership inference, where an attacker tests whether a specific record was in the training set, and model inversion or regurgitation, where outputs reconstruct training examples. Defences start with minimising and de-identifying personal data before training, then differential privacy, output filtering and rate limiting against extraction. And because removing a person later may mean retraining, you decide what personal data goes in very deliberately.
Which events the vendor logs — sign-ins, admin changes, data access, sharing — whether you can export them to your SIEM by API, how long they're retained, and whether that's an extra-cost tier. In SaaS you only get what the vendor exposes, so if data-access logs aren't available, you won't be able to scope a breach. That needs to be in the contract, because after an incident it's too late.