AI Data Privacy
Applying GDPR and KVKK to AI
The practice of making sure personal data in AI systems, whether in training, prompts, logs or calls to foreign APIs, is processed lawfully under GDPR and Turkey's KVKK, for a specific purpose and with as little data as possible.
In an LLM application personal data doesn't stay in one place; it spreads into training data, prompts, RAG documents, logs and the provider's servers. Data protection law applies to each of those points separately.
The core principles are similar in both regimes: lawfulness, a specified legitimate purpose (purpose limitation), only what is necessary for that purpose (data minimisation), accuracy and limited retention. GDPR Article 5; KVKK (Law No. 6698) Article 4. In practice: if you don't need a customer's national ID number to train the support bot, that number should never reach the model.
Lawful basis: GDPR Article 6 lists six bases (contract, legal obligation, legitimate interests, consent...). KVKK Article 5 starts from explicit consent but also recognises conditions that don't need it, such as performance of a contract, legal obligation or legitimate interest. The European Data Protection Board's Opinion 28/2024 (17 December 2024) covers how to demonstrate legitimate interest for AI models and holds that a model trained on personal data cannot be considered anonymous in all cases.
Automated decisions: GDPR Article 22 gives people the right not to be subject to decisions based solely on automated processing that produce legal or similarly significant effects; even under the exceptions, the right to human intervention and to contest the decision remains. KVKK's counterpart is narrower: Article 11(1)(g) is a right to object to an adverse outcome arising from analysis exclusively by automated systems.
Impact assessment: GDPR Article 35 makes a DPIA mandatory for high-risk processing (e.g. systematic, extensive evaluation based on profiling that produces significant effects on people). Law No. 6698 has no general DPIA obligation, but running a similar assessment for high-risk AI projects is good practice.
Transfers abroad: Most LLM APIs are hosted abroad. KVKK Article 9 was rewritten by Law No. 7499 (in force since 1 June 2024): first an adequacy decision; failing that, appropriate safeguards (standard contract announced by the Board, binding corporate rules, or a written undertaking approved by the Board); failing those, only incidental transfers. A standard contract must be notified to the Authority within five business days of signing. GDPR follows the same logic in Chapter V (Articles 44-49).
Provider side: OpenAI, for example, states that since 1 March 2023 API data isn't used for training unless you opt in, and that abuse monitoring logs are kept for up to 30 days by default; Zero Data Retention requires prior approval. Every provider's terms differ, so read the contract yourself.
In November 2025 the Turkish DPA published "Generative AI and Personal Data Protection Guide (in 15 Questions)", the first thing to read for projects in Turkey. (As of October 2026; not legal advice.)
Like sending paperwork to an accounting firm. If you want them to file your tax return, you send income documents; you don't put medical reports and family photos in the folder. If the firm is abroad, you check the contract first: how long will they keep the papers, will they use them for anything else, who can see them? The LLM API is that firm, and the prompt is the folder you send.
An e-commerce company builds an LLM assistant that summarises customer complaints. Complaint texts contain names, phone numbers, emails, national ID numbers and card numbers. The resulting flow:
1. Purpose and basis: The purpose is "summarise the complaint and route it to the right team". No identity data is needed for that. 2. Redaction: Before the text goes to the API, Presidio scans it and replaces names, emails, phones, TCKN and card numbers with placeholders. 3. Provider: A data processing agreement is signed, training on customer data is off, retention is written into the contract, and the standard contract for the transfer abroad was notified to the Authority within five business days. 4. Logs: Application logs only hold the redacted text and ticket ID, deleted after 90 days. 5. Decision: The assistant doesn't decide refunds, it only summarises, so automated-decision rules aren't triggered.
The model does its job, and not a single national ID number reaches the provider's servers.
import os
from presidio_analyzer import AnalyzerEngine
from presidio_analyzer.predefined_recognizers import TrNationalIdRecognizer
from presidio_anonymizer import AnonymizerEngine
from presidio_anonymizer.entities import OperatorConfig
# The default engine uses the spaCy en_core_web_lg model.
analyzer = AnalyzerEngine()
# The Turkish ID recognizer isn't loaded by default; add it
# (it validates check digits, so random 11-digit numbers don't match)
analyzer.registry.add_recognizer(TrNationalIdRecognizer(supported_language="en"))
anonymizer = AnonymizerEngine()
ENTITIES = ["PERSON", "EMAIL_ADDRESS", "PHONE_NUMBER", "TR_NATIONAL_ID",
"CREDIT_CARD", "IBAN_CODE"]
def redact(text: str) -> str:
findings = analyzer.analyze(text=text, language="en",
entities=ENTITIES, score_threshold=0.4)
return anonymizer.anonymize(
text=text,
analyzer_results=findings,
operators={
"DEFAULT": OperatorConfig("replace", {"new_value": "<PII>"}),
"PERSON": OperatorConfig("replace", {"new_value": "<PERSON>"}),
"TR_NATIONAL_ID": OperatorConfig("replace", {"new_value": "<TCKN>"}),
},
).text
ticket = ("Ayse Yilmaz (ayse.yilmaz@example.com, +90 532 123 45 67, "
"TCKN 10000000146) says her refund for card 4111 1111 1111 1111 is late.")
safe = redact(ticket)
# <PERSON> (<PII>, <PII>, TCKN <TCKN>) says her refund for card <PII> is late.
# Only the redacted text goes to the provider
# (call_llm: your provider SDK call)
summary = call_llm(model=os.environ["LLM_MODEL"], prompt=f"Summarise: {safe}")- Whenever you send text, documents or images containing personal data to an LLM
- When training or fine-tuning a model on customer or employee data
- When indexing documents with personal data into a RAG system
- When AI makes or shapes decisions about people
- Building a heavy process for purely synthetic or genuinely anonymous data
- Treating redaction as the only safeguard; contracts, access control and retention policy are needed too
- Using this page as a legal opinion; borderline cases go to your data protection officer
Confusing anonymous with pseudonymous
A record with the name removed but customer ID, address and order history intact is still personal data. If it can be re-identified, data protection law still applies.
Assuming redaction catches everything
NER models can miss Turkish names, addresses and IDs buried in free text. Presidio's default NLP model is English; for Turkish you need extra recognizers and a recall rate measured on your own test set.
Forgetting the logs
Redacting the prompt but writing the raw text to app logs, error trackers or observability platforms is a common leak. Logs need the same retention and access rules.
Burying the transfer in the contract
Under KVKK Article 9, failing to notify the standard contract to the Authority within five business days is a separate ground for an administrative fine. Confirm from the contract where the provider processes data and who its sub-processors are.