EU AI Act · GDPR · Retention
An AI system produces data that two bodies of law claim at the same time. The AI Act requires it to be kept. The GDPR requires it to be deleted. Both mean the same log file.
Most retention and deletion policies do not mention AI at all. They are organised by business process — personnel files, accounting, job applications — and were written before the company ran a system that logs every use automatically. The structure is built for records somebody creates deliberately, not for data holdings that accumulate as a by-product of operations.
That is exactly what high-risk AI systems produce. The logs they generate are not a technical artefact that IT may rotate at its own discretion, but an obligation with a retention period of its own. Carrying an existing deletion policy forward unchanged therefore means breaching one rule or the other sooner or later — and typically noticing only once somebody asks.
The way out is unglamorous: treat an AI system as what it is in data protection terms — not one holding with one period, but four holdings with four periods.
What the AI Act requires you to keep
The regulation contains four retention obligations with very different periods, and they do not fall on providers and deployers equally. The distinction is not academic: a company that buys a system and runs it carries a markedly shorter list than one that develops it or puts its own name on it.
The first and second rows concern the same kind of data but two different controllers. The provider holds the logs its system generates and that accrue on its side. The deployer holds those it actually controls — for locally installed software that is everything, for a hosted service possibly very little or nothing at all. This boundary is precisely where most policies turn vague, because nobody has established where the logs physically sit.
And what sits in such a log is no abstract event journal. It typically records the time of use, the identifier of the operator, the input data or a reference to it, the system's output and — for high-risk applications — the details meant to make a decision traceable after the fact. A log is therefore not metadata but often the densest trace a system leaves behind.
Important for deployers: six months is a floor, not a target. The regulation ties the period to the system's intended purpose and expressly leaves room for longer periods under other law. Entering six months as a deletion deadline reads the provision backwards — and will be hard to justify to a market surveillance authority.
What the GDPR requires you to delete
On the other side stands the storage limitation principle in Art. 5(1)(e) GDPR: personal data may be kept in a form permitting identification only for as long as the purpose requires. On top of that comes the data subject's right to erasure under Art. 17 — and the accountability principle, which requires not only meeting your own retention periods but being able to evidence them.
The point that tends to get lost where AI is concerned: logs are regularly personal data. A log recording which input produced which output, and who operated the system, contains data on at least two groups of people — the operator and the person affected. That puts it squarely inside the GDPR, with everything attached: a legal basis, purpose limitation, data subject rights, an entry in the record of processing activities.
In practice the log file is therefore the holding most often overlooked. It does not appear in the record of processing activities, because it is treated as a technical matter. It does not appear in the deletion policy, because it belongs to nobody's business process. And it does not surface on a subject access request, because nobody has it in mind as a source of personal data. Three gaps resting on the same misjudgement.
Where the two collide, and how that resolves
The conflict is real, but the law has already settled it. Art. 17(3)(b) GDPR withdraws the erasure obligation where processing is necessary for compliance with a legal obligation to which the controller is subject. The logging duty under the AI Act is such an obligation. Retaining logs because the AI Act requires it does not breach the GDPR — it uses an exception the GDPR itself provides.
The limit of the exception
The exception reaches exactly as far as the obligation does. It covers the logs, not the system's entire data holdings. Refusing an erasure request wholesale by pointing at the AI Act overstretches it — and invites the question of why a training dataset or a customer history should be covered by a logging duty. That is the point where a defensible position turns into an exposed one.
In practice that means an erasure request is not answered with yes or no, but category by category. The sequence that works has four steps.
Establish which holdings the person appears in
This is the step that takes weeks without a system register. The question is not "is this person a customer" but: do they appear in inputs, in outputs, in logs, in a training dataset? Four questions, four possible locations.
Separate out the holdings covered by an obligation
Logs within a running retention period stay. Everything else is deleted. The order matters: first determine which obligation actually applies, then hold data back — not the reverse, looking for a reason to keep a holding you wanted to keep anyway.
Restrict what you retain
Data kept under a legal obligation is no longer used for the original purposes. From that point the logs serve only the evidentiary function they are retained for — implemented through restricted access rights, not through a statement of intent.
Tell the person what remains and why
Not "we have deleted your data", but: these categories are gone, this one remains until this date, on this legal basis. That is the answer that withstands a complaint — and it is far easier to give when the periods were fixed in advance.
Four categories of data, four periods
A policy that treats an AI system as one thing will be wrong by construction. The four categories follow different logics, and only one of them is protected from deletion by law.
Logs
Statutory floor of six months, ceiling set by the purpose. This is decided by the legal basis, not by the business unit. Keeping them longer requires a named reason — ongoing proceedings, an open complaint, a sector-specific retention duty. "Might come in handy" is not such a reason, and indefinite retention turns the storage limitation principle into its opposite.
Inputs and outputs in day-to-day use
Prompts, uploaded documents, generated text and assessments. No separate retention obligation under the AI Act attaches to them, which makes them the fastest category to delete — and in practice the one that lingers longest, because nobody treats it as a data holding of its own. With assistant-type systems this becomes the largest holding of all: every uploaded file and every pasted excerpt stays in the history, often indefinitely and visible to everyone with access to the same workspace.
Training and test data
Relevant only if you train or fine-tune yourself. Requirements on data quality and provenance do not justify indefinite retention. The question to settle first is whether production inputs flow into this holding at all — often that is a provider default rather than a company decision. And this category has a property the other three lack: what has once gone into a model cannot be deleted row by row afterwards. Whether inputs are used for training is therefore not a detail but the decisive fork.
System documentation
Risk classification, role assignment, training records, conformity paperwork. Ten years for providers, and nothing a deployer should let disappear at the end of a project either — it is the evidence that a decision was taken, and that is needed retrospectively. This category is usually not personal data, with one exception: training records carry names, and the usual employment-law periods apply to those.
The SaaS case: when the data is not in the building
With hosted systems the question shifts from "when do we delete" to "who deletes at all". The deployer remains the controller under data protection law but has no direct access to the holdings. What it has is a contract.
The processing agreement under Art. 28 GDPR is therefore where a deletion policy for SaaS systems is actually enforced. Three points inside it decide whether it works: the binding of the processor to instructions on deletion, the turnaround within which the provider carries a deletion out, and the arrangement for the end of the contract — return or deletion, and in what format.
AI systems add a wrinkle classic software did not have. Where the provider uses inputs for its own purposes — improving its model, say — it is no longer acting as a processor in that respect. There is then no processing on instructions but processing of its own, for which it is the controller and needs its own legal basis. That boundary often runs through an account setting rather than through the contract text.
Practical note: one email before a system is introduced settles what otherwise costs weeks afterwards: which logs does the system produce, where are they stored, for how long, who has access, and how is a deletion triggered? A provider who will not answer those four questions in writing will not answer them when it matters either.
What the policy actually has to say
Three entries per AI system are enough to turn a statement of intent into something that survives an audit. They sound banal, which is exactly why they are usually missing.
Where the data sits
With the provider, in your own infrastructure, or both. For SaaS systems this is the question most often left unanswered — and without an answer no deletion can be promised, neither to a data subject nor to an authority.
Who deletes
Automatically by the system, manually by a named role, or on request to the provider. The third case needs a contractually guaranteed turnaround, otherwise it is not deletion but a request. And the second needs a deputy: a deadline resting on one person keeps running while they are on leave.
What the clock hangs on
The creation date, the end of the business relationship, or decommissioning of the system. Only the first can be automated; the other two need an event that somebody reports.
That last point deserves its own remark, because it surfaces too late with some regularity. When an AI system is replaced, its logging obligations do not leave with it. The periods keep running, on data holdings for which no live system remains to manage them. Anyone planning a shutdown therefore has to decide where the logs will sit until their period expires, who still has access, and who deletes them at the end. Skip that and an orphaned holding remains — formally still subject to the GDPR, practically looked after by nobody.
The actual point
A deletion policy for AI is not a new document. It is a missing column in the existing one. What is missing is the recognition that an AI system is not one data holding with one period, but four holdings with four periods — one of which is protected from deletion by law, and one of which, the training data, cannot be cleanly purged after the fact at all.
The effort of writing that down once is modest: four rows per system, plus settling with the provider where the logs sit. The effort of not having done it always lands at the worst possible moment — an erasure request with a one-month deadline, or an information request from the market surveillance authority.
Get it right and you answer two questions that otherwise arrive separately and create work separately: the data protection officer's question about retention, and the authority's question about logs. It is the same table.
With SimpleAct
Retention periods on the system, not in a spreadsheet.
EU AI Act and GDPR in one platform: register AI systems centrally, classify them rule-based, assign responsibilities and export audit-ready evidence at any time. Made in Germany, hosted in Germany.
Start your AI inventory →This article is for general information only and does not constitute legal advice. Last updated: 22 September 2026.
Tags
Ready to put EU AI Act compliance on autopilot?
SimpleAct helps you inventory AI systems, classify risk, and generate the required documentation automatically.
Kamill Jarzebowski | SimpleAct
Author · SimpleAct Team
