GDPR Did Not Create the Data Governance Crisis. It Exposed It.
Good governance reduces decision latency because it settles who decides before the difficult case arrives.
The request nobody could answer
At 9.10 on a Monday morning, a data protection officer forwards a customer access request to six functions. The request appears simple: show what information is held, where it came from, why it is used and with whom it has been shared.
Eighteen days later, the response is still being assembled. The Article 30 record names Customer Operations as the owner of the processing activity. In practice, the relevant data sits across a customer relationship system, a billing platform, a campaign database, an archive maintained by the records team and two analyst spreadsheets. One nightly feed was written in 2012 and is understood only by a contractor. A field labelled “consent” has three meanings depending on which system is consulted. Marketing can explain what it wants to do with the data; Technology can explain where some of it moves; Legal can explain the lawful basis. Nobody can explain the whole chain.
This is not an unusual failure of diligence. It is the predictable result of ownership that exists in policy but disappears at the point of decision.
As the first anniversary of the General Data Protection Regulation’s application approaches, many organisations are preparing to declare the compliance programme complete. Privacy notices have been rewritten. Processor contracts have been reviewed. Retention schedules have been debated. Data subject request procedures have been rehearsed. Those were necessary tasks, often completed under considerable pressure.
But the most important lesson is being missed. GDPR did not create the enterprise data governance crisis. It exposed it.
Compliance finished where governance should have begun
The common programme pattern was understandable. A fixed deadline encouraged mobilisation around visible deliverables: records of processing, consent wording, contract clauses, impact-assessment templates and breach procedures. Progress could be counted and reported. Senior committees saw traffic lights, completion percentages and lists of residual actions.
The mechanism was compliance by collection. A central team asked business units to describe their processing, recorded the answers and chased missing evidence. For a few months, privacy lawyers, information security specialists, records managers, system owners and operational managers worked together because the programme required them to. When the deadline passed, that temporary coordination was treated as the finished product.
Governance is different. It is not the collection of answers; it is the continuing ability to produce a reliable answer when the purpose, data, system or model changes.
A record of processing is not governance; it is evidence that governance ought to exist.
The difference becomes visible in four recurring symptoms:
- The named owner cannot make the decision. A senior manager appears in the register but cannot authorise a change to retention, access or use without negotiating across several functions.
- The system boundary is mistaken for the data boundary. Each application is documented separately, while the same customer attribute is copied, transformed and interpreted differently across the estate.
- The control is performed after commitment. A data protection impact assessment begins when the supplier is selected, the model is built or the campaign date is fixed, leaving only cosmetic choices.
- The data protection officer becomes the owner by default. Every ambiguity is routed to the privacy function, even though that function should advise, monitor and challenge rather than decide the commercial purpose of processing.
The real test of data ownership is not whether a name appears in a register. It is whether that person can stop, change or retire a use of data before harm, delay or regulatory exposure is created.
These symptoms existed before GDPR. The Regulation simply connected them to rights, deadlines and demonstrable accountability. A subject access request crosses organisational boundaries that an internal policy can ignore. A request for erasure forces the organisation to discover copies and dependencies. A challenge to an automated decision forces it to connect the model to its inputs, purpose and human oversight. Rights turn abstract governance into an operational test.
Machine learning is widening the gap
The next pressure is already visible. Enterprises are moving predictive models out of specialist teams and into ordinary decisions: fraud referral, credit risk, customer retention, claims prioritisation and campaign selection. The mathematics may be sophisticated, but the governance failure is usually prosaic. The model depends on data whose provenance, meaning and permitted use were never settled.
Consider a composite but recognisable example. A regulated services business develops a customer-retention model using 96 variables drawn from eleven sources. The model performs well in testing. Four weeks before deployment, an impact assessment reveals that 27 variables have no agreed business owner, nine are derived from operational codes whose meanings changed over time, and one web-behaviour feed is retained indefinitely because no team has authority to delete it. The data scientists can reproduce the model score, but they cannot establish why every input exists in the training set or whether its current use is compatible with the purpose for which it was collected.
Deployment pauses for eight weeks. The delay is blamed on privacy review. In reality, privacy review discovered decisions that the organisation had deferred for years.
This sequence matters because it shows the mechanism. Machine learning concentrates many small data ambiguities into one consequential decision. A poorly defined field in a management report may cause argument; the same field in a scoring model can affect thousands of customers before anyone notices. The more data a model consumes, the less credible it becomes to govern the model while leaving its inputs ownerless.
The emerging distinction is not between “regulated data” and everything else. It is between data that can be traced to an accountable purpose and data that is merely available.
| Question | Paper governance | Operational governance |
|---|---|---|
| Who owns the data? | A name in a policy or register | A person with authority over purpose, access, quality, retention and retirement |
| Where did it come from? | The source application is listed | Transformations, copies and derived fields can be traced |
| May we use it? | Legal review is sought near launch | Purpose and lawful basis are tested before design is committed |
| Who owns the model? | The analytical team maintains it | A business sponsor owns the decision and its consequences |
| When does control occur? | Evidence is assembled for review | Decision rights operate whenever use changes |
The serious objection: governance can become its own risk
There is a strong practical objection to this argument. Organisations faced an immovable deadline and had to prioritise evidence. They cannot now halt every project until a perfect enterprise data catalogue exists. Adding owners, stewards, forums and approval gates can create a bureaucracy that protects the organisation from making decisions at all. Line managers already carry accountability; another governance structure may only diffuse it.
That objection is correct about the danger and wrong about the remedy.
The answer is not a grand data-governance programme that attempts to classify every field before useful work can proceed. Nor is it a committee between every team and every decision. Both approaches confuse completeness with control. The answer is to make a small number of decision rights explicit for the data that matters most.
For each material dataset or processing activity, four questions must have an answer owned outside the privacy function:
- Purpose: Which business executive is authorised to decide why the data is used and to reject a proposed secondary use?
- Provenance: Which operational steward can explain its meaning, origin, transformations and known limitations?
- Permissible use: Which control functions must advise or challenge before the use is committed, particularly where profiling or significant automated decisions are involved?
- Retirement: Who can order deletion, archiving or model withdrawal, and who can execute that decision across dependent systems?
This is not extra administration. It shortens the route to a defensible decision. In the rights-request example, the delay was not caused by excessive governance; it was caused by the absence of a person able to resolve conflicting definitions. In the model example, the late review did not create uncertainty; it revealed uncertainty after most of the cost had already been incurred.
Good governance reduces decision latency because it settles who decides before the difficult case arrives.
Use the portfolio to reset accountability
Individual projects cannot repair this alone. Their incentives favour delivery within scope, cost and time; inherited ambiguity is usually logged as a dependency and passed onward. The portfolio is where recurring data failures can be seen as one management problem rather than dozens of local exceptions.
The reset should be selective and practical.
- Reopen the highest-consequence processing first. Prioritise activities with large data volumes, frequent rights requests, sensitive information, extensive sharing, profiling or significant automated decisions. Do not begin with an indiscriminate enterprise inventory.
- Convert records into decision maps. For each priority activity, attach the purpose owner, operational steward, key systems, material transfers, retention decision and escalation route to the Article 30 record. The record should point to living accountability, not substitute for it.
- Move privacy and data-quality judgement forward. Require the purpose, provenance and permitted-use questions to be answered before a model is trained, a supplier is contracted or a campaign design is fixed. A gate is useful only while there remains a genuine choice.
- Fund the unglamorous repair. Broken metadata, undocumented interfaces and inconsistent reference codes rarely attract sponsorship, yet they determine whether rights can be honoured and models trusted. Portfolio funding should follow repeated control failure, not only visible innovation.
- Measure exceptions, not documents. Track unowned material fields, unresolved lineage breaks, late impact assessments, repeated manual searches and time taken to answer rights requests. These measures reveal whether governance is becoming operational.
The practical ambition is not a single, flawless version of all enterprise data. In a complex estate, that promise is usually expensive theatre. The ambition is more disciplined: for every consequential use, the organisation should know whose purpose is being served, what data supports it, what limitations travel with that data and who has authority to intervene.
GDPR has given enterprises a rare forcing mechanism. It has made invisible dependencies visible, converted vague accountability into answerable rights and exposed the cost of data that is copied more readily than it is governed. Treating this as a completed legal programme would waste that advantage.
The organisations that benefit will not be those with the most polished privacy documentation. They will be those that use the Regulation to replace ownership on paper with authority in practice.