The Work Beneath Judgement
The work may have been doing work on us.
The work beneath the work
There are things I know now that nobody ever really taught me.
I know when an estimate feels too clean. I know when a beautifully written paper is carrying a weak argument. I know that a programme can be green on Friday and in serious trouble by Tuesday. I know that sometimes the sentence everybody agrees with in the room is the sentence that deserves the most suspicion.
None of this arrived as theory. I acquired it badly: through analysis that did not survive contact with the evidence, documents I thought were finished until somebody asked the question I had not considered, estimates that contained more confidence than understanding, and decisions whose consequences did not disappear when the meeting ended.
Something accumulates through those experiences. Not knowledge exactly. More like sediment: layers of error, correction, embarrassment, pattern and consequence that eventually become judgement.
For most of my career, I assumed the work was the work and the learning happened alongside it.
I am beginning to wonder whether I had that backwards.
Perhaps more of the work than we realised was the learning.
That matters because we have become exceptionally good at removing the work.
And much of it deserves to go.
There is nothing inherently developmental about spending half a day assembling material a machine can produce in seconds. Nobody becomes a better leader by manually reformatting another spreadsheet. A first draft is not morally superior because it took four hours instead of forty seconds. Some of what earlier generations called apprenticeship was simply inefficient work performed by junior people because senior people could make them do it.
We should be suspicious of nostalgia masquerading as wisdom.
There is also an obvious historical objection to the anxiety around AI. We have heard this argument before. The calculator would destroy numeracy. The spreadsheet would create managers who no longer understood the numbers. Search would destroy memory. The internet would make everybody shallow.
The tools changed. Some capabilities moved. New ones appeared. People stopped doing things machines could do faster or better, and professional judgement somehow survived.
Perhaps AI is simply the next version of that story.
Perhaps experienced practitioners are looking at a technology that makes the next generation faster and interpreting the disappearance of familiar struggle as the disappearance of learning.
That possibility needs to remain alive, because otherwise this becomes an elaborate defence of the apprenticeship I happened to have.
I have no interest in defending that.
But there is a question underneath it that I do not think the calculator analogy answers.
“What must a person have done before they can safely stop doing it?”
A calculator removed the act of calculation. A spreadsheet removed a great deal of mechanical modelling. Search removed the need to remember where information lived.
Generative AI can enter much earlier.
Give it the ambiguous problem. Ask it how to frame the issue. Ask what matters. Ask for the options. Ask it to structure the analysis, challenge the assumptions, produce the recommendation, write the difficult document, critique its own argument and then make the whole thing sound more senior.
The important difference is not that AI is cleverer than a spreadsheet.
It is that AI can increasingly participate in the encounter through which the practitioner would previously have discovered what they did not understand.
And that changes the question.
The apprenticeship we did not name
We tend to describe early-career work by its output.
The analyst produces the analysis. The consultant produces the deck. The programme manager produces the plan. The associate produces the first draft.
Seen that way, the productivity case for AI is almost embarrassingly straightforward. Reduce the production cost. Increase the quality. Move people towards higher-value work.
But perhaps we have described the job too narrowly.
The analyst was not only producing analysis. They were learning what weak evidence feels like when an argument depends on it. The person wrestling with the paper was discovering which ideas remained convincing only while they were allowed to stay vague. The programme manager building the estimate was learning that apparently small assumptions can carry enormous consequences.
The artefact was visible. The formation was not.
That distinction matters because organisations are very good at valuing the first and surprisingly bad at noticing the second.
The formation loop
Work → prediction → error → consequence → correction → judgement
The task matters because it repeatedly exposes the practitioner to the gap between what they expected and what actually happened.
Not every task creates that loop. Plenty of work is just work. Repetition alone does not create judgement, and difficulty is not evidence of developmental value.
But some work forces the practitioner to make a model of reality before reality answers back.
An estimate does that. So does a recommendation. So does a difficult diagnosis. So does writing an argument clearly enough that somebody else can attack it.
When the answer arrives before the practitioner has built the model, something changes.
We save the effort.
We may also lose the correction.
That is the possibility I think we have not taken seriously enough.
The lag we will not see
The difficulty is that the two effects arrive on different clocks.
Productivity arrives immediately.
If an early-career practitioner can produce in two hours what previously took a day, everybody can see the gain. The organisation gets more output. The practitioner appears more capable. The manager has more capacity. The case writes itself.
If some part of professional formation has been lost, the evidence may not appear for years.
It appears when that practitioner is no longer being asked to make the first draft but to decide whether somebody else’s first draft can be trusted.
When they are no longer building the estimate but approving the investment built on top of it.
When they are no longer conducting the analysis but deciding whether the analysis has missed the point.
By then, they may have become exceptionally good at the visible signals through which organisations recognise professional maturity: speed, polish, range, responsiveness, fluency across unfamiliar material.
AI strengthens exactly those signals.
This is the second-order consequence that concerns me more than the loss of any individual task.
We may begin promoting people on signals that used to correlate with experience after changing the process that created the experience.
The output improves first.
The career follows the output.
Only later do we discover whether the judgement followed too.
This is where the word hollow becomes useful, although I would use it carefully.
It does not mean unintelligent. It certainly does not mean young. Some of the hollowest leaders I have met were formed long before generative AI existed.
Hollow describes a gap between visible capability and the structure underneath it.
A practitioner can improve an argument but struggles to create an independent frame before seeing one. They can review an answer but cannot reliably reconstruct how the answer should have been reached. They can request alternatives without recognising which assumption makes each alternative dangerous. Their confidence remains high for too long as the problem moves outside familiar territory.
Everything looks impressive until something genuinely unfamiliar happens.
Then the borrowed structure runs out.
The supervision problem
The reassuring answer is that none of this matters because the new professional skill is supervision.
People will not need to produce everything themselves. They will become excellent at directing AI, interrogating its work and deciding what to trust.
I hope that is right.
But supervision contains a dependency that is easy to overlook.
To supervise something well, you need some model of how it fails.
If I have never struggled to construct the analysis, what tells me where an analysis tends to hide its weakness?
If I have never watched an estimate collapse, what gives me the instinct to distrust the perfectly plausible estimate in front of me?
If I have never tried to turn confused thinking into a coherent document, how do I know whether coherence in the document represents clarity in the thinking?
This is not an argument for human exceptionalism. AI will get better at detecting error, challenging assumptions, modelling uncertainty and recognising organisational patterns. Betting the future of leadership on machines never understanding context or politics seems increasingly unserious.
The issue is not whether AI can judge.
The issue is how a human develops the capability required to remain accountable for judgement when more and more of its production has been delegated.
Leadership eventually reaches that point.
The evidence conflicts. The pattern changes. The incentives matter. Several answers are plausible. The system recommends one.
Somebody still has to decide what to believe, what to commit to and what consequence they are prepared to carry.
The ability to make that decision cannot safely be inferred from the quality of the material surrounding it.
“The danger is not delegation. It is delegation before formation.”
That is the narrower claim I am prepared to defend.
Not that people must continue doing work machines can do.
Not that younger practitioners need to suffer through the same inefficiencies as earlier generations.
And not that AI necessarily weakens judgement.
The risk is that we delegate the work before understanding which parts of it were creating the capability later required to govern the delegation.
The strongest case against this
There is a serious objection, and it goes further than saying tools have always changed work.
AI may create a better apprenticeship.
A practitioner can make a prediction before asking the machine. They can compare their reasoning with another frame immediately. They can expose themselves to counterarguments that a busy manager would never have generated. They can simulate edge cases, ask why an assumption matters, reconstruct a decision from first principles and receive explanations at the moment confusion occurs.
Someone using AI that way might experience more deliberate cycles of prediction, challenge and correction in a month than I encountered in a year.
That would not be the destruction of apprenticeship.
It would be its acceleration.
And if that is what happens, much of this provocation falls away.
The old estimate, the painful first draft and the analysis that collapsed under review were never sacred. They were simply the machinery available to us for producing judgement.
AI may give us better machinery.
The mistake would be assuming that it does so automatically.
The incentives currently point elsewhere.
Organisations are rewarded for extracting productivity. A junior practitioner becomes faster, so the gain is captured. Fewer people are required, so the operating model adjusts. Work moves upward earlier because the person appears ready for it.
Nobody receives a warning saying that a useful developmental collision failed to occur.
We may therefore make a perfectly rational sequence of local decisions and create a formation problem nobody intended.
Remove the first attempt.
Remove the initial analysis.
Remove the uncomfortable reconstruction after the mistake.
Remove the necessity of committing to an answer before seeing a better one.
Then promote on the quality of the resulting work.
If a capability gap appears later, it will be tempting to blame the people who came through the system.
That would be convenient.
We built the system.
What would change my mind
A provocation should leave something that can be attacked.
If delegating the underlying work does not weaken judgement, the difference should be visible.
Practitioners who rely heavily on AI should still be able to detect consequential flaws in convincing work without asking the machine to find them. They should be able to reconstruct the reasoning beneath an output they supervised but did not produce. They should form an independent frame before seeing the machine’s frame. Their confidence should adjust when the problem moves outside familiar territory. Learning should transfer into genuinely novel situations.
And when the tool disappears, speed may fall. Judgement should not.
If AI-assisted practitioners match or outperform people who performed more of the underlying work themselves on those dimensions, I will change my mind.
If they develop those capabilities faster, the conclusion becomes more interesting still.
It would mean we confused apprenticeship with its historical form.
The tedious analysis, difficult document, failed estimate and painful mistake were not the source of judgement. They were one route to it.
That would be good news. We could remove the old work without preserving it sentimentally and design a better formation system deliberately.
But if the opposite happens—if the artefacts improve while independent reconstruction, error detection, calibration and transfer weaken—then the productivity story will have hidden a much larger exchange.
We will have removed work we thought was merely production and discovered, years later, that it was also producing the person.
The work may have been doing work on us.