AI is doing more of the work. Your job is moving upstream.
For almost as long as we’ve been alive and working, being good at your job meant being able to produce the answer.
There was always someone doing the research, building the analysis, comparing the options and working through the problem until they had something worth putting out.
AI can now do a large part of that work in seconds. And that pushes more of the human contribution into what happens around the answer.
If AI drafts the recommendation, someone still has to decide whether the recommendation makes sense. If it summarizes the research, someone has to know what deserves a verification check. If it proposes a course of action, someone has to judge whether that action fits the situation and whether the consequences have been thought through.
AI is already moving closer to the point of decision. Deloitte’s 2026 Global Human Capital Trends research found that 60% of executives regularly use AI to support their decisions. Nearly two-thirds said using AI to support decision-making were very important to their organization’s current success, yet only 5% believed their organization was leading in this area.
So the challenge is getting more specific. It is no longer enough to ask whether people know how to use AI. It is whether they know how to judge what AI generates.
Human judgment in AI is the differentiator.

Coursera’s 2026 Job Skills Report, based on nearly six million enterprise learners across more than 7,000 organizations, found a sharp rise in critical-thinking enrollments among people learning AI-related skills. Among learners focused on GenAI, enrollments rose 185%.
Coursera describes the emerging human role as that of an “expert validator”: someone who can evaluate what AI produces rather than simply generate the output themselves. And that validation is proving harder than it sounds.
Why do two equally capable employees get very different results from the same AI?
Giving everyone the same AI does not give everyone the same performance.
A July 2026 field study from KPMG and the University of Texas tested 523 early-career professionals on realistic business tasks using the same AI agent.
- Only 50.1% improved on the AI-only output.
- Another 25.8% performed at roughly the same level as the AI
- while 24.1% actually made the result worse.
The researchers looked at whether stronger performers simply knew more. More AI literacy helped explain part of the picture. So did critical-thinking knowledge and expertise. But the actual difference was how well they applied those capabilities while working with the tool.

The gap between knowing and applying becomes especially important as people hand more first-pass thinking to AI. Q Studio has explored a related problem in its piece on how AI use can affect cognitive capacity, where passive reliance on AI can reduce the amount of independent reasoning people bring to the task.
Knowing that AI can make mistakes is different from catching one.
Knowing that assumptions should be tested is different from noticing the assumption concealed inside a recommendation.
Knowing what critical thinking means is different from doing it under time pressure when the output already looks comprehensive..
Human judgment is becoming the differentiator.
The harder part of the work is moving upstream.
It is now thinking, what should I be asking in the first place?
What deserves to be trusted?
What has been left out?
What changes when this answer is applied to context – in this situation, with these people, under these conditions?
And finally: would I put my name behind it?
AI can give more people access to strong analysis. However, it cannot remove the need for someone to decide whether that analysis deserves to become action.
That is judgment.
And if judgment is becoming the differentiator, the next question is obvious:
If AI can increasingly produce the answer, what does good human judgment actually require?
The judgment paradox
Now here’s the paradox: the better AI gets at producing convincing work, the easier it becomes to stop interrogating what it gives you.
A Microsoft and Carnegie Mellon University study surveyed 319 knowledge workers who used generative AI at least weekly and collected 936 examples of real AI-assisted tasks. There was a clear pattern, that the more confidence people had in AI’s ability to perform the task, the less likely they were to report engaging in critical thinking.
The reverse was also true. When people had greater confidence in their own ability to perform the task or evaluate the AI response, they reported more critical engagement.
That gives us the first part of the problem.
AI makes judgment more valuable because it produces more of the first-pass work for us. At the same time, confidence in that output can reduce the amount of scrutiny we bring to it.
The risk is – it’s easy to miss errors because the human is still part of the process. They read the answer. They review the recommendation. They make a few edits. They may approve the final output.
But being in the workflow does not tell us how much judgment actually happened.
Sometimes the work has simply shifted from thinking through the problem to processing what AI produced.
What’s the difference?
Processing asks, Does this look reasonable?
Judgment goes further:
What is this based on?
What would make it wrong?
What has been missed?
Would I have reached the same conclusion if I had not seen this answer first?
Those questions take more effort. They also depend on:
Recognizing that deeper scrutiny is needed in the first place
Having enough time or cognitive capacity to evaluate the output
Knowing enough about the subject matter to challenge the output
And that is where the next problem begins: judgment does not happen automatically just because a human is involved.
What does good human judgment actually require?
The KPMG-University of Texas study showed that everyone worked with the same AI technology. The difference showed up in what people did with the output. Some improved it. Some barely changed the result. Nearly a quarter made it worse. That last number is an interesting one. These were people who were brought into the workflow to add human value. In almost one in four cases, human involvement reduced the quality of the result.
That shows that judgment is not something you inherently have just because you are a human. It has to show up in the moment.
You have to notice that an answer needs checking. You need enough reason to challenge it. And you need enough knowledge to know where it could be weak.
The Microsoft and Carnegie Mellon research grouped barriers around judgment into three broad conditions: awareness, motivation, and ability.

- Do you recognize that this needs your judgment?
The first condition is awareness.
AI output is often delivered looking like a finished product. The writing is clean. The recommendation sounds coherent. The structure makes sense. Nothing in the response necessarily warrants any interrogation.
But judgment actually starts before evaluation of the output. It starts with recognizing that evaluation is needed.
A weak answer with an obvious error gets attention. A plausible answer is harder.
With a plausible AI answer, you may read it, make a few edits, and move on without asking whether the logic underneath it deserves the same confidence as the language on top.
This is where human oversight can become superficial. Someone is technically reviewing the work. But the review is happening at the surface.
- Do you have enough time and reason to interrogate it?
Then there is motivation.
Judgment takes effort. It takes longer to question an assumption than to accept it. It takes longer to check a source than to trust the summary. It takes longer to build an opposing case than to go with the answer that already sounds reasonable.
And under a time crunch or competing demands, that effort starts to look optional.
That creates a workplace problem that is becoming increasingly common nowadays. If someone has ten AI-assisted outputs to get through before lunch, “good enough” starts to win.
That is why judgment also depends partly on stakes.
What will this answer be used for?
Who will act on it?
What happens if it is wrong?
An internal brainstorm vis-a-vis a client recommendation may not need the same level of scrutiny. People need to be informed when deeper judgment is expected, rather than being left to decide after the output has been created.
- Do you know enough to tell when it is wrong?
Ability is the third condition.
And this may be the least comfortable one. You cannot challenge an answer well if you do not know enough to recognize what good looks like.
The Microsoft research found that workers had more difficulty critically improving AI output when they were operating in unfamiliar domains. Confidence in their own ability to perform the task and evaluate the AI response was associated with greater critical engagement.
That means AI can create a strange illusion of competence. The output may be sophisticated enough for you to use, while sitting just beyond your ability to evaluate properly.
You can understand the recommendation without knowing whether the assumptions behind it are sound. You can follow the analysis without knowing what evidence is missing. You can edit the language without realizing the underlying argument is weak.
This is why prompting expertise cannot replace subject knowledge.
And it is why delegating more of the first-pass thinking to AI can create a new responsibility for the person using it, which is having enough understanding to know when the answer deserves further evaluation.
Where this leaves us
The three conditions above sound simple enough: recognize when judgment is needed, have enough reason to interrogate the output, and know enough to challenge it.
But access to those conditions depends on something underneath them.
Enough cognitive capacity to notice the weak signal in the first place. Enough steadiness to resist the urge to move quickly because the answer already looks plausible. Enough attention to hold the context, the stakes, and your own assumptions in mind while you decide what to trust.
That is where judgment starts to become a human-performance issue.
A person can know all the right questions and still fail to ask them when they are rushed, overloaded, or mentally scattered.
Part 2 of this Judgment in an AI world series looks at what helps close that gap: how judgment can be trained, which thinking skills matter most, and how to practice them in real work so they are available when the pressure is on.
A few things are worth carrying into Part 2:
- The same AI can produce very different results in different hands. The difference is often how people apply what they know, not simply whether they can work with the tool.
- Confidence can reduce scrutiny. The more reliable AI feels, the easier it becomes to stop questioning the output closely.
- Judgment needs awareness, motivation, and ability. You have to recognize when deeper review is needed, have enough reason to do it, and know enough to challenge what you are seeing.
- Human oversight can still be superficial. A person can review, edit, and approve AI output without ever testing the reasoning underneath it.
That leaves us with: if human judgment is becoming more important, can people actually get better at it?
The answer is yes. But it takes more than telling employees to “think critically.”
In Part 2, we look at what judgment is made of, the Mind Skills that strengthen it, and how those skills can be practiced inside real work.
Read Part 2: How to Train Better Judgment in an AI World

