Imagine a worker at 4:47 on a Friday afternoon. An AI system has produced a confident recommendation. The dashboard is green. The deadline is close. Something about the answer feels wrong. What, precisely, is this person supposed to do next?
That small moment contains the central problem in IBM’s new workforce research. The September 21, 2026 study, conducted by the IBM Institute for Business Value with Oxford Economics, surveyed 1,500 chief human resources officers and equivalent executives across 21 geographies and 23 industries, plus 8,800 full-time employees in 28 countries. The headline finding is not that workers dislike AI. It is that companies are installing powerful tools faster than they are redesigning the human decisions surrounding them.
Sixty percent of employees worry that AI is eroding their skills. Among that worried group, three in four say the erosion has already begun. Critical thinking is the ability they most often say is declining. The paradox is sharp: the more capable the tool becomes, the more valuable human judgment should become—and the fewer ordinary chances people may get to practice it.
of employees worry AI is eroding their skills
say the blame falls on them when AI goes wrong
of CHROs say AI creates invisible work
The 71–29 gap
Executives and employees are looking at the same workplace and seeing two different jobs. Seventy-one percent of CHROs call the ability to supervise, validate and override AI outputs the workforce’s most essential skill. Only 29% of employees rank judgment as important. That 42-point spread is not a training gap in the usual sense. It is a mismatch in the definition of the work itself.
Who says judgment matters?
Share identifying AI supervision, validation or judgment as important
Leaders appear to imagine an employee as an alert co-pilot: skeptical, informed and ready to seize the controls. Employees may experience something closer to a passenger seat. They are told to use the system, move faster and hit the target. If the organization rewards speed and compliance, then careful resistance can look like delay. A worker does not learn that judgment matters from a slide deck. The lesson arrives through incentives, permission and what happens after someone says, “The model is wrong.”
IBM’s data makes the stakes visible. Forty-nine percent of employees and 57% of CHROs say critical thinking and problem framing are among the skills that matter most in the AI era. But naming a capability is easier than creating the conditions in which it can survive.
Responsibility sits with the human, while authority quietly migrates to the machine.
The blame arrives last
Now return to that Friday-afternoon recommendation. If the employee accepts it and the outcome fails, 43% of employees say the blame falls on them. Yet 41% of CHROs believe workers may not feel safe challenging or overriding AI output. This is the accountability trap: responsibility sits with the human, while authority quietly migrates to the machine.
The hidden labor compounds it. Eighty percent of CHROs say AI creates “invisible” work—validating recommendations, correcting mistakes, supplying context and managing exceptions. Forty-two percent of employees say AI increases their workload or that this extra work goes unrecognized. The machine can appear productive partly because the human cleanup crew is missing from the measurement.
When accountability is unclear, adoption becomes brittle. Thirty-six percent of CHROs say that ambiguity complicates AI deployment. Nearly half of organizations—46%—do not involve the CHRO when AI strategy is defined, even though the hardest questions are about roles, incentives and confidence rather than software alone. Only 28% of CHROs report a joint roadmap with IT supported by a shared operating cadence.

Draw the lanes before accelerating
The study’s most useful finding is almost stubbornly simple. Organizations that clearly classify workflows as human-led, AI-assisted or AI-executed report 18% lower risk and 20% higher quality. The advantage does not come from adding another model. It comes from drawing the lanes.
Human-led
People frame and own the decision; AI supplies evidence or administrative help.
AI-assisted
The system proposes or performs a step; a named person reviews the result.
AI-executed
The system acts within defined limits, monitoring and escalation rules.
Those labels force managers to answer questions that vague enthusiasm can conceal: Who owns the decision? Who checks it? When must a person intervene? What evidence can they see? What happens when they disagree?
The confidence effect is substantial. Where judgment is designed into the work, 62% of CHROs report growing employee confidence in AI-enabled decisions. Where it is not, 57% report declining confidence. And where the CHRO shares responsibility for deciding which choices remain human-led, 76% of employees feel safe questioning or overriding AI recommendations. When HR is merely advisory, that figure falls to 43%.
IBM tries the experiment on itself
IBM’s answer is not only a study. The company has been running an internal experiment it calls Client Zero: using its own operations as the proving ground for AI, hybrid cloud and automation. In the 2025 IBM watsonx Challenge, employees proposed 15,000 AI agents. The scale matters because the ideas came from people close to the friction—the repeated searches, handoffs, tickets and reconciliations that rarely appear in a strategy presentation.
The program paired bottom-up proposals with top-level sponsorship. According to IBM’s case study, a CEO-led steering committee reinforced accountability and cultural change, while teams worked in rapid sprints and measured initiatives against business outcomes. IBM says it designed more than 155 AI use cases across core functions. Its examples include an HR agent that resolves common inquiries, automation of routine IT support, procurement time savings and tools for sellers.

IBM reports that this broader transformation unlocked USD 4.5 billion in productivity gains over three years. That number is IBM’s own Client Zero result, not an independently audited finding of the CHRO study, and the case study cautions that examples are illustrative and results vary. It is still useful as a picture of the operating model IBM is advocating: executive sponsorship above, employee invention below and measurement in the middle.
Fluency without judgment simply helps an organization make mistakes faster.Dr. Amit Das · Director and CHRO, Bennett Coleman & Co.
Productivity is not the whole ledger
The temptation is to count the visible gain and ignore the invisible cost. A support ticket disappears; a manager saves 15 minutes; a seller finds an answer without switching among five applications. Those are real benefits. But if a worker spends unrecorded time checking a confident error, or stops exercising a skill because the machine always takes the first pass, the productivity ledger is incomplete.
This is why the IBM research lands differently from a conventional reskilling report. It suggests that skill erosion is not solved by sending everyone to a course. Judgment is developed through use. People need repeated exposure to ambiguous choices, access to the information required to evaluate them, and organizational permission to be inconvenient. A company cannot automate away the practice and then demand mastery when the exceptional case arrives.
The better metric is not merely tasks completed. It is decision quality: whether the human-machine system catches errors, learns from exceptions and keeps the people inside it capable. That is slower to measure than hours saved. It is also closer to the risk leaders say they care about.
The new job description
Every important AI workflow now needs two descriptions. The first says what the system can do. The second says what the person must still notice. Most companies have invested heavily in the first and treated the second as common sense. IBM’s numbers show the cost of that assumption.
The redesign can begin with a plain-language map. Label each consequential workflow. Name the human owner. Define the evidence required before approval. Make override routes visible and record overrides as learning signals, not acts of disobedience. Count validation and exception handling as work. Reward people who catch plausible mistakes before those mistakes travel downstream. Then test whether workers still understand the underlying task when the tool is unavailable or wrong.
The point is not to preserve every old skill. Technologies have always changed what people remember and practice. The point is to choose deliberately which capabilities may fade and which must become stronger. If judgment is the safeguard, it cannot be an emergency lever behind glass. It has to be part of the daily job.
A machine can answer; a company must decide
The AI debate often turns on a theatrical question: can the machine do the task? The more consequential question is quieter: has the organization designed what happens around the answer?
IBM’s study offers both warning and direction. Workers already fear erosion. Leaders already know judgment matters. Employees already encounter blame without the confidence to override. Yet organizations that assign clear roles to people and machines report better quality and lower risk. The gap, then, is not inevitable. It is designed—and it can be redesigned.
At 4:47 on Friday, the employee should not need courage to challenge a machine. They should have a mandate, evidence, time and a known path for escalation. The quality of the AI matters. But the quality of the surrounding job may matter more.
Primary reporting
IBM Institute for Business Value CHRO study announcement · September 21, 2026
IBM Client Zero case study · Company-reported transformation results