A June 2025 analysis argues that artificial intelligence advances fastest where work already produces data, stable feedback and repeatable measures of success. For managers, the dividing line is less about whether a job looks creative or routine than whether its tasks can be observed, scored and improved at scale.
Harvard Business Review published the article by Christian Catalini, Jane Wu and Kevin Zhang on 19 June 2025. Its title is deliberately broad, but the argument is most useful as a framework rather than a deterministic law. Measurement can make a task easier to automate; it does not by itself prove that the output is safe, valuable or ready to replace an entire occupation.
Automation needs more than a capable model
The authors describe a three-part engine. A system needs examples or observations of the task, a reward or feedback signal that distinguishes a better result, and enough compute to iterate. If one element is weak, raw model capability may not translate into dependable work. A large archive with ambiguous outcomes is different from a clean record of actions and verified results.
This explains why operational traces matter. A customer-service workflow may record the question, response, resolution time and later complaint. A production process can capture sensor readings, defects and accepted parts. A design decision may leave files but no agreed score for whether the outcome was original, trusted or commercially durable. All three create data, but only the first two offer relatively immediate feedback.

A task is more exposed when it has
- many comparable examples captured under reasonably consistent conditions;
- an outcome that can be checked quickly and at low cost;
- a metric that reflects the real objective rather than an easy proxy;
- feedback that returns soon enough to improve the next attempt;
- limited consequences when an uncertain case is escalated or rejected;
- a workflow that software can observe and act upon without losing essential context.
ImageNet illustrates the power of a measurable contest
The authors' pre-publication draft uses ImageNet as an example. The large labeled image collection gave researchers a common training resource, while its competition and error measures made progress comparable. In 2012, a neural-network result changed expectations about computer vision because performance could be demonstrated against the same benchmark.
The lesson is not that every business needs a public leaderboard. It is that shared examples and a stable evaluation can turn vague improvement into an engineering loop. The caution is equally important: a benchmark can be optimized without covering rare cases, new environments or the consequences of a wrong decision. What is easy to score can crowd out what actually matters.
Cheaper sensing widens the automation frontier
Measurement once cost enough that companies reserved it for large problems. Lower-cost sensors, software logs and models running near devices can capture smaller events continuously. Synthetic data can also supply examples that are scarce, expensive or dangerous to collect in reality. These tools can make automation economical for tasks whose individual savings are tiny but occur millions of times.
However, synthetic examples inherit assumptions from the process that creates them. Sensors fail, logs omit informal work and users change behaviour when they know a metric is being watched. Before automating a newly measured task, a company needs to test whether the captured signal represents the environment and whether the evaluation detects harmful edge cases.
Jobs should be separated into tasks
Calling a profession “automatable” hides the variation inside it. An analyst may retrieve records, normalize data, calculate scenarios, explain assumptions and persuade a decision-maker. The first three tasks may have abundant examples and direct checks. The last two depend more heavily on context, trust and responsibility. Automation can therefore change the composition of a job long before it removes the role.
This task-level view also reveals new bottlenecks. If generation becomes cheap, verification may consume more expert attention. If junior workers no longer perform routine analysis, an employer must find another way to build the experience needed for later judgment. A short-term productivity gain can weaken the future talent pipeline unless training is redesigned.

Unknown probabilities create a different problem
The authors distinguish risk that can be modeled from Knightian uncertainty, where reliable probabilities cannot be assigned because the relevant events or outcomes are not yet known. Launching an unfamiliar category, responding to an unprecedented crisis or choosing a research direction may not provide enough repeated history to train or score a system in the usual way.
Humans do not automatically solve such problems well, but they can frame new questions, use analogy across domains and accept responsibility for a decision made without statistical comfort. As soon as a domain becomes measurable, part of that advantage can shrink. The boundary therefore moves rather than dividing work permanently into machine and human territory.
Management controls for the measurement frontier
- map roles into tasks and identify the data, feedback delay and consequence of error for each one;
- audit whether the chosen metric is a faithful objective or a proxy that can be gamed;
- measure verification time and exception volume as well as generation speed;
- preserve apprenticeship through reviewed practice, simulations and rotation into ambiguous cases;
- set escalation rules for low-confidence, novel or high-impact decisions;
- fund experiments whose value cannot yet be expressed by a precise near-term return.
What cannot be counted still needs management
Trust, taste, tacit knowledge and an ability to recognize a new problem are difficult to reduce to one target. That makes them easy to neglect in a measurement-heavy organization. Leaders can protect them without abandoning accountability by using portfolios, peer review, narrative evidence and staged experiments rather than forcing every idea into a premature score.
The June article offers a sharper automation question than asking which job titles disappear. Companies should ask where data and feedback already create a learning loop, where measurement is becoming cheaper, and where apparent precision hides missing context. The resulting map will change continuously, so the durable capability is not a one-time forecast but disciplined redesign of work and verification.



HOT NEWS INTERNATIONAL
Leave a comment