In early 2026, Meta pursued an ambitious plan to reorganize parts of its workforce around AI. Project OT, short for Organization Transformation, envisioned smaller teams supported by AI agents, with some groups potentially shrinking by as much as 60 percent.
The company moved forward with layoffs and organizational changes, but internal evidence showed that the technology hadn’t produced the expected productivity. Technical disruptions increased, employees pushed back, and Meta canceled a planned second round of workforce reductions. By July, Mark Zuckerberg acknowledged that agent development had progressed more slowly than expected. A Reuters’ investigation showed what happened when expectations for AI agents moved faster than their demonstrated performance.
I didn’t see Meta’s experience as an argument against AI agents. I saw it as evidence that companies had treated agents as technology deployments when they should’ve treated them as a new category of worker. Organizations gave agents assignments, system access, data, and decision-making authority. They didn’t always give them defined responsibilities, measurable standards, ongoing supervision, or structured performance reviews.
Deployment Isn’t the Finish Line
Traditional software either executes its programmed logic or produces an identifiable error. AI agents operate differently. They interpret objectives, select tools, make intermediate decisions, and adjust their actions as conditions change. Two executions of the same assignment can produce different results.
That variability makes uptime, speed, and token usage inadequate measures of success. An agent could remain available, respond quickly, and complete thousands of actions while still damaging the business. It might close support cases without resolving customers’ problems, generate sales leads that never converted, or produce code that increased the review burden on engineers.
The 2026 APEX-Agents benchmark evaluated agents across 480 realistic tasks in investment banking, management consulting, and corporate law. The agents worked across documents, spreadsheets, email, calendars, and other workplace tools. When the benchmark was released, even the strongest systems completed fewer than one-quarter of the professional assignments successfully on their first attempt.
A separate 2026 study, Towards a Science of AI Agent Reliability, evaluated 14 models using 12 measures across consistency, robustness, predictability, and safety. The researchers found that capability improvements had produced only modest gains in reliability. A model could perform well on average and still fail unpredictably across repeated runs.
I believe that distinction matters. Capability tells me what an agent could do under favorable conditions. Performance management tells me what it did consistently inside my business.
Every Agent Needs a Performance Contract
I would have started each deployment with a performance contract. That doesn’t mean an employment agreement. It means a precise definition of the agent’s job, the results it was expected to produce, the decisions it could make, and the conditions that required human intervention.
A customer-service agent, for instance, shouldn’t have been evaluated by the number of conversations it handled. I would’ve measured verified resolution, repeat contacts, escalation quality, customer satisfaction, policy compliance, and the cost per completed resolution. Those measures would’ve revealed whether the agent created value or just activity.
I also would have evaluated the path the agent took. An agent might reach the correct outcome after accessing unnecessary data, calling the wrong tools, repeating expensive actions, or bypassing an approval. The result could look successful while the process introduced risk. Performance management had to examine both the outcome and the conduct that produced it.
Agents Need Reviews and Consequences
Human performance management doesn’t end with an annual rating, and agent performance management shouldn’t end with a pre-deployment test. I would have reviewed agents continuously because models, data, tools, customer behavior, and business rules kept changing.
Those reviews would have compared actual results with defined thresholds. They would have identified recurring failure patterns, tested performance after system changes, and examined cases in which humans corrected or reversed an agent’s work. The findings could then guide better instructions, narrower permissions, additional training data, revised workflows, or a different model.
Agents need consequences. A high-performing agent could receive broader authority. An inconsistent one might require more approvals. An agent that violated safety or compliance limits might be suspended until the failure was understood. Autonomy should be earned and retained through demonstrated performance.
I expect AI agents to become an important part of how companies operate. That makes management more important, not less. The organizations that succeed won’t be the ones that deploy the most agents. They’ll be the ones that know which agents perform, which ones fail, why they fail, and when human judgment still produces the better result.




