July 29, 2026

AI Governance Failures & Healthcare

WHAT WE CAN LEARN FROM FOUR RECENT STUMBLES

Late springtime of this year was a rough period for Team AI. The headlines arrived in a cluster, the way they always do when a hype cycle meets operational gravity. In the case the headlines revealed what can be fairly described, in the aggregate, as more in the realm of AI governance failures than technology breakdowns. 

Starbucks scrapped an AI-powered inventory management system after the technology failed basic product identification — routinely confusing visually similar items like different milk varieties and missing stocked products altogether. A company promotional video from the launch period captured the problem plainly: a peppermint syrup bottle sitting on the shelf went unregistered as the system scanned the bottles on either side of it.

 → Reuters — “Starbucks Scraps AI Inventory Tool” (via Yahoo Finance)

A major Pizza Hut franchisee filed suit alleging that its AI-driven operations platform handed DoorDash delivery drivers unusual visibility into kitchen timing — allowing them to delay pickups, batch deliveries, and cherry-pick higher-tip orders. The result, per the complaint: pizzas sitting out longer, slower deliveries, and frustrated customers.

 → Fortune — “Pizza Hut Franchisee Lawsuit: AI Adoption, DoorDash Delivery Drivers”

Uber burned through its entire 2026 AI budget in four months after rolling out Claude Code to its engineering organization. By March, 84% of engineers were classified as agentic coding users. Roughly 70% of committed code originated from AI tools. About 11% of live backend updates were written by agents with no human in the loop. Monthly costs ran $150–$250 per engineer on average, with power users reaching $500–$2,000. The tools worked. The budget did not.

 → Startup Fortune — “Uber Burned Its Entire 2026 AI Budget in Four Months”

Amazon quietly shut down KiroRank — its internal leaderboard tracking usage of the company’s Kiro agentic AI platform — after discovering employees had learned to game it. The practice acquired its own vocabulary: tokenmaxxing. Engineers spun up unnecessary agents and ran pointless tasks to inflate their scores, burning real compute dollars in the process. Senior VP Dave Treadwell was eventually moved to state, explicitly: “Please don’t use AI just for the sake of using AI.”

 → 404 Media — “Amazon Shuts Down Internal AI Leaderboard After Employees Cheated”

If you have been waiting for the AI hype cycle to develop some visible cracks, this was your week.

Read carefully, though. These are not the same story. They are four different failures — each with its own mechanism, its own accountability gap, and its own lesson.

———

STARBUCKS: THE VALIDATION PROBLEM

Let’s be precise about what happened at Starbucks. This was not an adoption failure. It was not a change management failure. It was not a trust problem.

The model did not work.

A system deployed to manage inventory at one of the world’s largest retail operations could not reliably identify the products on the shelves it was scanning. A peppermint syrup bottle — sitting in plain sight — went unregistered while the system catalogued the bottles on either side of it. On milk inventory, performance was horrible in differentiating the distinctions of Oat, Soy, Almond, Skim, and Whole. The detail matters because it wasn’t a corner case. Reuters found errors were common. Imagine that, a company whose brand is tied to six versions of creamer rolls inventory AI that can't differentiate.

The question that interests me is not why the model failed. Models fail. The question is how a model that failed this visibly cleared every checkpoint between vendor demo and production deployment at scale.

That is a validation and oversight failure. Somewhere in the procurement and piloting process, the gap between “performs well on benchmark” and “works reliably on a Tuesday morning in a real store” was either not tested or not weighted appropriately. The technology broke. But something in the governance process let a broken technology get that far.

In healthcare: We are observing similar gaps in prior authorization AI. Clinical decision support tools and AI-assisted prior authorization platforms which pass vendor demonstrations and procurement committees every day. The edge cases that break them — the ambiguous diagnosis codes, the dual-eligible patient, the non-standard formulary exception — rarely appear in the validation set. The peppermint syrup that went invisible on the Starbucks shelf becomes the authorization criterion the model confidently misclassifies. Providers and patients absorb the consequences. The procurement committee has moved on to the next initiative.

 ———

PIZZA HUT: THE SYSTEM BOUNDARY PROBLEM

The Pizza Hut story is more structurally interesting than it first appears — because the technology worked exactly as designed.

The platform gave delivery drivers visibility into kitchen operations and order timing. That visibility was presumably intentional — better information theoretically supports better coordination. What the designers appear not to have modeled is what a rational DoorDash driver would do with that information.

A rational DoorDash driver, compensated per delivery and competing for higher-tip orders, will delay a pickup if batching it with another order improves their economics. They will cherry-pick. They will optimize for their own outcome, not the franchisee’s. This is not bad behavior. It is entirely predictable behavior from an actor whose incentives were never aligned with the system’s intended outcome.

The failure was a system boundary failure. The designers mapped the organization. They did not map every actor the system touched. Gig economy workers sitting entirely outside the org chart were handed a tool that let them arbitrage kitchen operations in real time — and they used it.

The lawsuit is the franchisee’s problem now. The lesson belongs to anyone deploying AI in an environment with multiple stakeholders who do not share the same incentives.

In healthcare: Watch OpenAI’s stated ambition to own the patient advisor role. A platform that sits between patients and care navigation gains structural visibility into medication adherence signals, care-seeking behavior, and Rx decision-making before the provider or payer does. Prescription data is the most sensitive edge of that exposure. The regulatory frameworks for what a patient advisor platform can do with that visibility — and who is accountable when it is exploited — do not yet exist at the scale the technology is moving. The Pizza Hut franchisee did not anticipate what DoorDash drivers would do with kitchen timing data. The healthcare industry should be asking harder questions about what a patient-facing AI platform will do with longitudinal health intent data.

 ———

UBER: THE FINANCIAL GOVERNANCE PROBLEM

The Uber story is, in some ways, the most uncomfortable of the four — because it is a success story with an unaccounted bill.

The Claude Code rollout worked. Adoption climbed from 32% of engineers in February to 84% classified as agentic coding users by March. Roughly 70% of committed code originated from AI tools. About 11% of live backend updates were produced by agents operating with no human in the loop. By any engineering adoption metric, this was a successful deployment.

And yet Uber burned through its entire annual AI budget in four months.

The structural flaw was organizational, not technical. The teams driving adoption — engineering, product — were not the teams managing the spend. Monthly cost per engineer ranged from $150 to $250 on average. Power users ran $500 to $2,000. Naga, Uber’s engineering leader, reportedly spent $1,200 in a two-hour personal demo. That was not anomalous. That was the tool working as designed for the workloads it was built to handle: parallel agent execution, large-scale codebase refactoring, automated test generation.

Nobody had modeled what successful adoption at scale would actually cost. And the leaderboard culture Uber built around Claude Code usage accelerated token burn directly.

“The direct link to new consumer-facing features remains a mystery.”

 — Uber COO Andrew Macdonald

Output went up. Outcomes remain unclear. That distinction — between measuring what the AI produced and measuring what the AI produced that mattered — is the load-bearing question that most AI business cases leave unanswered.

In healthcare: Health system and RCM leaders are pushing AI adoption — ambient documentation, coding assistants, denial prediction tools — with urgency that is culturally driven as much as strategically driven. The engineering and operations teams measuring adoption velocity are rarely the same teams accountable for financial outcomes. Clinical ROI is at least as hard to attribute to AI output as Uber’s consumer-facing features were to token consumption. The CFO will eventually see the bill. The question is whether the governance structure to connect spend to outcome was built before or after that conversation.

 ———

AMAZON: THE INCENTIVE DESIGN PROBLEM

Amazon’s KiroRank story is the oldest lesson in the group, wearing the newest clothes.

KiroRank tracked how much employees used Kiro, Amazon’s agentic AI developer platform. The intent was presumably to encourage adoption, surface engaged users, and create visibility into organizational AI maturity. What it created instead was tokenmaxxing — a practice so widespread it required a named Senior VP to intervene publicly.

Employees created unnecessary agents. They ran pointless tasks. They burned real compute dollars doing work that had no business value, optimizing for a metric that was supposed to measure business value. Dave Treadwell’s instruction — “Please don’t use AI just for the sake of using AI” — is the kind of sentence that only needs to be said when the measurement system has thoroughly decoupled behavior from intent.

This is not a technology story. It is not even an AI story. It is a Goodhart’s Law story, playing out in a new medium: when a measure becomes a target, it ceases to be a good measure.

The only thing surprising about it is that anyone was surprised.

In healthcare: RCM organizations run on KPIs, and KPIs have always had this property. Optimize for DSO and watch teams prioritize easy receivables over complex recovery. Optimize for denial rate without tracking appeal recovery and watch the wrong denials get written off quietly. Optimize for AI tool utilization — logins, queries, recommendations accepted — and watch staff find the path of least resistance through the workflow without actually changing how they work. The downstream cost in healthcare is not compute spend. It is revenue leakage, compliance exposure, and the quiet erosion of outcomes that nobody’s dashboard is measuring.

 ———

FOUR FAILURES: FOUR DIFFERENT LESSONS

It is tempting to collapse these into a single narrative: AI is overhyped, organizations are underprepared, management is failing. That narrative is not wrong. It is just not precise enough to be useful.

The taxonomy matters:

Starbucks was a validation failure. A model that could not perform its core function cleared procurement and reached production. The question to ask before the next deployment: what would it take to break this in production, and did we test for that?

Pizza Hut was a system boundary failure. The designers modeled the organization. They did not model every actor the system touched. The question to ask: whose incentives does this system change, and are any of those people outside our org chart?

Uber was a financial governance failure. A successful rollout produced an unaccounted bill because adoption accountability and cost accountability lived in different parts of the organization. The question to ask: who is responsible for connecting AI spend to AI outcomes, and do they have the authority to act on that?

Amazon was an incentive design failure. The measurement system created the behavior it was intended to track — and that behavior was the wrong one. The question to ask: if someone wanted to game this metric, how would they do it, and is that what we are currently rewarding?

The organizations that navigate this decade well will not necessarily be those with the most sophisticated models. They will be the ones that asked these questions before the go-live date rather than after the lawsuit, the budget overrun, or the SVP memo.

 ———

WHAT THIS MEANS IF YOUR SETTING IS HEALTHCARE

Healthcare did not invent any of these failure modes. But it has a well-documented talent for making them more consequential.

The validation gap that let a broken Starbucks model reach production is the same gap that allows AI-assisted prior authorization tools to confidently misclassify edge cases that experienced reviewers would catch. The procurement committee approved the accuracy rate. Nobody modeled the failure mode.

The system boundary failure at Pizza Hut has a direct analog in any patient-facing AI platform with ambitions at the scale OpenAI , for example, is targeting. When a platform gains longitudinal visibility into patient health intent — including Rx behavior, care navigation, and adherence signals — before the provider or payer does, the question of who is accountable for how that visibility is used is not a future regulatory problem. It is a present governance problem.

The Uber financial governance failure maps precisely onto AI rollouts where clinical and operational leaders are measured on adoption velocity and CFOs discover the cost structure in a quarterly review. Output is measurable. Outcomes, in healthcare as at Uber, are considerably harder to attribute.

And the Amazon incentive design failure is native to RCM. DSO. Denial rates. Authorization turnaround. Clean claim rates. These metrics exist because they are measurable. What gets written off, delayed, or missed because the incentive pointed in a slightly wrong direction is considerably harder to see — and AI tools do not fix that problem. They accelerate it.

The harder, slower, less glamorous work is the same across all four failure modes:

1. Validate before you deploy. Not on benchmarks. On the edge cases that will actually break the system in production.

2. Map every actor the system touches. Not just the ones on your org chart.

3. Connect spend to outcomes before adoption runs ahead of accountability.

4. Design incentives for the behavior you want, then ask how someone would game them.

None of that is an AI capability. All of it is a leadership responsibility.

===================================

Larrian Martin is the Chief Information Officer at GoSB, a specialty healthcare RCM company, He was formerly Senior Vice President at Envision Healthcare and holds an Engineering Science degree from Penn State University, where he studied in the Artificial Hearts Lab.