← Writing

We Measure AI Adoption and Productivity. Are We Measuring AI Dependency?

One half of the ledger has frameworks, vendors and a live argument. The other half has neither a definition nor a number.

TL;DR: The industry is having a serious, well-funded, unresolved argument about how to measure AI's impact on productivity. Frameworks, vendors, research programs, and people behind DORA and SPACE are actively revisiting how developer productivity should be measured in the AI era. That argument is messy, but it exists. There is not yet an equally mature measurement discipline on the other side. Most organizations cannot say which workflows stop if AI becomes unavailable, how much productivity is lost, how long recovery takes, or whether their fallbacks work. One half of the ledger is contested and improving. The other half has barely been opened.

There is a thread I came across with the title: "Claude went down and I went down with it."

It is funny. It is also a fairly precise description of an operational dependency that most organizations have never measured.

The half we are at least arguing about

AI productivity measurement is not a settled discipline. The published estimates do not merely differ in degree, they differ in sign. A 2023 controlled experiment by Peng et al. found developers completed a targeted task 55.8% faster with Copilot. METR's 2025 randomized trial, using real repositories the developers already maintained, found them 19% slower. Vendor case studies claim 2x and higher. DX's own measurements across tens of thousands of developers land at 5 to 15%.

Two randomized controlled trials, opposite conclusions, using radically different settings: a tightly controlled coding task in one case and real work in mature repositories in the other.

That is what an unresolved question looks like. But a serious attempt is under way, and the scale of it is worth appreciating.

Stack Overflow's 2025 Developer Survey found 84% of respondents are using or planning to use AI tools, and 51% of professional developers use AI daily. It found 52% agree AI tools or agents have had a positive effect on their productivity, and among AI-agent users, 69% agree agents have increased productivity.

DORA runs an annual research program on AI-assisted development. DX has published an AI Measurement Framework built with GitHub, Dropbox, Atlassian and Booking.com; DX reports average time savings of 3 hours 45 minutes per week and estimates a 5-15% productivity boost across its research, rather than the 50 to 100% figures sometimes seen in headlines. And there is a commercial market: Jellyfish, LinearB, Faros AI, Swarmia and Waydev all sell AI impact measurement.

Now the honest part. Even people who helped define modern developer productivity measurement argue that many existing metrics can be misleading in the AI era. Nicole Forsgren, who created both DORA and SPACE, discussed exactly this in a 2025 interview, including why coding gains do not automatically translate into equivalent developer or organizational productivity gains. DX's own guidance opens by noting most engineering leaders cannot answer basic questions about their AI investments, warns that suggestion acceptance rate is flawed because accepted code is often heavily modified or deleted, and describes the industry-standard impact metric as self-reported time savings.

So: contested, unfinished, actively being rebuilt by credible people with real money behind it.

That is exactly my point. Even unfinished, it has frameworks, vendors, benchmarks, longitudinal studies and a live argument. And it all points one way: what do we gain?

The other half nobody is arguing about

Now try answering these about your own organization.

  • Which engineering workflows would stop, or materially slow, if your primary AI provider were unavailable for four hours?
  • How many production code paths call a model? What is your measured productivity during an AI incident, rather than your assumed productivity?
  • When your primary model degrades, what percentage of requests successfully fall back, and to what? How long until normal throughput returns?
  • What latency threshold turns your AI-assisted workflow from faster into slower?

I cannot answer all of them for every environment I have worked in, and I have not yet met the engineering leader who can. That is not a failing of the people asking. These questions are not yet part of a broadly adopted measurement framework, so most organizations have not built the instrumentation to answer them.

We are not even measuring the half we think we are

In a randomized controlled trial using early-2025 data, METR asked 16 experienced open-source developers to complete 246 real tasks in their own repositories, with AI allowed on some and disallowed on others. They took 19% longer with AI. They estimated afterwards that AI had made them 20% faster.

Then METR ran it again with newer tools and a larger cohort. In February 2026 they reported that they could no longer measure it: "we have observed a significant increase in developers choosing not to participate in the study because they do not wish to work without AI." Among those who did participate, "30% to 50% of developers told us that they were choosing not to submit some tasks because they did not want to do them without AI."

One participant:

my head's going to explode if I try to do too much the old fashioned way because it's like trying to get across the city walking when all of a sudden I was more used to taking an Uber.

METR now believes developers are probably being sped up in early 2026, and is redesigning the experiment. A separate METR survey of 349 technical workers found a self-reported median change in work value of 1.4 to 2 times. So this is not an argument that AI does not help. It is an argument about instrumentation. A rigorous attempt to compare work with and without AI became harder precisely because some developers increasingly did not want to perform selected tasks without AI.

That is not a footnote about study design. That is the dependency, showing up as a measurement obstacle before most organizations have even named it.

This is not a hypothetical exposure

I tried to compare provider incident histories for the first 18 days of August 2026. The exercise exposed a measurement problem before it produced a reliability comparison: providers expose different scopes, different definitions, different levels of detail and sometimes different status surfaces altogether.

Anthropic, OpenAI, GitHub and Google expose incident information differently. GitHub publishes detailed postmortems with impact windows and root causes. Anthropic publishes impact windows on many incidents but rarely root cause. OpenAI's entries frequently carry generic recovery language with no window at all. Google exposes Gemini status across multiple surfaces, with the public Cloud feed focused on major incidents.

So the counts are not comparable, and publishing them as a league table would be dishonest. A provider that discloses more looks worse. A provider that discloses less, or splits disclosure across three dashboards, looks better. Reliability and transparency get scored as the same number.

That is the finding. You cannot compute the cost of downtime you cannot measure. You cannot set an SLO against a dependency whose impact windows are unpublished. You cannot compare providers on availability, which means availability cannot be a procurement criterion, which means there is little market pressure to improve it.

Cloud availability is not perfectly standardized either. Providers define "unavailable" differently, measure over different windows, and carve out different exclusions, and anyone who has argued a service credit knows the headline percentage is the least interesting part of the contract. But there is a shared vocabulary and a contractual structure to argue inside. Uptime percentages, service credits, exclusion clauses, defined measurement periods.

There is no broadly standardized, cross-provider way to express AI capability availability and its business impact.

I have spent a large part of my career on this problem in its older forms. Keeping infrastructure and databases available, and designing and running the disaster recovery plans behind them, for database estates and for cloud providers. That work has a settled vocabulary. RTO and RPO. Replication lag. Failover testing on a schedule rather than on hope. Uptime percentages that mean the same thing to a vendor, a buyer and an auditor.

None of it arrived by accident. It was built over years, mostly by people who had already been through an outage they could not explain to their board.

What strikes me about AI dependency is how little of that has carried over. We are running a critical dependency in production, and in our engineering workflow, with none of the vocabulary we would consider mandatory for a database.

One incident illustrates the deeper problem. In a Copilot incident on August 3, users hit errors in chat and agent features because a rate limit on an internal authorization lookup surfaced as failures. The postmortem notes: "The underlying AI models themselves remained healthy throughout."

The models were fine. The capability was unusable. Any metric asking "was the model up?" scores that as a success. I have cited GitHub most here for one reason, and it is worth stating plainly: they publish the most. That is a disclosure ranking, not a reliability ranking, and the two should not be confused.

Naming it: AI Dependency Debt

AI Dependency Debt is the gap between how effectively a person, team or system operates with AI and how effectively it operates without it. Not an established metric. A name for a risk we are carrying unmeasured.

A developer who needed eight hours now needs five. Call baseline 100 and AI-assisted 160. The intuitive assumption is that an outage returns you to 100. I am not sure that is always true. The developer has not forgotten how to code, but the workflow now assumes AI for exploration, boilerplate, testing, debugging, documentation, SQL and code review. Those things still happen without it. They may take longer, and re-learning a discarded workflow is not free. So the number might be 90.

A crude starting metric:

AI Productivity Resilience = Productivity During AI Unavailability / Baseline Productivity

That team scores 0.9. It will not be precise. It needs to exist, be tracked, and be uncomfortable enough to start a conversation.

Concentration multiplies all of this. Standardizing on one provider, one gateway, one assistant and one toolchain is often reasonable, but it means the obvious mitigation may not be one. An Anthropic incident on 18 August named five different models in a single status update. Anyone whose fallback plan was "switch to another model from the same provider" had no plan.

What to actually do about it

Here is the part I think gets missed: the application side of this has well-established engineering patterns, even though those patterns are certainly not implemented correctly everywhere. The developer side is much less mature.

If AI sits in your product, the tooling exists today.

Failover is a well-established engineering pattern, and AI gateways now provide practical ways to implement it. LiteLLM lets you define a priority fallback list, OpenRouter fails over automatically when a model is rate-limited or unavailable, Portkey adds enterprise governance on top, and Cloudflare AI Gateway supports fallback routing and retries to backup providers. If your application calls a model directly with no gateway in front of it, that is a decision you should reconsider.

Observability is also available. Langfuse is open-source and self-hostable, Helicone sits in the request path capturing latency and cost, and Datadog's LLM observability puts token and latency metrics alongside your infrastructure metrics. These give you the degradation signal that provider status pages will not.

Provider status can be aggregated. StatusGator and IsDown track AI provider status pages and will alert you before your users do.

Fault injection for AI is starting to appear, though mostly as research rather than shipping tools. ChaosLLM is a proof-of-concept from a 2025 ISSRE workshop paper that sits between an agent and its tools and injects four failure modes: unreachable, slow, hanging, and plausible-but-wrong. Its finding is worth more than its maturity. Across 225 runs, slow responses caused almost no failures, obviously garbled output was caught and discarded, but subtle corruption, a right-shaped answer with a wrong value, produced 42% of all silent wrong answers. Degradation was more dangerous than outage, measured rather than asserted.

On the practitioner side, agent-chaos injects faults into agent workflows and Steadybit shipped a chaos engineering MCP server in 2025. All of it tests whether your agent survives a failure.

Which is where the actual gap starts. These tools solve pieces of the problem. None of them measures the organization-level human dependency I am describing.

I have not found an established tool that measures the organization's exposure through the people using it: which engineering workflows are AI-dependent, what the team's throughput actually was during a provider incident, and whether the fallback has been tested. If it exists, I would like to be corrected.

So my suggestions, in order of how cheap they are:

  1. Capture a baseline and capture it before you need it. Git, Jira and similar systems can reconstruct throughput, cycle time and revert rate. What you cannot reconstruct is how the team perceived its productivity at the time. More importantly, a baseline chosen after an outage invites questions about bias. Capture it beforehand, and report revert rate alongside throughput; more output with more rework is not necessarily more productivity.
  2. Inventory AI-dependent workflows. Identify where developers now rely on AI for coding, testing, debugging, documentation, code review or incident response. You cannot manage a dependency you have not identified.
  3. Measure the next outage. Treat the next provider incident as a natural experiment. Record the outage window, compare throughput and cycle time with the baseline, and ask what developers did instead. One real data point is more useful than a framework you never implement.
  4. Test the fallback. Having a second provider is not the same as having a working fallback. Route real traffic to it periodically and measure what breaks. Chaos engineering has taught us the same lesson for infrastructure. Netflix open-sourced Chaos Monkey in 2012 to terminate instances at random on a schedule, precisely because untested resilience is not resilience. I haven't found a mature equivalent that measures both the application and developer-workflow impact of removing the primary AI dependency.
  5. Put AI availability into procurement. Ask providers for impact windows, incident history and postmortem commitments, not just uptime percentages. If enough buyers ask for comparable information, disclosure standards can emerge.

The ask

DORA's 2025 research describes AI as an amplifier: it magnifies the strengths of organizations with strong foundations and the dysfunction of those without, and the largest returns come less from the tools than from the quality of the underlying system. If that holds for productivity, it holds for resilience. But you cannot amplify what you have not measured.

We built a measurement discipline for AI adoption. We are rebuilding one for AI productivity. We do not yet have an equally mature discipline for measuring AI availability, dependency, recovery time or fallback success. There is no widely adopted standard definition, benchmark or comparable cross-provider disclosure model to build on.

That is not an unsolvable problem. It is an under-invested one.

To be clear about my ask: the productivity work is valuable and should continue. What I am arguing is that the two sides deserve the same seriousness. The same budget, the same rigour, the same place on the dashboard, the same line in the board pack. Right now one side has frameworks, vendors and a live industry argument. The other has neither a definition nor a number.

We measure AI adoption. We measure AI productivity. AI dependency deserves equal weight.

If you have started measuring any of this inside your own organization, I would like to know more about it. That is the part of the conversation I think we need to develop next.

References

Research and frameworks

Provider status histories (compiled 18 Aug 2026, covering 1-18 Aug 2026; not comparable across providers, and not a reliability ranking)

Chaos engineering

Other

  • r/claude, "Claude went down and I went down with it" — reddit.com

Also published on LinkedIn.