
For the past couple of years, “AI Strategy” and “AI Implementation” have become phrases no investor or board meeting is a stranger to. CEOs launched pilot programs with the urgency of military mobilization. As a consequence of this rush to move on AI, many organizations have found their body has outrun their shadow and now no clear plan exists to answer a seemingly simple but important question: Is it actually working?
It’s worth pausing to consider just how difficult that question really is to answer. In many cases, the difficulty is increased by a lack of available data, but buried beneath is the deeper question of what “working” itself actually means.
Part of this confusion is categorical. When we reach for analogies to understand AI, we tend to grab ones that tend toward the grandiose. Is AI like the printing press, democratizing access to creative pursuits, allowing anyone to produce what once required specialized training, talent or skill? Is it more like electricity, transformative at a civilizational level, but only after an enormous infrastructure is built around it? Or Is it like the calculator, taking something humans already knew how to do but making it faster and way more scalable?
The answer matters because each category implies a different measurement framework. You don't evaluate an accelerant the same way you evaluate labor replacement, and you don't evaluate either the same way you evaluate a new infrastructure layer. Organizations that collapse these categories end up measuring the wrong things or measuring nothing at all.
Let’s get one thing clear first: At the end of the day, artificial intelligence is intended to solve a problem, whether it's solving an existing pain point or it's improving a certain workflow, or it's leveraging a capability that was previously inaccessible. Calculators didn't create mathematics: they unlocked a speed and scale of computing that was enabling for people who understood math. Large language models are doing something similar with unstructured data in the form of text, notes, documents, and recorded conversation that previously sat inaccessible to scale. That's a real unlock, and it's not necessarily the same thing as replacing headcount. In other words, companies that treat AI primarily as a cost-reduction lever may be missing a more interesting question.
Here’s a useful lens to consider: Do you consider AI your product, or is it an improvement tool for your operations? In some cases, AI really is the product. Ambient documentation tools for clinicians are a clear example. The AI is delivering clinical value directly. Whether or not it's working–and the data for confirming this–is relatively easy to determine: adoption rates, time saved, clinician satisfaction, documentation quality. The measurement question more or less answers itself.
The more common view of it that companies take, though, is operational: AI helps a marketing team produce content faster, or it acts as a coding copilot for an engineering team. Here, "working" is murky. In theory, operational efficiency gains should be straightforward to measure. In practice, most organizations didn't have the measurement infrastructure in place before they adopted AI in the first place, such as velocity metrics. It means they have no counterfactual to compare against, like they’ve been handed a supposedly faster engine but have never seen a speedometer.
This isn't a small problem. Implementing AI before implementing the means to measure AI is like an engineering team that adds headcount with no point assignments on their sprints. You can feel like things are moving faster without being able to demonstrate it, and you certainly can't tell what's driving the change.
There's a further complication that disrupts the standard return-on-investment logic most companies apply to their AI decisions: AI is not a fixed software cost. The tokenization model underlying most generative AI means that usage scales directly with consumption such that costs may scale unpredictably with use. Early signals from the organizations that have moved furthest on AI-enabled engineering suggest something surprising: in some cases, replacing an engineer with AI tools is not, in fact, cheaper than just keeping the engineer. The economics that made AI feel like obvious margin improvement may not hold at scale, or may depend on a footprint large enough to amortize what turns out to be a variable (and sometimes significant) per-use cost. This doesn't mean the economics are necessarily bad, but it does mean the financial modeling that underpinned many AI decisions was built on assumptions that deserve to be revisited now that there's more usage data.
The question of whether AI is working and where its true value lies may even be counterintuitive. There's an apocryphal story about a photography professor who divided his class into two groups at the start of the semester. One group would be graded purely on the volume of photographs they produced. The other group would be graded on a single photograph, i.e. their “best.” At the end of the term, the strongest images came overwhelmingly from the volume group. They shot more, they experimented more, they failed faster and learned from it more often. Quantity, counterintuitively, had produced quality.
There’s a possible parallel here for AI: the most measurable value AI generates may not be operational efficiency or cost reduction, but the ability to iterate faster. Faster iteration means more experiments per unit of time. More experiments means more learning, faster. The organizations pulling ahead may not be the ones that implemented AI most carefully but rather the ones that used it to compress their feedback loops the most. Critically, this is only achievable if an organization firmly ascribes to a learning culture of continuous improvement.
This reframe doesn't eliminate the measurement question; instead, it sharpens it. If speed of iteration is the core value, then what you should be measuring is how quickly your teams can test, learn, and change course - and, of course, whether AI is visibly accelerating that cycle. And, maybe one sign that the foundational learning culture described above already exists is that those metrics were already in place and monitored before AI even arrived.
What's clear is that "we deployed it" was not the finish line. The real work is building the measurement infrastructure, clarifying which category of AI value you're actually pursuing, and making sure human ingenuity remains coupled with technological enablement.