top of page

Measuring AI ROI in an SME: why 95% of pilots return nothing, and how to be in the 5%

Nov 12, 2025
3 min read

Updated: 5 days ago

The figure did the rounds of boardrooms this summer: according to the report "The GenAI Divide: State of AI in Business 2025", published in August 2025 by MIT's NANDA project, 95% of organisations that launched generative AI pilots get no measurable effect on their P&L. Only 5% obtain a tangible return. We do not read this figure as a verdict on AI. We read it as a finding about how projects are set up, and above all about how they are measured.


What the MIT study actually says


Three lessons from the report seem directly useful for an SME.

The first: the gap does not come from the quality of the models but from their integration into real work. Generic tools, used individually, please employees but do not change processes; the projects that succeed are those anchored in a precise workflow and that learn from the company's context.

The second: organisations that buy a specialised solution from a vendor and adapt it succeed roughly twice as often as those that build in-house. For an SME without a technical team, this is good news: the shortest path is also the most reliable.

The third: the most solid gains are not in visible uses (marketing, sales) but in the back office, where hours of outsourced or repetitive processing get replaced. It is less spectacular, and it is where the money is.


The three measurement mistakes we see most often


When a business owner tells us "we tested AI, it did nothing", we ask one question: what did you measure, and against what? The answers almost always fall into one of these three categories.

  • Satisfaction is measured, not results. "The teams love it" is not a return on investment. The right indicator is an hour saved on an identified task, a shorter lead time, a falling error rate, an external invoice that disappears.

  • There is no baseline. If nobody timed the processing of a file before the tool, nobody can prove it went down afterwards. Measurement starts before the pilot, not after.

  • Hours saved are counted without saying what they become. Three hours per week per employee are worth something only if they are reallocated: to additional clients, to a hire avoided, to a response time that becomes a sales argument. Otherwise they dissolve.


A simple grid to cost a use case before launching it


Here is the grid we use with our clients. It fits on one page and takes an hour to fill in with the owner of the process concerned.

  1. The process, in one sentence. "Prepare renewal offers for maintenance contracts." Not "AI for sales".

  2. Current volume and time. How many occurrences per month, how many minutes each, measured over two weeks. This is the baseline.

  3. The full cost of the solution over 12 months. Licences, integration, internal configuration time, training. SMEs almost always underestimate internal time: count it at its real hourly cost.

  4. The target gain and its destination. A quantified objective (for example, from 45 to 15 minutes per offer) and the planned use of the freed hours.

  5. The decision point. A date, within three months, when measured is compared with targeted and a decision is made: scale, adjust or stop. Without that date, the pilot becomes a subscription nobody questions.


A worked example: 45 minutes that become 12


A technical services SME in French-speaking Switzerland, 35 employees, prepared around 90 maintenance contract renewal offers each month. Each offer took an account manager 45 minutes: re-reading the intervention history, checking prices, drafting the letter. Measured baseline: 67 hours per month.

The chosen tool, an assistant connected to the field service software and the price list, now prepares a draft offer that the account manager checks and signs. Time measured after three months: 12 minutes per offer, or 18 hours per month. The 49 hours freed were reallocated to client visits, and the renewal rate rose by four points over the half-year, which management attributes partly to offers being sent earlier. Full cost over 12 months, internal time included: the equivalent of about 5 months of the hourly gain. The return is there, and above all it can be demonstrated.


What we take away


The 95% in the MIT study are not short of tools; they are short of a precise process, a measured baseline and a decision date. The 5% have those three things before choosing a technology. This is what we call starting from the outcome, not from the tool.

To identify the processes in your SME where this reasoning applies first, our AI Barometer gives a first reading in a few minutes. And if you would rather discuss it directly, get in touch.

 
 
 

Comments


bottom of page