Everyone Says to Pilot Your AI Project First. Here’s the Case for Simulating It Instead.

The Playbook Is More Sophisticated Than "Try It and See"

Search "how to measure AI ROI" and what comes back is genuinely thorough. I discovered the following very important pieces of the approach: define clear objectives and KPIs, establish a pre-AI baseline, quantify the fully loaded cost of the initiative, track both hard ROI like cost savings and time saved and soft ROI like decision quality and customer experience, and use controlled experiments or A/B tests to isolate what the AI actually caused versus what would have happened anyway regardless. Good frameworks even warn against specific pitfalls, like measuring at a single point in time instead of continuously, or evaluating an initiative in isolation instead of accounting for what it does to the rest of the portfolio.

That's a real methodology, built by people who have clearly done this before, not a strawman I'm setting up to knock down. But read through every step of it and there's a thread running underneath all of it, which is that every single piece of that framework describes how to measure an AI initiative once it already exists, once it's live in production or running inside a controlled test with real data moving through it. None of it tells you how to know whether the initiative is worth building in the first place.

What Real-World Testing Actually Costs You

Controlled experiments and pilots are the right tool for confirming a result you already suspect. They're an expensive and slow tool for finding out whether there's a result worth confirming at all, and that distinction matters more than it gets credit for. Standing one up means picking a team or a segment, building or configuring the tool, training the people who have to use it, and then waiting, often for weeks or months, until enough real volume has passed through it for the comparison to actually mean something.

And even then, a single test window only tells you about the conditions it happened to run under during the course of the pilot. If your test period landed on a quiet stretch, or a busy one, the numbers you walk away with don't necessarily generalize to the rest of the year. A well-designed A/B test is still measuring one slice out of a much wider range of possible outcomes and treating that slice as the answer, unless it runs long enough and across enough real variation to actually be representative, which is exactly what makes it expensive in the first place.

The Missing Option: Simulate the After, Before You Build It

Here's what the standard methodology skips over....you can model the "after" state computationally, before you've committed to a pilot, an A/B test, or a production rollout of any kind.

Take the process you're thinking about automating or handing to an AI tool. Build a discrete event simulation of it, with realistic arrival rates, task durations, and resource constraints, the same inputs you'd already be gathering for any process improvement engagement. Add the proposed change, whatever step gets automated, whatever task the AI tool now handles, however much time or capacity it's supposed to free up. Then run the simulation thousands of times instead of once, so it accounts for the variability a single real-world test window simply cannot show you, the slow weeks, the volume spikes, the edge cases nobody thought to plan for.

What comes out is a distribution of likely outcomes, quantified in the same terms the standard playbook already uses, cost per outcome, cycle time, throughput, resource utilization, but across a range of conditions instead of whatever conditions happened to occur during one test window. I want to be clear that this doesn't replace ever running a controlled experiment. What it does is let you find out, before you've built or bought anything, whether the idea even clears the bar of being worth testing, and it lets you walk into that test with a specific, quantified prediction instead of a hopeful one.

The Bottleneck That Moves, Not Disappears

The most rigorous AI-ROI frameworks already sense a version of this problem, even if their answer to it stops short. They warn against evaluating an initiative in isolation instead of accounting for its effect on the rest of the portfolio, and I think that's the same insight I keep coming back to, just stated at the portfolio level instead of the process level. Most guidance answers it by saying to consider the portfolio effect after the fact, once you have a few initiatives running and can compare notes. Simulation lets you see the same effect before the fact, inside a single process, before you've committed to anything.

Here's what I mean in practice. Say you automate document intake, and what used to take days now takes minutes. Measured on its own, that's an easy win to report, faster intake, lower cost per document, a clean number for the board. But intake was never the whole process, it was one station in a longer one. If the review step downstream was only keeping pace because intake had been feeding it work slowly and steadily, that same review team is now buried, and the bottleneck that used to sit at intake has simply relocated to review instead. Total time in system may barely move, or it may get worse, even though the piece you automated is genuinely, measurably faster than it used to be.

This is exactly the kind of effect a test scoped narrowly to the automated step tends to miss, because that test's own metrics look great in isolation. The dashboard says success but nobody's watching what happens three steps downstream, because that was never what the test was set up to measure. It takes looking at the process as a full system, not just the one point you changed, to see that the win didn't move the needle on the outcome that actually mattered, or worse, that it quietly created a new problem somewhere less visible. That's precisely what a process-wide simulation is built to catch, not just what happens at the point of automation, but what that change does to every resource and every queue connected to it, including the ones nobody thought to watch in the first place.

Where Controlled Experiments Still Earn Their Place

I want to be fair to the standard methodology here, because nothing replaces the reality check of an actual controlled experiment, especially for the parts of a system a simulation genuinely can't capture, like how people adapt to a new tool once it's real, or the soft ROI that only shows up once actual employees or customers are interacting with something. The strongest approach isn't simulation instead of a pilot or an A/B test. It's simulation first, to decide whether a real-world test is even worth running and to set a specific expectation for what it should show, and then a controlled experiment, scoped and informed by what the simulation already told you to expect going in.

The Bottom Line

The standard AI-ROI playbook is sound but it just starts one step too late, after you've already committed to building or testing something real. Quodsi fills the step before that, turning the process diagram you've already built into a simulation that tests an AI or automation change, and its effect on the rest of the process, before implementation begins.  Quodsi is built for business professionals who aren't simulation experts and simply want to take their static process diagrams and turn it into a dynamic, decision making tool. 
Next
Next

Systems Thinking: Why Process Consultants Are More Valuable Than Ever in an AI-Driven World