The Pilot Chose Itself
A pilot program is not a random sample. Nobody runs a pilot by picking ten sites out of a hat. Someone picks the ten sites that already have a competent IT staff, a director who wants this on their resume, and enough slack in the budget to eat six months of friction without anyone getting fired. Then the pilot succeeds, and the writeup treats the success as a property of the software.
It isn't. It's a property of the site that got chosen, filtered through software that happened not to break it.
Think about what selection into a pilot actually requires. Someone has to apply, or get nominated, or say yes when a vendor calls. That someone is already paying attention, already has some baseline of working infrastructure, already has staff with the bandwidth to learn a new system on top of their existing job. None of that is evenly distributed. The places with broken infrastructure and burned-out staff don't apply to pilots. They're too busy running the fires they already have.
So the ten sites in the pilot are the ten sites that were closest to not needing the pilot in the first place. A hospital that already has a stable EHR backbone and a full-time integration engineer is the one that gets picked to test the new charting module, because the vendor needs a clean demo, not a case study in triage. The module performs well there. It would perform well almost regardless of what it did, because the site was built to absorb exactly this kind of change.
Then the module ships to every hospital in the network, including the one running on a server that hasn't been patched since a previous administration and a nursing staff already at capacity. The results are worse. Sometimes much worse. And the postmortem calls it an adoption problem, a training gap, a communication failure, anything except what it actually is: the second site was never in the sample the first result came from.
This is the part that gets skipped in most retrospectives on failed rollouts. The failure isn't a surprise if you look at who got tested and who didn't. It's the predictable consequence of measuring a joint quantity, technology times readiness, and reporting only the technology half. A pilot report that says "adoption reached 90% within eight weeks" is really saying "at a site already primed to hit 90%, this software didn't stop that from happening." That's a weaker claim than the headline number suggests, and it's a different claim entirely from "this software will get you to 90% wherever you deploy it."
The honest version of a pilot report would include a second column: what was already true about this site before the software arrived. Existing broadband. Staff turnover rate. Whether there was a champion internally pushing the project regardless of results. Most pilot reports don't carry that column, because the column would make the number look smaller, and smaller numbers don't get funded for phase two.
None of this is an argument against piloting things. You have to start somewhere, and starting with a site that can absorb the mistakes is reasonable. The argument is against reading the pilot's number as if it travels. It doesn't travel. It was measured at one point in a distribution of readiness, and the interesting, expensive, unglamorous work is finding out what the same software does at the other end of that distribution, at the site that didn't apply, didn't have a champion, and can't spare anyone to learn a new interface this quarter.
Nobody funds that study. It would just tell you what you could have guessed: that a tool built and tested where things already worked meets a place where things don't, and the place wins.