Health systems have spent years running AI pilots. The question now is what comes after scaling them.
Four leaders working at organizations with populations ranging from tens of thousands to 5 million patients say the answer requires something most organizations haven’t yet built.
The challenge is no longer proving that AI can work in a controlled environment. It’s making AI work reliably at the scale of a real health system across dozens of facilities, thousands of clinicians and patient populations that don’t behave the way pilot cohorts do. At Pittsburgh-based UPMC, Somerville, Mass.-based Mass General Brigham, Oakland, Calif.-based Permanente Federation and California-based Stanford Health Care, leaders are confronting that problem differently but they’re arriving at a shared standard.
At the 11th Annual IT + Revenue Cycle Conference: The Future of AI & Digital Health, four leaders spoke on a panel about what AI at scale actually looks like.
Mass General Brigham has been measuring the impact of ambient AI on clinician well-being for nearly four years, and Rebecca Mishuris, MD, vice president and chief health information officer, said the results changed how the organization thinks about proof.
“When we started this journey almost four years ago, there was no data,” Dr. Mishuris said.
The system started with 10 physicians on her team testing the technology to confirm it was safe and producing accurate documentation before moving to a broader pilot. The goal of that pilot wasn’t operational efficiency, but it was clinician well-being.
“I have a very enlightened CFO who said, ‘I don’t care about the finances. There are real outcomes here,’ and so that was our goal for this: can we improve the well-being of clinicians with this technology?” Dr. Mishuris said. “We measured that. We measured on a regular basis our clinician well-being and professional fulfillment, as you know, the national standard to do that. And we did that on these patterns. We measured everybody at the start, and then at six weeks, and then at 12 weeks, and within a year.”
The results were significant. Mass General Brigham documented nearly a 20% reduction in clinician burnout, a figure Dr. Mishuris said has no parallel in the broader burnout literature. What surprised her team was that the technology didn’t need to be sold.
“This was not a technology I had to sell,” Dr. Mishuris said. “This was a technology that I had to slow down, safety quite yet, right? Because it sold itself, and the reason it sold itself was because it truly transformed how we were working. It took the technology out of the room. It let me just talk to my patient and not have to worry about the administrative burden of writing my note.”
With burnout reduction now proven and sustained, the organization has shifted its measurement focus. The next question isn’t whether clinicians feel better, but whether patients do.
“What we’re interested in now is what’s the patient experience, and are we actually changing the quality of care that we’re delivering through this technology?” Dr. Mishuris said.
At UPMC, the prerequisite for scaling AI is different: a unified data foundation. CIO Chris Carmody said the system just completed a milestone that makes the next phase of AI work possible.
“We spent the entire weekend converting about 1.4 million patient records,” Mr. Carmody said, describing UPMC’s recent Epic go-live. “Our goal is really for the technology to fade into the background, so our doctors and nurses can actually turn their attention to our patients and provide better care.”
With approximately 120,000 technology users across the system, UPMC is now building on that unified platform. Mr. Carmody said the system is moving toward AI tools that handle the administrative layer, freeing clinical staff to focus on care, and has set a concrete internal target for every IT professional at UPMC will use AI agents to do their job more efficiently by the end of 2027.
One early example is prescription refill processing. UPMC uses natural language processing to handle several million refill requests per year, a task that previously required staff time spent on phones and manual verification. Mr. Carmody said the shift has moved people away from administrative work and toward higher-value functions.
He cautioned, though, that the pace of AI adoption creates real cost risk. Token pricing is uncertain, and the temptation to deploy broadly without governance structures in place can lead to runaway spending and inconsistent outcomes.
“It can get out of hand very quickly if you open up these tools without structure,” Mr. Carmody said. “We have to be measured, we have to be collaborative, we have to manage that. Having solid processes and structure in place — that’s what will help us, no matter what model is leading in 2027 versus 2025 or 2026.”
For Permanente Federation, scale is where the team starts. The federation serves approximately 5 million members, and Khang Nguyen, MD, chief medical officer of care navigation, said the work of scaling AI is fundamentally about patient care navigation.
Permanente Federation’s care navigation engine integrates clinical understanding, operational capacity and patient factors to match patients to appropriate care settings. A study examining the tool’s performance showed approximately 90% accuracy in routing patients to the appropriate appointment type. The system processes free-text inputs and draws on multiple signals simultaneously to make those determinations.
“The notion that it is scaled is important,” Dr. Nguyen said. “We are involved in our governance structure around what to do.”
Michelle Mello, JD, professor of law at Stanford Law School and professor of health policy at Stanford University School of Medicine, said the question health systems most often skip is whether a tool that works in testing will actually work in the workflow it’s being deployed into and whether anyone is measuring that after deployment.
“We can’t just have a room full of experts sitting around and imagine what it would be like to be at this tool,” Dr. Mello said. “We go out and do work with front-end users. We have a learning community, patients that we engage in governance, a patient-centered panel. We talk to developers and ask hard questions. It’s only through this process of stakeholder consultation that we really surface some of these issues.”
Stanford has made its governance framework publicly available, including templates, toolkits and case studies, for health systems building their own processes.
Across all four organizations, the measurement question sits at the center of what it means to scale responsibly. Dr. Mishuris framed it as a shift in what accountability actually requires.
“AI adoption should proceed not at the pace of availability, but at the pace of accountability,” Dr. Mishuris said. “There is so much that we turn on; the approaches we take after we turn on to make sure it’s delivering value for patients matter.”
No comments:
Post a Comment
Note: Only a member of this blog may post a comment.