
Clinical AI on Trial: Shadow Tools and Governance Gaps | Newsday with Yaw Fellin and Matt Troup
Questions Answered in This Episode
- Should health systems trust benchmark studies claiming general AI models beat specialized clinical tools?
- Are clinicians unknowingly deploying shadow AI tools that bypass governance entirely?
- Who bears liability when an AI tool makes a clinical error—the provider or the vendor?
- How do health systems establish governance frameworks faster than AI tools evolve?
- What does clinical-grade AI actually mean when validating tools for patient care?
About This Episode
July 20, 2026: A new Nature of Medicine study claims general-purpose LLMs outperform specialized clinical AI, but is that the right benchmark to begin with? Bill and Drex are joined by Yaw Fellin, General Manager of UpToDate at Wolters Kluwer, and Matt Troup, Clinical Strategy Principal at Abridge, to break down what those benchmarks actually measure, why shadow AI is already spreading through health systems, and what "clinical grade" really means. Plus: the first-ever look at UpToDate's new drug dosing feature in Expert AI, and why automaticity bias may be the next liability crisis hiding in plain sight.
Key Points:
03:18 Nature Study Breakdown
10:25 Shadow AI Reality Check
18:37 Transparency and Provenance
23:40 Human in the Loop Wrap
Keep up to date on the latest in health IT:
https://thisweekhealth.com/news/
Donate: Alex’s Lemonade Stand: Foundation for Childhood Cancer
Contributors
People featured in this episode — open a profile for more.
Transcript
This transcription is provided by artificial intelligence. We believe in technology but understand that even the smartest robots can sometimes get speech recognition wrong. I'm Bill Russell, creator of this week Health, where our mission is to transform healthcare one connection at a time. Welcome to Newsday, breaking Down the Health it headlines that matter most. Let's jump into the news. All right, it is news day, and, uh, exciting conversation today joined by Drex, the incomparable Drex DeFord, um, with the 229 project. Yaw Fellin, who is, general manager of, uh, UpToDate, Wolters Kluwer. And then, Matt Troup with, uh, clinical strategy principal with, uh, with Abridge. So look forward to this conversation. How are you guys doing? Doing great. Great to see you, Bill Yeah. Doing fantastic. It's a pleasure to, join you again, Bill Well, I, I'm looking forward to this. Drex and I have been together for the last couple of days. We've been up here in, the wine country. Up here in Napa, we had a CIO summit, we had an analytics, a chief analytics officer summit, and we had a chief nursing, informati- or informatics, uh, CNIO, chief nursing information officer thing. Or is it informatics? No, it's information, It could be either way, I think, probably. If you polled the people who were in the room, I bet you'd get about a 50/50 going on there We're gonna touch on some of the stuff that's, we were talking about. I, I know you guys are gonna find this hard to believe, but AI came up quite a bit in, uh, our conversations. both of them, all again, huh? the thing I wanna talk about with you guys because you're, you're both positioned and, and, um, uh, Wolters, you guys just released a white paper and obviously Abridge, you guys had a bunch of really interesting announcements. So a, a new debate is forming the role AI in clinical settings. I mean, we're, we're having these discussions and, uh, one of the, one of the articles that came up, I'm sure you guys are very familiar with the recent study in, uh, Nature of Medicine found that general-purpose frontier, uh, large language models outperformed specialized clinical AI tools, including open evidence and up-to-date, expert AI across medical benchmarks. Um, open evidence has strongly disputed this. And I, I wouldn't be-- And actually a bunch of the clinicians who were in my room this past weekend strongly disputed this. Uh, but the, the study does raise the questions you know, what, what does this look like? Uh, the larger issues isn't really simply which AI model won the benchmark. The real question for, for these health systems is how clinical AI should be e-evaluated, governed, uh, deployed um, you know, at, at, at the same time we're, seeing AI adoption is, is already moving faster than these institutions are ready to accept it. So, uh, we have unauthorized AI tools that's, that's in the, the white paper from Wolters Kluwer as well. and then, you know, I, I love the fact that we have the two of you on here. Uh, Abridge just last week announced, "Hey, you're gonna build your own foundation model." So there's, there's a whole bunch of things coming up. Uh, let me, let me throw out a first question here. Um, the, the Nature of Medicine study suggests that general-purpose large language models may outperform specialized clinical AI tools on certain medical benchmarks. I guess the question is: What should we make of this? Is this a meaningful signal for health systems or are benchmarks still, removed from clinical reality? I think this conversation is, uh, as you noted, Bill, it's absolutely, you know, kind of top of mind for, for everybody. Um, we've been looking at it deeply with our clinical and research teams and, you know, I think there's a handful of things that are important to draw out, right? Uh, nu- number one is just, you know, to, to what extent, um, you know, are some of these tools actually trained on the tasks that they're taking? Um, and I think that's one of the... You know, when people talk about, you know, was the methodology or the study design, you know, appropriate or strong enough, I think that's some of the consideration that's coming into the picture. Mm-hmm. Um, and, you know, th- there's a kind of a wealth of data out there that, you know, basically shows that in some way, shape, or form, and this is not my area of expertise, um, that, you know, these foundational models have been trained on the tasks. Uh, so I've seen, you know, some statistics that say, you know, 96, 97% of the time they're delivering truly like an exact verbatim answer. Um, and I think that's a, that's a challenge, right? Uh, you know, if you're training specifically on the test, um, you know, what does that demonstrate as you think about applying this to real world clinical applications? And I think that gets to the deeper question, which is when you get into, you know, these more sophisticated, uh, point of care clinical scenarios, um, it, you know, what is the best approach, uh, to support high quality care and, and patient outcomes? Um, and is that one that could be based, you know, solely off of a foundational model, or do you need actually a true, you know, kind of AI system that can be governed and, um, you know, has the right level of oversight, uh, for the sensitivity of, of clinical practice? So I'd, I'd probably start there. Matt, I don't know, uh, as a clinician if you would correct or edit anything I said. Well, I think it's-- I mean, this is the f- probably the, the first paper of many that are gonna continue to make this the, uh, uh, yeah, top- topic of discussion, um, across the industry. So in that way, it's good in some ways too. Like, there needs to be much more than, than this one paper and, and better understanding of what benchmarks we all should be holding ourselves to and for those companies that are developing in this space. I do think, though, it, it, it continues to, uh, shine a light on a core need, which is validation and just benchmarking in general. And I think where Abridge is, and I'm sure this is true of Wolters Kluwer and UpToDate as well. But where Abridge is sitting right now is that we are, we are building on top of enterprise partnerships, and by doing that, there are expectations are set that when we are, when we are partnering with a large, complex healthcare system, there are expectations and there are trust that are part of every part of that partnership that we build. And because of that, we are held to a pretty high standard. Um, versus if you were just a clinician who was picking up a cons- you know, a consumer app and have no idea whether this has been validated and there's nobody, know, that might know that you're using this and keep, you know, have the awareness of how it's being used in your clinical practice. But by partnering with systems and rolling this out across the enterprise, there is an expectation that's being set. And, uh, Bill, I know you- you've heard this from Shiv, but one of the things we talk about is moves at the speed of trust, and innovation is at an all-time fast pace. It's moving at light speed compared to anything in the, in the past couple decades. And I think for us to do this work well, we have to hold ourselves to that high standard, develop the whitepapers, which I know both of our organizations are, are doing, trying to be as transparent as possible in how we build. And then for the end user clinician who is told by their healthcare system, "Hey, you can use this product now," they have the confidence that somebody else has, has done the, the work, the governance committees have authenticated this, that there's a confidence level in using it. Because at the end of the day, that end user clinician is just trying to get through that clinic with the best insight possible and doesn't wanna worry about whether or not they can trust or not trust the technology that they're being given. Oh, yeah. know, uh, and, and Drex, I wanna throw you a question 'cause you're neither a clinician nor... But I, I, but I wanna throw you this question, which is- Uh ... I think every time I open up OpenAI or open up Claude, there's an update. There is Like, when we, when we, when you and I as CIOs used to bring out updates to software, we had a, a rigorous process to make sure was, it was validated. And these guys are treating this like consumer software and then saying, "Hey, it can make clinical decisions." I'm a little... I, I mean, I understand what they're saying, but I'm a little concerned Uh, trust was a big part of the conversation and the, uh, chief analytics officer, um, group too. And the, I think the challenge of, um, yeah, when you're w- as CIOs, when our core software systems update, we have a very careful change control testing, make sure everything works the way we expect it to work, and then we, um, we release it into the environment to be used in operations. And when stuff like this happens with Claude or ChatGPT, it just updates every two days or something. There's some update that is radically different from the update that was released two days ago. Like, it's gotten, uh, much better, or at least we hope that it's gotten much better. Um, the, the other interesting part of this is the, are we holding AI to a higher standard than we would hold, Mm-hmm. you know, Matt, you kind of referenced, like, maybe somebody's using something on their, on their iPhone, and they're using it in their clinical decision-making, but nobody knows. And so we don't hold them to that higher standard that we're holding AI to today. And that's another whole thing that I think we're struggling with in the governance process and what comes on board and how we use it. And are we, I mean, yeah, we probably should. There are patients and families involved, and so we need to, we need to hold it to a higher standard. But, um, you know, are we also being maybe a little unrealistic? And that rolls into the whole conversation about, and so who's holding the bag when something goes wrong? Is it the provider? Is it the company who produced the product? Is it, you know... So there's a bunch of questions that kind of spring from this issue. we still don't have a liability framework for this. Well, actually we do, it's the clinician. It Yeah but I, I wanna throw out some of these numbers from the, uh, from the Wolters Kluwer whitepaper. We have, uh, only 27% of clinicians know their org's AI policy. I mean, that doesn't surprise me at all. I, I-- my guess is the CIO, CIOs who knew their org's AI policy wouldn't, would, would, would still be below 50, yeah because they're, they're, they're still being written. Uh, nearly 20% u- uh, of clinicians used unauthorized AI, 10, 10% for direct patient care. So we have this shadow AI thing in the clinical setting and, uh, absorb- o- observed unauthorized AI use. Um, the paper says 40% of healthcare professionals, admin staff have seen unauthorized AI tools used in their And we're not even touching on the patient side yet. The patients are absolutely using AI. Uh, we had, uh, I about 15 to 20 people s- uh, s- standing around, uh, the other night and I, I just asked, "How many of you personally have used AI for your, for your own health, and questions you have around health?" it was unanimous. It was every single one of them had, had used it at some point. Um, so I, I wanna go back to what Drex is saying, which is, clearly these tools are being used for health and, and, for health questions and those kind of things on the, on the patient side. From the CIO side, we should be terrified by the fact that they're taking their phone out and taking a picture of the screen, of their Epic screen and sending it to OpenAI or, uh, OpenAI or whatever tool they're using. you know, how, how do we establish this, this... Uh, it's-- I guess as I'm listening to these CIOs talk, they're, they're, they're kind of like, man, this, this area is moving so quickly. We need to establish the, the framework and the foundation for where these tools are gonna be used and how they're going to be used. They recognize the probabilistic nature of these models and how scary that is, but they also, you know, they're, they're trying to remain relevant in a, in a world that's moving very fast. I mean, what, what does it look like to have these conversations and to establish, how a, how a health system would approach the use of AI in their, in their environment? I mean, I think it's in a, a pretty exciting time now to be, uh, uh, working on this side of enterprise AI software because so much of this is a collaboration effort. Um, two, three years ago, when I first started at Abridge, and we were talking about deploying our first note generation product just to going from a conversation to a note output, um, very few partners had AI governance in place. Um, it was a CMIO making a decision or a couple folks making a decision, and you would-- have some, some metrics at the time or some benchmarks we were, we were, uh, tracking. But there wasn't a AI governance committee or oversight. And now fast-forward two years, and so many of the decisions, um, so many of the, the features and products we're rolling out are going through these processes. And it's been, um, incredible for us to be part of giving feedback on maybe how to do this, how to do this well, and, um, getting feedback from them on what they need from us to feel as though some of these features can get through governance stronger. Matt, what, what does clinical grade mean? I mean, when we're, we're working to use these tools in a clinical setting, what, what should we, how should we be thinking about that? From my clinical point of view, it's, is, is this, is this, uh, AI application solving a specific problem or are we just throwing AI at things that don't really still need AI to solve them? I think that is, I think that is the overwhelming nature of the current state that we're in, is that there are a lot of AI technology being thrown at clinicians or even at health systems, and not all of them are solving specific pain points in the clinical workflow. And then all you're doing is just making that clinician more paralyzed by the decision about what should I use when. And so I think from our point of view, it's if we're rebuilding the clinical workflow as, you know, what is most substantial to think about for a clinician before a visit, during a visit, after the visit, where can we apply AI to best solve the problems in that workflow for the current state for a positive impact, whether it's I need a, I need to prep for the visit I'm about to have. the visit, I need easy access to, you know, uh, an active listener in the room who can help with clinical decision support questions in real time. Or after the visit, it's populating artifacts like notes and codes and orders, or it's, uh, additional clinical decision support. It's matching the clinical trials. Like these are real-world problems we can solve with AI rather than just saying, know, "Here's a tool that can-- you can use if you want to, and it may or may not help and, know, but it's available to you." And, and then all that does is it reinvents some of the problems of the past two decades where clinicians just feel like, "I'm overwhelmed by choice," and that becomes really confusing , I mean, so, I mean, no, no, no tool has been more used than ambient listening at this point. I mean, accuracy, um, uh, usability, uh, a-addressing real workflow or real problems, bringing the information into the workflow. Um, and in, in, in so doing has really relieved the clinician burden. I mean, that's what we've seen in the Yeah. Oh, listening space Yeah. And Bill, if I could, if I could build on that, you know, back to your question of, you know, what is or what should the definition of clinical grade be? I think a key part of that definition is that, you know, clinical grade applications should be governed, right? And, you know, you talked, uh, about some of the research that we've done at Wolters Kluwer on Shadow AI. actually, you know, uh, some of the work that Matt and the Abridge team did really did establish, you know, what are governance frameworks because Ambient was one of the first spaces where this was all adopted. And I think what we've seen in the clinical decision support space is it's, you know, extending those governance frameworks to higher stakes domains like clinical decision support is a really critical part of the process. I mean, you know, casually, if you think about it, um, you know, a physician practicing in a hospital needs to be credentialed. There's a way that you govern, you know, which physicians get to come into the four walls uh, and interact with patients. And, and I think if you, you know, extend that concept into these clinical, you know, AI, uh, solutions, it's-- you know, that, that's the first step. Now, I don't think that's sufficient, right? So governance is thing number one. I would actually give, you know, CIOs and CMIOs a tremendous amount of credit. I think the, the speed with which they have, you know, built and adapted, uh, and that they're still discovering clinical decision support frameworks is, you know, kind of unparalleled in my twenty-five plus years in healthcare. Um, so, you know, I, I see those teams working exceptionally hard. As, as we've rolled out our expert AI application, one of the things we are most proud of is we are governance first, uh, you know, governance friendly. Um, and with our enterprise hospitals and health systems, um, we are getting, you know, darn close to seventy percent adoption of our expert AI features. We'll, we'll probably easily cross that threshold a-across this summer. Um, and each one of those is governed, and we've put that in front of committees. Uh, we've, you know, followed through on, on all of the, the processes. You know, for some organizations that didn't have stronger processes, we've been able to share best practices. Um, but I, I-- you know, just back to the core, I think governance is, like, the first ingredient, and then there's probably a whole bunch of other, um, you know, nuance that we could draw out beyond that. Right? And within that governance is the, the two words I keep hearing from these, these healthcare leaders as they move forward, which is observability and transparency. Yes in healthcare, we have to know why the AI made the decision it did, why it made the recommendation it did So Bill, this is such a critical point. Um, and actually we have not done a, a, you know, a massive press release on this yet, although it, it's coming. So, uh, one of the pl- places until last week that we had not gone yet, we had not yet included drug dosing in Expert AI in our clinical decision support, uh, utilities. And, and why, you know, why were we more conservative there? Well, um, you know, when you get drug doses wrong, that can have, you know, very catastrophic, you know, human and patient impact, right? Um, and so our clinical teams, uh, spent a tremendous amount of time developing a true systematic approach to drug dosing. and, uh, you know, that's not just, you know, how do you do it transparently, but how do you report on potential harm? Uh, you know, how do you actually govern and be able to measure, uh, in a real-time way how the system is performing? Mm-hmm. last week was the first time that we got it to a stage where we felt like this could be introduced, you know, to our customer base, um, which is a massive, uh, you know, milestone from, from our perspective. And, you know, a big part of that is the system is performing tremendously well and, you know, we're able to manage out, uh, you know, the types of harms and errors that everybody fears. Uh, but the other part of that is we can now sit down with any, you know, CMIO, CMO, uh, you know, clinical leader, uh, and we can show how the system performs. We can tell you what percentages at a time do we answer with model knowledge. We can stratify the types of errors and risks that, you know, occur within the product, and we can actually give those, uh, you know, leaders agency, uh, in terms of is this something that is ready to be deployed at your organization. And, and I will also tell you there's a spectrum of leaders. Uh, you have some leaders, uh, you know, that are looking for, you know, however many nines, 99.99, you know, 9% accurate. Um, and that's gonna be the governance standard, and then you have other leaders that, um, you know, that really recognize that there is some, some error within the system and actually arming people with better tools and the right checks and balances, um, you know, could still reduce overall harm. Uh, but the core has been we've had to work that up in a, you know, a very rigorous, transparent, you know, systematized way. Uh, and then we've had to align that with, uh, you know, the governance approaches of the organizations that we're serving And I would add to that, Bill, that a lot more of the governance committees are requiring a lot, uh, requiring that to be a core part of the future, which is the ability just to have provenance for where this information came from. and I think you know well, Bill, from, from the start at Abridge, it's been about trying to track back to what we call linked evidence and how do you find the ground truth for how Abridge, um, or AI in general arrived at the answer. I was helping my, my son with homework over this past weekend, and I kept telling him, um, you know, "Show your work," right? It's like, how did you arrive? Did you just guess or did you actually do like the, you know, the equation to get there? And I think that we have to treat... I mean, it's a simple, a simple analogy, but we certainly have to treat and what we build at Abridge or across any tool the same way as like, where did this come from? How did you arrive here? Even if it's probabilistic, let's do the heavy lifting to help the clinician trace back to source of truth and in that way, you know, develop again that, that cycle of, of trust. Um, the more we just create more of a part of that as like our core ethos, then that continues to help governance councils, I believe, say, you know, when Abridge brings them the next product or next feature we want to run through that governance committee, oh great, they have this same sort of feature that allows the clinician to trust and verify the, the content Yeah. Trust, transparency. I mean, the transparency really leads to the trust. If I can see what's happening inside the black box, it's not a black box, which means I can see, I can give feedback, I can close the loop on the things that I think are good and bad about this, and we can... You know, we're operating in the clear. Everybody knows what's going on. There's a lot of apps that we use today, even in our personal lives, where at some point we click the license agreement, and somewhere in the license agreement it says something like, "And the stuff we provide to you might be wrong." But it, but there's no, uh, way to actually know that or feedback on it or anything else. So the more transparency, the more trust. I love that. I, I wanna exit on this 'cause we had a, a pretty robust conversation on human in the loop and, and liability. And, uh, the, the thing that they're noticing is that as they use these tools, especially ones they trust that have AI in them, they just get into this habit of saying yes, yes, yes, yes. Because it was right the first seven times, of course it's gonna be right the, the next seven times. Um, and that's true of notes, uh, especially, uh, they were s- they, they did talk about the, uh, the success rate of ambient listening and, and the creation of the notes, and they have trusted partners. A lot of your clients were in there, and they're saying like, it, it, it's right, uh, you know, most of the time. And so they just, they just start to get into this pattern. But they're accepting liability when they hit that yes. Um, and that's one of the things they're saying is, "This is the reason we have to drive them closer to, to five nines and not 95," 'cause 95 is unacceptable, um, because we're gonna take the liability and they're gonna click, We, click, same problem, though, back in the early EHR days when alerts, uh, drug-drug interaction alerts or other things would pop up and, and, you know, clinicians would just, uh, there were so many, they would just wind up bypassing them. Uh, so it's not the first time we've had this dance I, I mean, we think of this as the automaticity bias, right? It's, uh, you know, people, you know, become comfortable with the systems. They become used to them being correct, and then they, they follow the guidance. And, and, and Bill, I do think there's, you know, part of that challenge is just how accurate can you make, you know, the system or the answer. I think the other part that we've spent a lot of time on in the clinical decision support space with UpToDate is, you know, how do you clarify intent? Uh, how do you nudge, uh, and reinforce the clinical reasoning skills of the clinician, right? And, and one of the other big themes that we spend a lot of time talking with our, uh, you know, our customers and, and physicians and the academic community about is this concept of, you know, de-skilling or in some cases, uh, never skilling, um, I think is kind of the new term that's coming out from a, from a clinician perspective. So, you know, particularly in the clinical decision support domain, you know, we, we do feel like it's important to have some safeguards built into the system, right? Um, you know, if we've made an assumption about a patient, you know, let's state that assumption so that the end user can see it and decide if they need to change what the system has assumed. Um, if there's ways to, you know, present some of the clinical reasoning or clinical logic, you know, let's make those transparent so that, uh, you know, the end user kind of gains skills through the application. And so I do think it's gonna take a holistic, you know, set of approaches to overcome some of those biases or some of those risks that are, uh, you know, just Drex, as you said, kind of inherent in these systems, whether we're talking about AI or, uh, or even some of the, the earlier stage applications. Oh, yeah. I mean, I think of, uh, I think of a decade or two ago when you're thinking about copy and paste and in EHR and how often I, I showed up to see a patient that I was gonna start rounding on and, know, there's an error from four days ago that just gets carried forward every day because nobo-nobody has the time and bandwidth and cognitive load. We're all overwhelmed. And a single line of text that was copied from four days ago, it just has never gotten flagged or corrected, and it perpetuates. And of that is just ensuring that we give clinicians the, the time and space to do the work well, um, so we can... You know, the reduction in cognitive load that we've seen, I'm hoping, is giving, myself included, just more bandwidth to sit and think critically about what I'm writing, what I'm putting in the chart, how I'm processing information UpToDate in my workflow to get to the right answer at the right time. But I, I also agree that as we develop tools, we need to not, um, recreate the wrongs of the past and now fall into a, a place where AI is generating these notes or these outputs, and similarly, I'm just rushing through it, signing it, and moving on to the next one. And so there's a lot we can do here. I mean, certainly we're, we're, you know, by kind of tracking the workflows and where there is risk for some of this, but also just ensuring again that that evidence is there if there's ever a question, that we're always pushing the envelope on how accurate we can be. I think back, Bill and Drex, to your earlier questions of what clinical grade even means. It's understanding that when you use AI workflows in a clinical setting, that the bar, again, just has to be high for reducing the amount of the-- I would say the, the expectation for how few errors get, end up in, in the outputs and that, um, we also have the metrics and data to, to track along with what are the edits that clinicians are making. How, how often are they making edits? How often are they taking the CDS answers and actually incorporating them or making decisions based on those? And using that to maybe also help us understand some of the human behavior now with the, the age of AI and AI in our workflows Well, gentlemen, I want to thank you. This is a great conversation. Uh, you know, it still feels like we're very early on in this AI discussion, even though we're a couple of years in. I mean, we're, we're more than that because AI's been around for a while. But, but just this, this latest explosion and sprint, um, know, we're a couple of years in, but still feels like we're really on the, on the front end of this. I mean, there's awful long, a-a awful long way to go and an awful lot of, uh, questions still to be answered as we bring this in, especially into the clinical setting. Administrative setting, okay making mistakes. Hopefully, hopefully not too many. But, um, Yeah you know, that was another story I, I thought about bringing up, but we're, we're out of time. But just this, this whole idea of the, the, the difference between the two, because on the administrative side, we just saw that United Healthcare appropriated, uh, like, a billion dollars for AI development and automation and all that other stuff, which, Drex, we've talked about this. This is the-- my AI bot's bigger than your AI bot, and they're just gonna Yeah and forth. It's gonna be, it's gonna be really interesting to, to watch. Uh, gentlemen, hey, thanks for your time. Really appreciate it Thanks, Bill. Appreciate it you, Bill That's Newsday. Stay informed between episodes with our Daily Insights email. And remember, every healthcare leader needs a community they can lean on and learn from. Subscribe at this week, health.com/subscribe. Thanks for listening. That's all for now.





