As AI scribe adoption grows, researchers at Suki challenge the industry's quality playbook

As adoption of AI ambient scribe technology rapidly grows, the healthcare industry's standard method for evaluating clinical AI notes may be fundamentally flawed, failing to detect significant errors or omissions in patient notes.

That's the position among researchers at Suki based on a study of evaluation rubrics. Suki offers an ambient clinical artificial intelligence solution for providers, and it works with more than 400 health systems.

Suki researchers argue that the healthcare industry's standard tool for evaluating AI-generated clinical notes, the Physician Documentation Quality Instrument (PDQI-9), which was developed in 2012, is poorly suited to assessing modern ambient AI scribes. That tool, the PDQI-9, was originally validated on a very small sample of inpatient notes and focuses on overall note quality—such as organization and conciseness—rather than identifying LLM-specific errors, the researchers wrote in a white paper.

The core PDQI-9 evaluation domains are up-to-date, accurate, thorough, useful, organized, concise, synthesized, internally consistent and cohesive.

But base quality rubrics "need a refresh in the post-LLM, ambient world," said Kevin Wang, M.D., Suki's chief medical officer.

Suki's white paper analyzes the limitations of holistic Likert-based scoring instruments, like PDQI-9, in reliably evaluating AI-generated clinical notes and detecting specific errors such as hallucinations. Suki provided Fierce Healthcare with an early look at the white paper.

In a study of 84 paired notes across four specialties, Suki researchers found significant inconsistencies among reviewers, including disagreement over what constitutes a hallucination, unreliable assessments of accuracy and varying scores depending on which AI model generated the note. They contend these issues reflect a fundamental mismatch between traditional note-quality rubrics and the types of errors LLMs make.

PDQI-9 scores holistically, looking at factors like organization, conciseness and cohesiveness, which is exactly the wrong lens for LLM errors, Suki researchers contend. Traditional evaluation frameworks may fail to reliably detect more consequential errors such as invented medication doses, missing diagnoses or omitted clinical details.

The white paper raises concerns about how health systems assess ambient AI scribes and suggests the industry needs more rigorous methods for measuring AI documentation quality and patient safety risks.

Suki's research was driven by a belief that healthcare needs a more nuanced way to evaluate the quality of ambient AI-generated clinical notes, Wang said. After reviewing prior studies and industry standards, the company concluded that existing measures may no longer be sufficient for assessing AI documentation. Suki researchers argue that older note-quality studies and evaluation frameworks, including PDQI-9, were developed before the era of ambient AI and were designed to assess EHR-based inpatient documentation rather than AI-generated clinical notes. The company contends those measures do not adequately capture the factors that matter most for ambient AI, such as factual accuracy, consistency, alignment with clinician intent and physician acceptance.

The white paper traces the evolution of note-quality and factual-accuracy assessments and lays the groundwork for a larger study Suki plans to release later this year, Wang noted. The company is exploring a new evaluation rubric aimed at addressing existing gaps and better measuring the quality of AI-generated documentation.

"It's going to show why there are limitations to existing frameworks and why something better is needed to prove real quality in ambient technology," he said.

Suki researchers also looked at the alternatives frameworks as well (PDSQI-9, SCRIBE, FActScore, VeriFact, CREOLA and others). Newer tools attempt to address limitations but still face low reliability, researchers concluded. None of these tools combine sentence-level error detection, inter-rater reliability as a first-class metric and a statistical procedure for model release decisions, according to the white paper.

The researchers argue that valid evaluation requires sentence-level, error-specific metrics with validated inter-rater reliability and statistical testing procedures.

Wang outlined some of the potential patient safety risks that can arise when LLM errors aren't detected in AI-generated patient visit notes. "There's a big difference in tuberculosis by saying latent tuberculosis, that means you have tuberculosis, it's just no longer active. That is very different than previously treated tuberculosis that's resolved. If I'm a doctor talking to a patient, I'm not going to give all these technical jargon terms. But imagine the note output one of those two things, and it's statistically due to variance. One says active TB, but not active right now. The other says you don't have it. Clinically, very different, medication-wise, very different," he explained.

"You, as the patient, want to be sure that your provider is outputting a correct note that matches your current clinical condition. And that's just one type of risk," he said, adding that there can be downstream impacts to reimbursement and insurance coverage due to upstream challenges with the quality and accuracy of the notes.

The difference in the quality of a clinical note provided to an insurer can mean the difference between a procedure like a colonoscopy being coded as a screening colonoscopy, which is often completely covered by an insurer, versus a diagnostic procedure, which requires patients to pay coinsurance or a deductible.

Wang contends that there needs to be more transparency among AI scribe vendors about how they evaluate the accuracy and quality of the technology's output.

"We'd love to see the statistical output. We've seen other scribe companies talk about hallucination. We'd love to see inter-rater reliability. We think those are things where it shouldn't be a black box," Wang said.

Showing the quality behind an AI scribe's technology will become "table stakes," he noted.

"More and more of our partners, as we sign them up, are doing their own quality evaluations. I think the industry is moving towards, hey, the health system has to convince themselves that what they're purchasing is something of higher quality than what they don't have, and I think that's going to be the new bar that is being set," he said.

There's an imperative to evaluate quality as AI adoption in healthcare surges ahead, Wang noted.

"In a year or two's time, one health system could have dozens, if not hundreds, of AI vendors. There are electronic health record AI vendors. There could be specialty pharma. There could be hardware. There could be televirtual. In all of these, imagine the number of times that quality isn't evaluated enough. We think you have to have a high-quality bar. At the end of the day, it's doctors and providers and clinicians treating real patients, and so that's where I hope to have a lot more of these conversations in the future, and a lot more clinically academic ones," he said.

He added, "I think we're going to see this movement where AI is pushing innovation and good user experience, and it's also going to push a new quality and evaluation experience. It's not the most glamorous, sexy headline, but there needs to be a refresh because so much of healthcare is based on something old or outdated. We have a chance to change that going forward."