Skip to main content
Research

Comparison of Quality, Accuracy, and Empathy of Physician and AI Responses (PAIR) to Real Patient Questions

Haley Jeffers, Emma Sacks, Arjun Williams, Steven Williams MD, Justin Sacks MD MBA, Lynn Jeffers MD MBA

Oral abstract presentation, Research and Technology session — Plastic Surgery The Meeting 2024 (American Society of Plastic Surgeons), San Diego Convention Center, Saturday 28 September 2024. Abstract no. 42376. Presented by Haley Jeffers on behalf of the research team.

Official record

ASPS abstract no. 42376 — Plastic Surgery The Meeting 2024

About this page. PAIR was presented as an oral scientific abstract at a medical conference. Conference presentations are not indexed in PubMed or Crossref, and this abstract does not appear in the meeting's published supplement — so the ASPS record above is the only external listing, and it carries the submitted abstract rather than the full results.

This page exists so the study can be checked rather than taken on trust. Full methods, sample sizes, statistical tests and results are below. This work has not been peer reviewed or published in a journal, and should be cited as a conference presentation.

Why the study was done

In April 2023, a study in JAMA Internal Medicine compared physician answers with ChatGPT answers to patient questions posted on Reddit's r/AskDocs, and found the chatbot responses were rated higher for quality and empathy. That result was widely reported.

The finding rested on three methodological gaps:

PAIR was designed to close all three.

Methods

Data source

Patient questions and physician answers were drawn from RealSelf.com, where physician accounts are verified and answers are publicly attributed — a closer analogue to a real patient–physician interaction than an anonymous forum.

Design

Sample

240 ratings per group (physician, ChatGPT, Bard) — 720 ratings in total, across 180 matchup comparisons (60 per pairing).

Results

Mean ratings (1–5 scale)

MetricPhysicianChatGPTBard
Quality3.913.733.47
Accuracy3.873.603.27
Empathy3.373.143.02

95% confidence intervals

MetricGroupSD95% CI
QualityPhysician0.8033.806 – 4.010
QualityChatGPT0.7693.631 – 3.827
QualityBard0.8383.364 – 3.577
AccuracyPhysician0.8533.758 – 3.975
AccuracyChatGPT0.9003.481 – 3.710
AccuracyBard0.9263.149 – 3.384
EmpathyPhysician0.8123.263 – 3.470
EmpathyChatGPT0.6043.065 – 3.218
EmpathyBard0.6692.936 – 3.106

One-way ANOVA

Metricp-valueResult
Quality2.42 × 10-8Significant
Accuracy4.35 × 10-12Significant
Empathy3.89 × 10-7Significant

Tukey's HSD — pairwise comparisons

ComparisonQualityAccuracyEmpathy
Physician vs ChatGPT0.0390.0030.001
Physician vs Bard1.17 × 10-83.95 × 10-122.60 × 10-7
ChatGPT vs Bard0.0010.0001780.142 — n.s.

Physicians rated significantly higher than both AI models on all three measures. ChatGPT rated significantly higher than Bard on quality and accuracy, but not on empathy.

Head-to-head matchups — which answer was chosen as better

MatchupWinsp-valueResult
Physician vs ChatGPT34 – 260.197Not significant
Physician vs Bard44 – 160.001Significant
ChatGPT vs Bard36 – 240.155 (exact binomial, two-sided)Not significant

Overall win rates: physician 65%, ChatGPT 52%, Bard 33%. The ChatGPT vs Bard p-value is a two-sided exact binomial test on the 36–24 win count, recomputed to correct a transcription error — see the disclosure table below. The other two rows are as originally reported.

The central finding

Physician answers scored significantly higher than both AI models on quality, accuracy and empathy. But in direct head-to-head matchups, expert raters chose the physician answer over ChatGPT only 57% of the time — a difference that was not statistically significant (p = 0.197).

ChatGPT was measurably less accurate, yet close to indistinguishable in preference. That gap — between how good an answer is and how good it looks — is the reason this project exists. A patient without training has no way to see it.

The submitted abstract and the final analysis differ — here is why

The abstract listed in the ASPS record was submitted months before the meeting, and it says so in its own words: it reports "initial data" and states that "MedLM is currently being tested." The results on this page are the completed analysis presented at the meeting. Anyone comparing the two will find three differences, and all three are accounted for:

PointSubmitted abstractFinal analysis (this page)
Models tested ChatGPT and Bard, with MedLM named as still under test ChatGPT 3.5 and Bard. MedLM was not included in the completed analysis
Physician vs ChatGPT, mean ratings "trended higher… but did not show significant difference" Significantly higher on all three measures (Tukey HSD: quality 0.039, accuracy 0.003, empathy 0.001)
Physician vs ChatGPT, 1:1 matchups "physician answers were preferred over chatbot responses" Preferred 57% of the time — not statistically significant (p = 0.197)
ChatGPT vs Bard, 1:1 matchups Earlier versions of this page reported p = 5.89 × 10-8, "significant" That value was a transcription error (it resembles the ANOVA-scale values elsewhere on this page). A two-sided exact binomial test on the 36–24 split gives p = 0.155 — no significant ChatGPT-vs-Bard preference

Preliminary results changing once the full dataset is analysed is ordinary in conference research. It is documented here rather than left for a reader to discover, because a curriculum about verifying sources should hold its own sources to the same standard.

What this study does not show

How to cite this study

Citation
Jeffers H, Sacks E, Williams A, Williams S, Sacks J, Jeffers L. Comparison of Quality, Accuracy, and Empathy of Physician and AI Responses (PAIR) to Real Patient Questions. Oral abstract no. 42376 presented at: Plastic Surgery The Meeting 2024, American Society of Plastic Surgeons; September 28, 2024; San Diego, CA. Abstract: https://ww6.aievolution.com/asps/Events/viewEv?ev=10170 — Full methods and results: https://verifyproject.org/pair

Questions about methods or requests for the underlying data can be directed through The Verify Project.