Comparison of Quality, Accuracy, and Empathy of Physician and AI Responses (PAIR) to Real Patient Questions
Oral abstract presentation, Research and Technology session — Plastic Surgery The Meeting 2024 (American Society of Plastic Surgeons), San Diego Convention Center, Saturday 28 September 2024. Abstract no. 42376. Presented by Haley Jeffers on behalf of the research team.
About this page. PAIR was presented as an oral scientific abstract at a medical conference. Conference presentations are not indexed in PubMed or Crossref, and this abstract does not appear in the meeting's published supplement — so the ASPS record above is the only external listing, and it carries the submitted abstract rather than the full results.
This page exists so the study can be checked rather than taken on trust. Full methods, sample sizes, statistical tests and results are below. This work has not been peer reviewed or published in a journal, and should be cited as a conference presentation.
Why the study was done
In April 2023, a study in JAMA Internal Medicine compared physician answers with ChatGPT answers to patient questions posted on Reddit's r/AskDocs, and found the chatbot responses were rated higher for quality and empathy. That result was widely reported.
The finding rested on three methodological gaps:
- Length was not controlled. The chatbot answers were substantially longer than the physician answers, and longer answers tend to score better.
- Accuracy was never measured. Raters assessed quality and empathy, not whether the information was correct.
- Reddit is not a patient–doctor relationship. Answers came from an anonymous public forum.
PAIR was designed to close all three.
Methods
Data source
Patient questions and physician answers were drawn from RealSelf.com, where physician accounts are verified and answers are publicly attributed — a closer analogue to a real patient–physician interaction than an anonymous forum.
Design
- Each patient question was submitted to ChatGPT 3.5 and Google Bard.
- Each model was instructed to answer within 10 words of the physician's answer length, controlling for verbosity.
- Responses were assembled into blinded head-to-head matchups: physician vs ChatGPT, physician vs Bard, and ChatGPT vs Bard.
- Every response was rated by three board-certified plastic surgeons on quality, accuracy and empathy (1–5 scale). Raters also selected the better answer in each matchup.
Sample
240 ratings per group (physician, ChatGPT, Bard) — 720 ratings in total, across 180 matchup comparisons (60 per pairing).
Results
Mean ratings (1–5 scale)
| Metric | Physician | ChatGPT | Bard |
|---|---|---|---|
| Quality | 3.91 | 3.73 | 3.47 |
| Accuracy | 3.87 | 3.60 | 3.27 |
| Empathy | 3.37 | 3.14 | 3.02 |
95% confidence intervals
| Metric | Group | SD | 95% CI |
|---|---|---|---|
| Quality | Physician | 0.803 | 3.806 – 4.010 |
| Quality | ChatGPT | 0.769 | 3.631 – 3.827 |
| Quality | Bard | 0.838 | 3.364 – 3.577 |
| Accuracy | Physician | 0.853 | 3.758 – 3.975 |
| Accuracy | ChatGPT | 0.900 | 3.481 – 3.710 |
| Accuracy | Bard | 0.926 | 3.149 – 3.384 |
| Empathy | Physician | 0.812 | 3.263 – 3.470 |
| Empathy | ChatGPT | 0.604 | 3.065 – 3.218 |
| Empathy | Bard | 0.669 | 2.936 – 3.106 |
One-way ANOVA
| Metric | p-value | Result |
|---|---|---|
| Quality | 2.42 × 10-8 | Significant |
| Accuracy | 4.35 × 10-12 | Significant |
| Empathy | 3.89 × 10-7 | Significant |
Tukey's HSD — pairwise comparisons
| Comparison | Quality | Accuracy | Empathy |
|---|---|---|---|
| Physician vs ChatGPT | 0.039 | 0.003 | 0.001 |
| Physician vs Bard | 1.17 × 10-8 | 3.95 × 10-12 | 2.60 × 10-7 |
| ChatGPT vs Bard | 0.001 | 0.000178 | 0.142 — n.s. |
Physicians rated significantly higher than both AI models on all three measures. ChatGPT rated significantly higher than Bard on quality and accuracy, but not on empathy.
Head-to-head matchups — which answer was chosen as better
| Matchup | Wins | p-value | Result |
|---|---|---|---|
| Physician vs ChatGPT | 34 – 26 | 0.197 | Not significant |
| Physician vs Bard | 44 – 16 | 0.001 | Significant |
| ChatGPT vs Bard | 36 – 24 | 0.155 (exact binomial, two-sided) | Not significant |
Overall win rates: physician 65%, ChatGPT 52%, Bard 33%. The ChatGPT vs Bard p-value is a two-sided exact binomial test on the 36–24 win count, recomputed to correct a transcription error — see the disclosure table below. The other two rows are as originally reported.
The central finding
Physician answers scored significantly higher than both AI models on quality, accuracy and empathy. But in direct head-to-head matchups, expert raters chose the physician answer over ChatGPT only 57% of the time — a difference that was not statistically significant (p = 0.197).
ChatGPT was measurably less accurate, yet close to indistinguishable in preference. That gap — between how good an answer is and how good it looks — is the reason this project exists. A patient without training has no way to see it.
The submitted abstract and the final analysis differ — here is why
The abstract listed in the ASPS record was submitted months before the meeting, and it says so in its own words: it reports "initial data" and states that "MedLM is currently being tested." The results on this page are the completed analysis presented at the meeting. Anyone comparing the two will find three differences, and all three are accounted for:
| Point | Submitted abstract | Final analysis (this page) |
|---|---|---|
| Models tested | ChatGPT and Bard, with MedLM named as still under test | ChatGPT 3.5 and Bard. MedLM was not included in the completed analysis |
| Physician vs ChatGPT, mean ratings | "trended higher… but did not show significant difference" | Significantly higher on all three measures (Tukey HSD: quality 0.039, accuracy 0.003, empathy 0.001) |
| Physician vs ChatGPT, 1:1 matchups | "physician answers were preferred over chatbot responses" | Preferred 57% of the time — not statistically significant (p = 0.197) |
| ChatGPT vs Bard, 1:1 matchups | Earlier versions of this page reported p = 5.89 × 10-8, "significant" | That value was a transcription error (it resembles the ANOVA-scale values elsewhere on this page). A two-sided exact binomial test on the 36–24 split gives p = 0.155 — no significant ChatGPT-vs-Bard preference |
Preliminary results changing once the full dataset is analysed is ordinary in conference research. It is documented here rather than left for a reader to discover, because a curriculum about verifying sources should hold its own sources to the same standard.
What this study does not show
- It did not measure AI confidence, hedging or calibration. Claims about AI sounding equally confident whether right or wrong are not PAIR findings and should be cited to the calibration literature instead.
- It is not a clinical outcomes study. Ratings of written answers are not measurements of patient harm.
- It is limited to aesthetic and plastic surgery questions on one platform, rated by plastic surgeons. It does not generalise automatically to other specialties.
- It tested ChatGPT 3.5 and Google Bard as they existed in 2024. Newer models may perform differently; the team has noted work extending the study to current systems.
- It has not been peer reviewed.
How to cite this study
Jeffers H, Sacks E, Williams A, Williams S, Sacks J, Jeffers L. Comparison of Quality, Accuracy, and Empathy of Physician and AI Responses (PAIR) to Real Patient Questions. Oral abstract no. 42376 presented at: Plastic Surgery The Meeting 2024, American Society of Plastic Surgeons; September 28, 2024; San Diego, CA. Abstract: https://ww6.aievolution.com/asps/Events/viewEv?ev=10170 — Full methods and results: https://verifyproject.org/pair
Questions about methods or requests for the underlying data can be directed through The Verify Project.