Last Tuesday, a patient sat across from me and slid a stack of portal messages to the edge of my desk. She had printed them out, all twelve pages, because she wanted me to see what the last six months felt like from her side. “I know you’re busy,” she said, “but I need a person to answer me, not a wall of text.”
I remember the pause that followed. I had opened the visit thinking about medication changes and imaging. She was talking about attention, about being seen, about whether a system could still feel like care when so much of it had become interface.
Three weeks earlier, in a different room, a vendor had demoed an AI inbox assistant that could draft replies to patient messages. The dashboard looked reassuring. The promise sounded clean. Yet what I saw in the pilot mattered less than the pitch. The tool could save time on wording. It could not decide which message deserved a same-day call, which complaint was safe to answer in writing, or which sentence might calm a worried patient and which one would accidentally sound like dismissal.
The clinic teaches a different lesson
I used to think the central problem in medicine’s AI era was information overload. Then I spent more time watching what clinicians actually do with these tools. Now I think the deeper problem is judgment under compression, the moment when a machine speeds up the obvious work and exposes the human work underneath it.
I call this the compression gap. It is the distance between what software can complete quickly and what a clinician still has to hold in mind, namely context, risk, tone, prior history, and the social meaning of the encounter. The gap shows up in tiny ways. A draft note can be beautiful and still miss the one detail that changes the assessment. A drafted reply can sound polished and still fail the patient who is asking, in plain language, whether the doctor is paying attention.
This is where the current conversation about AI in medicine often gets lazy. People count minutes saved and declare victory. I care about minutes too. I live inside them. But the clinic does not reward speed alone. It rewards the kind of speed that preserves judgment, and that is a much narrower claim.
The best-known evidence points in that direction, but not as far as the headlines suggest. In Melnick et al., published in Mayo Clinic Proceedings in 2020, 870 physicians rated EHR usability at 45.9 on the System Usability Scale, a score the authors placed in the “not acceptable” range. That is a hard number, and it matters because it reminds me that clinicians did not suddenly become exhausted when AI arrived. The friction was already there. AI is entering a system that was already asking doctors to do too much on screens and too much in their heads.
In a 2024 study in JAMA Network Open, 162 clinicians used AI-generated draft replies for patient inbox messages. The mean utilization rate was 20%, and the study found lower burden scores without a meaningful change in time spent reading, writing, or replying. That is a useful result, but it is also a warning. Clinicians liked the relief more than the clock savings. The value lived in the feeling of friction easing, not in a dramatic reorganization of labor.
What I would not do
I would not hand an AI system the authority to answer messages, triage symptoms, or interpret ambiguity without a clinician in the loop. I would not let a hospital treat draft quality as clinical quality. I would not use a tool that makes the inbox look efficient while quietly pushing uncertainty downstream to nurses, residents, or patients themselves.
That line matters to me because I have seen what happens when a message sounds efficient but lands wrong. A patient interprets brevity as dismissal. A colleague interprets certainty as overreach. The result is not merely inconvenience. It is a leak in trust, and trust is the currency of the profession.
The strongest counterargument is obvious: if a tool reduces documentation time, lessens message burden, and gives clinicians back even a fraction of an hour, why resist it? I do not resist that. In a 2024 ambient scribe study discussed in JAMA Network Open, documentation time fell by roughly 13 minutes per clinic session, and after-hours EHR time did not change much. That tells me the benefit is real but partial. It helps at the point of writing. It does not repair the whole day.
That limitation is the story. AI reduces certain forms of clerical drag while leaving the moral and relational work intact. The profession still has to decide, in real time, what deserves an interruption, what deserves reassurance, and what deserves humility.
Physician identity in the AI age
The obvious reading is that AI threatens physician identity because it can imitate parts of our output. The clinic teaches otherwise. Output was never the whole identity. My identity as a doctor lives in the choices that happen before the note is signed, in the risk I am willing to tolerate, in the questions I ask when a symptom does not fit, and in the way I speak when I have to admit I do not yet know.
I remember a case that embarrassed me a little. I was sure a patient’s fatigue was mostly medication related. The data looked tidy. The story did not. The patient came back later and said, quietly, “I told you I was getting worse, not just tired.” She was right. The work I had to do was not algorithmic correction. It was professional correction. I had to absorb that I had overfit the tidy explanation.
That is the part AI cannot borrow from me, and I do not say that defensively. I say it because medicine trains a kind of fallibility management. We learn how to be wrong without becoming reckless, and how to be confident without becoming careless. A model can surface patterns. A clinician has to live with the consequences.
For readers who want the broader context of how I think about the profession and the work behind the work, I keep a running set of essays at my editorial site on medicine, technology, and practice, and my background is summarized on Dr. Sina Bari’s clinician profile and Stanford training page.
The human part remains the bottleneck
I have come to think that the scarce resource in medicine is not information, and not even time. It is interpretive attention. AI can draft, sort, summarize, and pattern-match. It cannot take responsibility for the meaning of a painful sentence or the social cost of a delayed reply. Those still belong to clinicians.
That is why the most useful deployments are the ones that stay modest. A draft response. A note skeleton. A routing suggestion. A signal, not a verdict. When the software stays in that lane, the clinician stays awake inside the work. When the software claims more than it can hold, the physician becomes a supervisor of confident mistakes.
I do not think that is a small distinction. I think it is the profession.
Back at the desk
At the end of that clinic day, I came back to the patient who had brought in the printed portal messages. We read through the pile together. I answered some immediately. I flagged others for follow-up. I ignored one draft reply I might once have trusted because it sounded fluent but felt too thin.
She looked relieved, but not because the system had become smart. It had not. She was relieved because someone had made room for the part of the visit that software does not own. The messages were still there. So was the work. What changed was the judgment applied to both.
That is my strongest view now. AI will keep taking pieces of medicine’s repetitive labor. Fine. Useful, even necessary. The profession will remain, though, where it has always been at its best, inside the compression gap, translating speed into care without letting one cancel the other.
FAQ
Can AI draft patient portal replies without changing the quality of care?
Yes, but only if a clinician reviews the content before it is sent. The best results so far show lower burden for clinicians, not a replacement for judgment, and the risk is that a polished reply can still miss urgency or tone. The draft should be a starting point, not a clinical decision.
What did the 2024 JAMA Network Open study actually find about AI inbox tools?
In that study, 162 clinicians used AI-generated draft replies and the mean utilization rate was 20%. The tool lowered burden scores, but time spent reading, writing, and replying did not meaningfully change. That suggests the main benefit was cognitive relief, not a dramatic time savings.
How does Dr. Sina Bari think about AI and physician identity?
Dr. Sina Bari’s view here is that physician identity lives in judgment, accountability, and the ability to admit uncertainty. AI can reproduce parts of output, such as summaries or drafts, but it cannot own the consequences of a clinical decision. That distinction matters most when the case is ambiguous.
Why do AI scribes help documentation time but not always after-hours work?
Because the burden is spread across the day. A scribe can shorten the time spent typing during a visit, but inbox messages, chart review, orders, and follow-up still consume attention later. The 2024 evidence shows real gains, but only at selected points in the workflow.
What is the compression gap in clinical AI?
The compression gap is the distance between what AI can finish quickly and what a clinician still has to judge carefully. It shows up when a draft, summary, or triage suggestion saves effort but leaves the human to manage risk, nuance, and trust. That gap is where the profession still lives.