All posts

The Ten-Minute Conversation: What Clinical Supervision of AI Use Teaches Law Firms

A resident in clinic takes out a phone between patients, asks a chatbot for a differential diagnosis and a management plan, and pastes the answer into the patient’s note. The attending physician, watching from across the room, has to decide what to do about it. That scene opens a review article published in the New England Journal of Medicine in August 2025 by Raja-Elie Abdulnour of Brigham and Women’s Hospital, Brian Gin of the University of Illinois College of Medicine, and Christy Boscardin of UCSF, on how clinical educators should supervise trainees who use AI. The attending’s first thought in the vignette is “Now what?”, and the authors’ answer is a structured conversation with the resident about what the resident asked the model, what came back, and what the resident did with it, before any teaching about the case itself.

The article does not address the privacy problem in its own vignette. The resident in the illustration has used, in the words the authors give the resident, “the free version of ChatGPT on my phone,” which means a patient’s presentation has gone into a consumer tool on a personal device, and the authors say nothing about it. A law firm cannot leave that out, because an associate who pastes client facts into a consumer chatbot has a confidentiality problem under Rule 1.6 before the quality of the output is even in question, a problem I have taken up before and one that a firm platform used under commercial terms exists to avoid. This post assumes the associate used the firm’s platform and takes up the question the article does answer, which is what the supervisor should do about the reasoning.

The rest of the scene has an obvious law-firm counterpart. The partner’s version arrives by email as a research memo, a first draft of a motion, or a diligence summary, produced by an associate who may have used the firm’s AI platform for some or all of it, and the partner has to decide what to do about that too. The authors’ answer matches a habit I have been encouraging at my firm: before you review a junior lawyer’s work product, spend ten minutes talking with the junior about it. It belongs in the ordinary course of an assignment, because it teaches the associate, by making them reason aloud about work the tool may have reasoned for them, and because it makes the partner’s review faster and more focused. The NEJM article supplies the structure of that conversation and the evidence for why it works, from a profession that trains its juniors the way we do, by supervised practice on live matters.

Deskilling, never-skilling, mis-skilling

Off-loading rote work to a tool can help, the authors note, because it frees working memory for harder tasks. But a learner who off-loads the reasoning itself risks what they call deskilling, never-skilling, or mis-skilling. Deskilling is the loss of a skill the learner once had. Never-skilling is the failure to build a skill at all, because the learner got the tool before doing the practice that would have built it. Mis-skilling is the acquisition of a wrong skill, reinforced by a model’s errors or biases that the learner accepted as correct. The distinctions are useful because each names a different person in a law firm.

Deskilling is the mid-level associate, or the partner, who used to read every case and now reads the model’s summary of it. Never-skilling is the first-year who has never built a research path from a statute to the controlling case, because the platform returned an answer with citations attached from the first assignment onward. That associate can confirm a citation exists but cannot tell whether the model found the right line of authority, since they have never found one themselves; I described them in June as the associate who turns in clean work for years while never acquiring the ability to evaluate what the model gave them. Mis-skilling is the associate who absorbs the model’s characteristic errors as doctrine: who learns that a standard is the standard because the model stated it confidently in three separate memos, or who learns, from a tool that validates whatever framing it is handed, that their analysis was sound because the model agreed with it. The authors’ figure of skill over time places never-skilling below the floor of competence required for independent practice, and that is where the associate who never learned to research stays until someone notices.

The evidence the authors assemble comes from medicine and adjacent fields, and the findings most relevant to law need no methodology to follow. In a field experiment with several hundred management consultants, Dell’Acqua and colleagues found that on a task chosen to sit beyond what the model could do well, consultants using AI were less likely to reach the correct answer than consultants working without it; the drop, as the NEJM authors summarize it, came from users accepting the output in place of their own judgment. When final-year medical students at Imperial College London were shown chatbot answers to clinical scenarios, half of them wrong, more than a third of the students missed the errors. And a study of radiologists reading chest X-rays with AI assistance found that when the tool was wrong, the radiologists were more likely to be wrong too, and that the lower-performing readers were not reliably the ones the tool helped. The Bednar trial showed the same pattern in law students from a different angle: the tool’s benefit ran through the quality of an intermediate document, not through any improvement in the students’ reasoning, and where students deferred to the tool at the revision stage, strong work got worse. The tool’s errors fall hardest on the people least equipped to catch them, and those people are the supervisor’s responsibility.

The medical setting has another feature law firms share. The attending in the vignette may know less about the chatbot than the resident does, and the authors treat that inversion as normal, comparing it to earlier moments when faculty had to learn a new system of practice alongside their trainees. Many partners supervising AI-assisted work have used the tools less than the associates producing it, and some have not used them at all. The instinct in that position is to supervise the only thing one can see, the finished document, and to leave the tool to the associate. The authors’ answer is that the supervisor’s inexperience is a reason to ask about the interaction, since the conversation is also how the supervisor learns what the tools are doing in their own practice.

The moment the supervisor cannot see

The most useful idea in the article is a definition that never mentions technology. The authors define an AI interaction by its effect on the person using it: a moment when a computational system supplies a judgment the user cannot retrace, and the user has to decide whether to take it on faith. The definition does not depend on how the system works, and it identifies what a supervisor has to look for. On this account, AI literacy begins with the ability to notice such a moment, name it, and pause.

In a law firm the associate decides whether to trust the output, usually alone, and the finished memo does not record the decision. A paragraph the associate wrote after reading the case and one the model wrote from its summary look identical on the page, because the tool writes fluently and nothing marks which passages it produced. Berkeley’s AI rule for students forbids far more than it can detect, and the partner reviewing a memo faces the same problem from the other side. A review of the output can find the fabricated citation, but it cannot see where the model chose the line of authority, framed the issue, or supplied the standard, because the memo presents those choices as the associate’s own. I made the same point about agentic tools, whose intermediate decisions are invisible in the output by design. To the reviewing partner, an associate working with a model is an agent whose intermediate decisions are just as invisible, but the associate, unlike the agent, can be asked.

That is the case for holding the conversation before the review rather than after it. Read first, and the partner is grading the output and inferring the process from it. Talk first, and the partner learns the process from the person who ran it, then reads the output knowing which parts the model produced. The framework in the article follows the same sequence, opening with the learner’s account of their reasoning and their use of the tool and only then testing the evidence behind them, because the reasoning cannot be evaluated until the supervisor knows which parts of it were the learner’s.

The framework, translated

The framework the authors propose, DEFT-AI, adapts a Socratic model that clinical teachers already use in the few minutes between patients: diagnosis, evidence, feedback, teaching, and a closing recommendation for how the learner should engage with AI next time. DEFT is small because it refines the “one-minute preceptor,” a bedside teaching method built for attending physicians with a queue of patients, and the group that proposed it described the four steps as taking “usually in the order of several minutes.” The NEJM authors expect their AI adaptation to be recognizable to any frontline educator. A partner with a queue of matters has no hour for a teaching conference either, but the authors’ premise is that supervision has to fit inside the time the supervisor has.

Each step of the framework has a counterpart in the conversation a partner can have with an associate.

DEFT-AI step The attending asks the resident The partner asks the associate
Diagnosis, discussion, and discourse What is your assessment? Which tool did you use, and what did you ask it? Did the output inform your reasoning or replace it? Before I read it: what is the answer, and why? Which authority controls, and what did you consider and reject? Where did the model come in, what did you ask it, and what did you do with what it gave you?
Evidence How did you verify the output? What evidence supports trusting this tool for this task? Which of these cases did you read yourself? What did you check the model’s statement of the standard against? Where is your corroboration thinnest?
Feedback How do you evaluate your own use of AI here, and how could you improve it? What would you do differently next time, with the tool and without it?
Teaching Focused teaching on the reasoning and on the tool The two or three things the partner now knows the associate needs: the missed authority, the misread holding, the prompt that invited the error
Recommendation for AI engagement How the learner should use AI on this kind of task going forward Which parts of this task the model may draft and which it may only inform, and what verification each requires

One question in the article belongs in every one of these conversations. The authors suggest having the learner reason through a modified version of the case without AI, to see whether the learner can work the problem unaided and to catch overreliance early. The legal version is to ask what the answer is, and why, without the memo in front of either of you. An associate who can answer has done the reasoning, whatever tool they used along the way. If they cannot, they have delegated the judgment along with the task, and the partner knows it before the review begins.

The evidence step is the one lawyers most often skip. The authors want the learner to justify not only the output but the tool: what is known about its accuracy on this kind of task, and where that knowledge comes from. The resident in the article’s illustration answers that they “keep seeing social media posts about how great it is at making diagnoses,” and the associate who says the platform came with the firm’s license has given the same answer. Opinion 512 treats the same question as one of competence under Rule 1.1. The verification the rule requires “will necessarily depend on the GAI tool and the specific task that it performs,” and the opinion’s own example is a lawyer who tests a summarization tool on a sample of contracts, compares the summaries to the documents, and only then relies on it for the rest of the set. The opinion is describing the evidence step, but the honest answer for most legal tools is that nobody at the firm has run the test for the task at hand. The opinion itself cites the Stanford study of the major legal research platforms, which found that even the retrieval-based tools marketed as hallucination-free produced fabricated or mischaracterized authority at rates no partner would accept from an associate.

Why it should be an expectation

The habit I have been encouraging is for the conversation to be expected: something an associate anticipates on every substantial assignment and a partner treats as part of the review. The case for treating it that way rests on what the conversation does that reading the document cannot.

Start with the supervisory duty. Rule 5.1(b) requires a lawyer with direct supervisory authority over another lawyer to make reasonable efforts to ensure that the other lawyer conforms to the Rules, and I have argued that the obligation is matter-specific: it attaches to each filing and each supervised task, and a firm-wide policy does not discharge it. ABA Formal Opinion 512 brings a lawyer’s use of generative AI within that framework. Managerial lawyers “must establish clear policies regarding the law firm’s permissible use of GAI,” and supervisory lawyers “must make reasonable efforts to ensure that the firm’s lawyers and nonlawyers comply with their professional obligations when using GAI tools,” which includes seeing that subordinates are trained “in the ethical and practical use of the GAI tools relevant to their work.” The opinion’s concern is the use, and the only way to supervise the use is to ask about it. The sanctions record points the same way. Sullivan & Cromwell had mandatory training, tracked completion, and a written verification requirement, and its motion still reached the court with dozens of corrupted citations; the sanctioned lawyers of the first quarter had none of those things and produced the same result. In either case, ten minutes with the drafting lawyer would have established whether anyone had checked the citations.

The pedagogical case is where the medical article and my own earlier posts converge. The model produces an answer, and only a person can make the associate defend it. Explaining your reasoning to someone who will push on it is the effortful step that turns information into judgment, the desirable difficulty the tool otherwise removes, and it is the Socratic function that only a human instructor performs. An associate who knows the partner will ask which cases they read, and what they checked the model’s standard against, reads the cases and checks the standard; the expectation changes how the memo gets written before the conversation happens. It also supplies the challenge the tool will not. A model tends to affirm the framing it is given, and an associate who has been working with an agreeable collaborator arrives at the conversation with a conclusion nobody has pushed back on, and the partner is the first person to do it.

The efficiency case persuades partners, because ten minutes with the associate tells the partner whether the associate understood the assignment, what the conclusion is and whether it is right, which parts the model drafted, and where the verification is thin. The partner then reads the document knowing where to look: the section the associate wrote from the cases gets a normal read, the section the model drafted from a summary gets a close one, and the standard the associate could not source gets checked first. The conversation catches the memo that answered the wrong question before the partner has read it, and it catches the associate who is in over their head, who will not announce the fact and whose model will not announce it either. Against a review that can run to hours, ten minutes is the cheapest supervisory step a partner has.

Centaur and cyborg on a deal team

The framework ends with a recommendation about how the learner should work with AI going forward, and the authors use two figures for the choice. A centaur divides the work: the model gathers, summarizes, or drafts, and the human decides, evaluating the output before relying on it. A cyborg interleaves: the human prompts, corrects, asks for justification, and refines the output jointly with the model, step by step. The authors do not rank them: centaur work is for high-stakes or uncertain tasks and for any task the tool has not been validated to perform; cyborg work is for low-risk, well-defined, or creative tasks where the tool’s performance is known. The competence they want to build is the ability to switch, task by task, on the basis of the stakes and the tool, and they add that a learner working in cyborg mode must be able to justify the approach to a supervisor when asked. Opinion 512 makes the same distinction when it says that a tool built for a discrete legal task “may require less independent verification or review, particularly where a lawyer’s prior experience with the GAI tool provides a reasonable basis for relying on its results,” but that lawyers “may not abdicate their responsibilities by relying solely on a GAI tool to perform tasks that call for the exercise of professional judgment.”

The pairing matches the distinction I drew between delegating a task and delegating judgment, and adds the stakes, which the earlier framing left out. Drafting a client status email or producing a first cut of a closing checklist is cyborg work, where an associate can iterate with the model and a light review suffices. Centaur work is the research memo on a dispositive issue, the diligence summary that becomes a disclosure schedule, and the brief that will go out under the firm’s name, where the model’s output is a lead to be verified against authority the associate consulted independently, which is the standard I proposed in June. Until more tools have been tested for the task, most substantive legal work is centaur work, and the closing step of the conversation is where a partner says so. The associate leaves knowing which mode the next assignment calls for.

Verify and trust

The authors close on a phrase that reverses the familiar maxim: verify and trust, with the check put first. Physicians reached the verification standard independently, and they reached the same supervisory method, a structured conversation held at the moment the trainee took the leap of faith. The apprenticeship that trained lawyers for a century ran on work the tools now do, and the training has to follow the work to the verification. Ten minutes of conversation is where a partner can watch that happen.


This post draws on Raja-Elie E. Abdulnour, Brian Gin & Christy K. Boscardin, Educational Strategies for Clinical Supervision of Artificial Intelligence Use, 393 New Eng. J. Med. 786 (2025), and three of the studies it relies on: Fabrizio Dell’Acqua et al., Navigating the Jagged Technological Frontier (Harvard Bus. Sch. Working Paper No. 24-013, 2023); William J. Waldock et al., Which Curriculum Components Do Medical Students Find Most Helpful for Evaluating AI Outputs?, 25 BMC Med. Educ., art. 195 (2025); and Feiyang Yu et al., Heterogeneity and Predictors of the Effects of AI Assistance on Radiologists, 30 Nature Med. 837 (2024). The DEFT model itself comes from Michael C. Savaria et al., Enhancing the One-Minute Preceptor Method for Clinical Teaching with a DEFT Approach, 115 Int’l J. Infectious Diseases 149 (2022). On legal research tools, it draws on Varun Magesh et al., Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, 22 J. Empirical Legal Stud. 216 (2025). The professional-responsibility discussion relies on Model Rules of Professional Conduct r. 5.1 and 5.3 and A.B.A. Committee on Ethics & Professional Responsibility, Formal Opinion 512 (2024). It builds on earlier posts on privilege and the consumer chatbot, the verification standard, the delegation framework, sycophancy, agentic AI and supervision, the Sullivan & Cromwell errata, Q1 2026 citation sanctions, answer quality versus learning, the Bednar trial, and three schools’ AI policies.