EconBase
← Back to paper

Overcoming Medical Overuse with AI Assistance: An Experimental Investigation

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

67,585 characters · 12 sections · 37 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Overcoming Medical Overuse with AI Assistance: An Experimental Investigation

abstractThis study evaluates the effectiveness of Artificial Intelligence (AI) in mitigating medical overtreatment, a significant issue characterized by unnecessary interventions that inflate healthcare costs and pose risks to patients. We conducted a lab-in-the-field experiment at a medical school, utilizing a novel medical prescription task, manipulating monetary incentives and the availability of AI assistance among medical students using a three-by-two factorial design. We tested three incentive schemes: Flat (constant pay regardless of treatment quantity), Progressive (pay increases with the number of treatments), and Regressive (penalties for overtreatment) to assess their influence on the adoption and effectiveness of AI assistance. Our findings demonstrate that AI significantly reduced overtreatment rates—by up to 62% in the Regressive incentive conditions where (prospective) physician and patient interests were most aligned. Diagnostic accuracy improved by 17% to 37%, depending on the incentive scheme. Adoption of AI advice was high, with approximately half of the participants modifying their decisions based on AI input across all settings. For policy implications, we quantified the monetary (57%) and non-monetary (43%) incentives of overtreatment and highlighted AI's potential to mitigate non-monetary incentives and enhance social welfare. Our results provide valuable insights for healthcare administrators considering AI integration into healthcare systems.

\\

quote\footnotetext[1]{We thank Mengyuan Wu, Shapeng Jiang, Yanyu Li, Xiangyu Gao, Qiuhao Chen, Wenjun Li for their assistance with the experimental sessions. Lijia Wei acknowledges support from National Science Foundation China (72173093), and the Center for Behavioral and Experimental Research (CBER) at Wuhan University. Lian Xue acknowledges support from the Research Funds for Youth Academic Team in Humanities and Social Sciences of Wuhan University (413000425). The authors declare no conflict of interest.} “The psychologist Gerd Gigerenzer has a simple heuristic. Never ask the doctor what you should do. Ask him what he would do if he were in your place. You would be surprised at the difference” --- Nassim Nicholas Taleb, Antifragile: Things That Gain from Disorder

Introduction

Overtreatment in healthcare, defined as unnecessary medical interventions that can harm patients, represents a significant economic and health concern. It leads to billions of dollars in wasted spending and exposes patients to unnecessary physical and psychological risks muskens2022overuse. This phenomenon is well-documented in both the clinical and health economics literature. Examples of overtreatment include the excessive use of antibiotics for viral infections --medical overuse or overprescription-- and the overuse of imaging tests, such as CT scans and MRIs, for conditions that could be managed with less invasive methods morgan2015setting. These practices not only strain healthcare budgets but also expose patients to unnecessary physical and psychological risks. A recent Lancet study highlights the pervasiveness of overtreatment, attributing 29% of healthcare spending in the US to it, with costs reaching as high as 89% in certain populations globally, and rising in low- and middle-income countries wennberg2002geography,korenstein2012overuse,brownlee2017evidence,albarqouni2023overuse.

A major hurdle in studying overtreatment is the difficulty of measuring overuse accurately brownlee2017evidence. Existing methodologies employed both direct and indirect measures. Direct approaches involve using patient registries and medical records guided by evidence-based guidelines or expert consensus panels, such as the RAND Appropriateness Method fitch2000rand. These require a precise definition of the “appropriate care”, which often lacks clarity in many clinical settings. Furthermore, guidelines typically do not provide the necessary details for individual patient care, challenging the case-specific flexibility in measuring overtreatment chassin1998urgent,korenstein2012overuse. Indirect approaches involve the identification of unexpected variations in healthcare service utilization both within and across countries in the healthcare system. For example, excessive use of insulin --compared to the means-- for patients with diabetes could be considered as medical overuse brownlee2017evidence. However, these methods struggle with standardization and often yield inconclusive results.

This paper presents a novel experimental approach designed to measure the medical overuse tendencies of prospective physicians (medical students) from a behavioral perspective. We construct a virtual doctor-patient consultation scenario, using cases from standard practice exam questionnaires that simulate the patient's illness presentation. Given the description of the illness scenario, participants were instructed to select the most appropriate medicine from a set of five options. They were informed that there was only one correct answer but could choose up to two responses (medicines) for each question (illness). To incentivize accurate diagnosis and mimic real-world interactions, each choice is associated with payoffs for both the participant (physician) and a stylised patient. Specifically, participants' choices are linked to a donation to a patient-regarding charity, where overuse of medicine is associated with a lower amount of donation compared to one exact correct treatment. We refer to this task as the medical prescription task. This setup has several advantages and complements existing measurements of medical overuse: it can be easily generalized to all medical services with category-specific questions; it is anonymous and harmless, with respect to physicians' reputations and patients' well-being.

Utilizing the medical prescription task framework in randomized controlled trials, we aim to disentangle the underlying motives for physicians' propensity to overuse medicines (causes) and explore effective methods to curb medical overuse (interventions), which represent the two primary objectives of this research. The causes of medical overuse can be summarized into two categories: monetary and non-monetary. Monetary incentives, such as fee-for-service payment models, often encourage physicians to the volume of care delivered rather than its quality pauly2012handbook,morgan20192019,lu2014insurance. Non-monetary incentives may involve defensive medicine, where doctors prescribe unnecessary treatments to avoid litigation summerton1995positive, a lack of adequate knowledge, leading to decisions that do not align with the latest clinical guidelines price1986doctors,bishop2010physicians,jatoi2019clinical, and out of demand pressure from patients when there is price reduction in the medicines lopez2018contribution. Systematic reviews indicate that medical overuse and its underlying incentives are deeply embedded within healthcare systems, necessitating further exploration of identification and intervention methods korenstein2012overuse,tung2018factors,segal2022factors.

In our experiment, we exogenously vary three types of incentive schemes for prospective physicians: a flat scheme where participant payoff remains unaffected by the choices made in the medical prescription task; a progressive scheme, where payoff increases with the number of choices selected; and a regressive scheme where overtreatment resulted in reduced payoff, contingent upon whether the correct medicine was selected. Patients' payoffs remained constant across all three treatments, with their welfare maximized when exactly one correct medicine was selected by the (prospective) physician (\ding{51}), followed by `one correct plus one incorrect medicine' (\ding{51}\ding{55}). `One incorrect' (\ding{55}) and `two incorrect' (\ding{55}\ding{55}) choices progressively decreased welfare. Notably, in the Regressive treatment, physicians' payoffs followed the same pattern as patients'. In this case, we say that physicians and patients have aligned interests, mirroring the quote from `Antifragile' at the beginning of this paper taleb2014antifragile.

Additionally, we explore the impact of AI decision-support on physician choices within each treatment. For each medical prescription task, participants made an initial decision on their own. Subsequently, they were presented with a descriptive analysis labelled “AI diagnosis” along with a specific treatment suggestion termed “AI choice” before making a second choice. This AI advice --diagnosis and choice-- was generated in advance using the large language model, ChatGPT 4.0, which achieved an accuracy rate of 73.93% across more than 800 evaluations openai. In real life, AI has the potential to significantly mitigate overtreatment by offering evidence-based recommendations that enhance clinical decision-making jiang2017artificial,topol2019high. AI systems can analyze vast amounts of data to identify the most appropriate treatments, potentially helping to avoid unnecessary medical procedures pine2012harnessing,obermeyer2016predicting. Existing literature on AI-assisted decision-making has demonstrated that these technologies can enhance the accuracy of diagnoses and tailor treatment plans to individual patients niel2019artificial, davenport2019potential,sung2024artificial. However, implementing AI in healthcare still faces challenges such as ethical and equity concerns, patient's trust towards AI diagnosis and physicians' acceptability of algorithms dietvorst2015algorithm,char2018implementing, emanuel2019artificial,glauser2020ai,al2023review. Our paper contributes by providing the first experimental evidence of the effectiveness of AI on overtreatment and corresponding welfare consequences.

Our results were centred around two key performance measures: the quantity of prescribed medicines and the quality of task accuracy. For quantity measures, exploiting the exogenous variations in incentive schemes and the provision of AI assistance in medical prescription tasks, we could identify both the monetary (incentive schemes) and non-monetary (AI-enhanced decision accuracy and others) influences on overtreatment. We observed that overtreatment rates among participants ranged from 24-58% across all treatments. Further analysis revealed that more than half (57%) of the overtreatment was due to monetary incentives, which can be reduced when overtreatment is penalized with a regressive incentive scheme. AI significantly reduce overtreatment in all settings, with the most pronounced effect observed in the regressive treatment, achieving a reduction rate of 62%. We refer to this phenomenon as the Synergistic AI Effect, highlighting how AI's effectiveness in reducing medical overuse is maximized when physicians and patients have aligned interests.

Regarding quality, we found that a progressive incentive encouraging more prescriptions could increase accuracy because participants included more choices in their answers compared to a fixed payment-- i.e., overuse as a conservative treatment. The introduction of AI significantly boosted the accuracy rate across all treatments, with the largest effect in the Flat treatment with constant payoffs --an enhancement rate of 37%. \footnote{almog2024ai finds that umpires lowered their overall mistake rate after the introduction of Hawk-Eye review in top tennis tournaments. The AI assistance in our paper is functionally similar to Hawk-Eye.} Intriguingly, the positive impact of AI can be further differentiated into two distinct channels: an increase in the provision of `one exact medicine \ding{51}' (termed the Optimal Treatment Adjustment) and the refinement of `one correct plus one incorrect medicine \ding{51}\ding{55}' (referred to as the Precision-Compromise Strategy). These effects vary significantly across different incentive schemes, with the regressive treatment most favoring the Optimal Treatment Adjustment over the Precision-Comparomise Strategy. In other words, the effect of AI on reducing redundant inaccurate prescriptions is most pronounced in the regressive treatment.

This study enhances the existing body of research by empirically examining the effectiveness of AI interventions in curbing overtreatment in healthcare decisions. By analyzing overtreatment and accuracy rates before and after the deployment of AI recommendations, our findings offer robust evidence supporting AI's role in minimizing unnecessary medical practices and improving diagnosis accuracy kelly2019key,chen2023framework. Furthermore, this research advances our comprehension of AI's potential integration into clinical workflows, thereby improving decision-making processes wang2019ai,recht2020integrating. Specifically, it aligns with the literature that discusses the modification of physician incentives, suggesting that AI can mitigate the incentive-driven overtreatment in flat or pay-for-service incentives (progressive treatment) and an advanced effect in the regressive incentive scheme, thus addressing issues related to incentive-driven overutilization clemens2014physicians,scott2002incentives,ettner2012role,xi2019does,vlaev2019changing. Additionally, it contributes to digital healthcare literature by demonstrating how technology-driven solutions can effectively supplement traditional healthcare practices, leading to more efficient and patient-centered care sheikh2021health, mesko2017digital,iqbal2022regulatory,ibeneme2022strengthening.

Policy considerations should prioritize exploring the interaction effect of AI and incentive schemes on physician behavior. This research highlights the potential for AI to mitigate overtreatment, and its effectiveness could be further amplified by aligning physician incentives with AI adoption. Policymakers can explore alternative payment models that reward physicians who utilize AI tools and make treatment decisions aligned with AI recommendations and patients' welfare, particularly when such recommendations discourage overtreatment. Furthermore, given the challenges of implementing new incentive schemes in the healthcare industry agrawal2022power, our results suggest an alternative intervention that is as effective as introducing a pay-by-performance ({Reg}) incentive scheme at ensuring healthcare quality -- i.e., to incorporate AI within existing flat (pay-by-visit) incentive structure. Finally, our experiment provides an example of clinician-AI collaboration, which advocates policies that foster trust and ensure that AI complements, rather than replaces, physician expertise.

The remainder of this paper is organized as follows: Section (ref) presents the experimental design and procedures; Section (ref) reports results, and Section (ref) concludes with a discussion.

Experimental Design and Procedure

We conducted a lab-in-the-field experiment at a medical school affiliated with Central South Hospital, utilizing a two-by-three factorial design. Participants in this study were prospective physicians recruited from this institution, where they are being trained as medical doctors.

\paragraph{Incentive Schemes (Between-Subject)} The experiment varied the incentive structures for participants during the main task. We tested three types of incentives: {Flat}, {Prog}ressive, and {Reg}ressive in a between-subject design.

\paragraph{AI Assistance (Within-Subject)} We also varied the decision-making process by introducing an AI assistance feature. Participants made initial choices independently and subsequently could remake their decisions after viewing AI-generated recommendations. The AI advice, powered by ChatGPT 4.0, included both a specific recommendation and an extended explanation of the optimal medication choice based on patient symptoms.

\paragraph{Control Variables} To account for individual fixed effects and other potential confounders affecting the propensity for overtreatment, we implemented pre- and post-experiment assessments. These assessments evaluated various attributes, including physicians' professional abilities, cognitive reflection frederick2005cognitive, cognitive ability carpenter1990one, risk preference blais2006domain, altruism preference and trust falk2018global, algorithm literacy, trust in algorithms, awareness of algorithms, and perceptions of algorithm fairness. \footnote{We also conducted a standard array of questionnaires, including gender, major of study, age, grade, and education. Instructions for the experiment and for the control question are available in the Appendix (ref). }

Medical Prescription Task

Participants were asked to answer 20 multiple-choice questions that simulated a real-life doctor-patient consultation, with basic patient information and disease descriptions from the “China's National Qualifying Examination for Medical Practitioners” \footnote{Medical students in China need to obtain the “National Practicing Physician Qualification Certificate" through the examination before they are qualified to practice medicine and can officially join the profession.} question bank, which consists of 817 questions. Each question had only one correct answer, but participants could select up to two answers to simulate potential overtreatment scenarios. Feedback on their accuracy was withheld until the end of the experiment. An example of such a question is provided below:

quoteFemale, 54 years old, with a history of hyperthyroidism. Recently, due to overwork and emotional stress, she has experienced insomnia, and heart and chest discomfort. Physical examination: Heart rate 160 beats per minute, electrocardiogram shows clear signs of myocardial ischemia, and sinus rhythm irregularity. The best choice would be: A. Amiodarone \\ B. Quinidine \\ C. Procainamide \\ D. Propranolol \\ E. Lidocaine

To examine participants' innate ability at the medical prescription task, they were instructed to complete a separate set of 10 medical prescription tasks as a practice at the beginning of the experiment, with each correct choice yielding a 1 yuan bonus.

Treatment Design

We vary how participants were incentivized and whether AI assistance was provided during the medical prescription task. The treatments regarding incentive schemes were structured as follows:

itemize• {Flat}: Participants received a constant monetary payoff of 3 yuan for each medial prescription task completed, regardless of their choice. • {Prog}ressive: Participants received a higher payoff for selecting two options (4 yuan) versus only one (2 yuan) option, irrespective of correctness. • {Reg}ressive: Payoffs varied depending on the accuracy and number of choices: 6 yuan for one correct choice (\ding{51}), 4.5 yuan for two choices including the correct one (\ding{51}\ding{55}), 1.5 yuan for one incorrect choice (\ding{55}), and 0 yuan for two incorrect choices (\ding{55}\ding{55}).

{Patients' Payoffs. } Upon successful completion of each medical prescription task, we donated the patients' payoff generated by the participants to a patient-regarding charity, specifically the Tencent Public Welfare “Love Angel” project, which supports children suffering from leukemia and cancer. The payoff structure for the patients was designed as follows: a base payoff of $P$ yuan was established, with an increase of $B$ yuan for correctly prescribed medicine and a decrease of $C$ yuan for incorrectly prescribed medicine. For the purposes of our experiment, the values were set at $P = 4$, $B = 4$, and $C = 2$. Notably, in the Regressive treatment, physicians' and patients' payoffs were exactly collinear. This means that physicians made decisions for patients as if they were making decisions for themselves from a payoff perspective. The detailed experimental design, including incentive schemes and patients' payoffs, is summarized in Table (ref).

table[table omitted — 1,332 chars of source]

\noindentAI assistance. Participants made the choice twice in the medical prescription task: an initial choice without AI and a second choice with AI assistance. The AI advice, powered by ChatGPT 4.0 in early 2024, entails two parts: an exact recommendation of the optimal medication choice given the patient's symptoms and an extended explanation of the recommended choice. An example of the AI assistance is as follows:

quoteThe AI recommends option: D. \\ Analysis: Based on the patient's symptoms and physical examination results, the optimal medication choice is usually D. Propranolol. This beta-blocker is commonly used to control rapid heart rates and arrhythmias, particularly in hyperthyroidism cases. It helps slow down the heart rate, reduce the cardiac workload, and alleviate myocardial ischemia symptoms. However, a specific treatment plan should be consulted with a physician.

To best incentivize participants' effort at both the initial and revised decisions, participants' payoffs were determined by either their initial or second choice, randomized at the session level by the computer.

Procedures

The experiment was conducted at the affiliated medical school of Central South Hospital. A total of 120 medical students were recruited for the study, with 40 students participating in each of the three treatment groups. These students were part of the subjects pool for the Center for Behavioral and Experimental Research (CBER) from Wuhan University. Upon arrival, participants were allocated to one of two seminar rooms on campus, with each session comprising 5 participants. Each session lasted between 50 to 60 minutes. All tasks were computerized using the oTree platform chen2016otree. The average payoffs were 104.7 yuan: 95.9 yuan for the {Flat}, 93.8 yuan for the {Prog}, and 124.3 yuan for the {Reg} treatment respectively.

Results

Overview

This section outlines the two key outcome variables that assess the performance of doctors in terms of quantity and quality.

\paragraph{Quantity} We define overuse as the selection of two answers in each medical prescription task. Given that there is a single correct answer and participants were fully aware of this fact, any response with more than one answer is categorized as overtreatment.

\paragraph{Quality} Our main measurement of quality is the accuracy of the diagnosis, which hinges on whether participants chose the correct answer. Without loss of generality, let us denote the correct choice as \( A \). The accuracy rate is then defined as the proportion of choices that include at least one \( A \). This can be mathematically expressed as:

equation[equation omitted — 81 chars of source]

where \( \complement_{A} \equiv U \setminus A = B \cup C \cup D \cup E \), and \( U \) represents the universal set of all possible choices.

In addition to the performance measures, we examine the extent to which participants adopt or trust AI advice subsequent to their initial choices. Participants initially made a choice without AI guidance, after which they were presented with the AI recommendation. Subsequently, they were prompted to make a second choice. Compensation for the session was determined randomly, based on either the first or second choice, as selected by the computer.

\paragraph{AI adoption} is characterized by whether participants altered their choices following the AI advice. We employ two measures of AI adoption: a flexible measure and a more restrictive measure. Under the flexible measure, maintaining the initial answer despite AI's recommendation to change is classified as not following AI advice, whereas all others are considered following. For instance, if the first choice was \( B \), AI suggested \( A \), and the second choice remained \( B \), this is deemed as not following AI advice. Conversely, any other alteration is categorized as following. The more restrictive measure defines only active changes that align with AI advice as AI adoption. All other instances are considered non-adoption. For example, a sequence of $B-A-B$ is defined as not following AI advice as with the main measurement. However, a sequence of $A-A-A$ is defined as not following since there are no active changes in choices. Table (ref) presents an exhaustive list of scenarios and their corresponding classifications regarding AI adoption. Table (ref) reports the summary statistics of the three key outcome variables before and after AI intervention across treatments.

table[table omitted — 1,645 chars of source]
table[table omitted — 1,876 chars of source]

The effect of AI on medical overuse

result[AI on Overuse] The use of AI can significantly reduce the propensity of medical overuse among medical students, with an overall reduction effect of 37% across all treatments. This effect is the most pronounced in the {Reg} treatment.
proof[Support.] Support for Result (ref) can be seen in Table (ref), Figure (ref) and Table (ref). Table (ref) highlights the substantial effect of AI on reducing overuse across different incentive scheme treatments. Notably, the transition from pre-AI to post-AI conditions shows marked improvements: in the {Flat} treatment, overuse decreased from 24.3% to 16.0%, a deduction rate of 34%. Similar positive shifts are observed in {Prog} and {Reg} treatments, with significant reduction rates of 26% and 62% in overuse, respectively, all supported by statistically significant $p$-values. \footnote{Across all treatments, the rate of overuse is decreased from 36.1% to 22.9% from pre- to post-AI decisions. The overall deduction rate is at 36.7%.} Figure (ref) visually complements these findings by illustrating the deduction of overuse rate in post-AI conditions. The left figure depicts the average instances of overuse times per 20 rounds, and the right figure further decomposes the average overuse proportion among participants by round number. Overall, we see a consistent deduction in overuse after the integration of AI in decision-making throughout the whole experiment. Finally, Table (ref) provides a detailed regression analysis showing the statistical impact of AI on medical overuse. The coefficients for PostAI across various models consistently indicate significant decreases in overtreatment. Model 5 further reveals the nuanced interaction between incentive schemes and the effect of AI; it can be seen from the table that the deduction effect of AI is almost twice as much as in {Reg} compared to the {Flat} treatment ($p<0.1$). The effects remained robust after controlling for individual characteristics and professional ability. \footnote{Recall that professional ability or innate ability is measured by the accuracy rate in the 10 practice questions at the beginning of the experiment.} \begin{figure}[h!] \caption{Medical overuse by treatment} \end{figure} \begin{table}[h!] \def\sym#1{\ifmmode^{#1}\else\(^{#1}\)\fi} \caption{The effect of AI on frequency of medical overuse: Random effect models} \begin{threeparttable} \begin{tabular}{l*{5}{c}} \toprule &\multicolumn{1}{c}{(1)}&\multicolumn{1}{c}{(2)}&\multicolumn{1}{c}{(3)}&\multicolumn{1}{c}{(4)}&\multicolumn{1}{c}{(5)}\\ Dep Var: Overuse freq &\multicolumn{1}{c}{\textit{{Flat}}}&\multicolumn{1}{c}{\textit{{Prog}}}&\multicolumn{1}{c}{\textit{{Reg}}}&\multicolumn{1}{c}{Pooled}&\multicolumn{1}{c}{Pooled}\\ \midrule PostAI & -1.650***& -3.050***& -3.250***& -2.650***& -1.650***\\ (\textit{Baseline: PreAI}) & (0.514) & (0.921) & (0.651) & (0.413) & (0.532) \\ [1em] Progressive & & & & & 6.818***\\ (\textit{Baseline: Flat}) & & & & & (1.229) \\ [1em] Regressive & & & & & -1.008 \\ & & & & & (1.525) \\ [1em] Progressive $\times$ PostAI& & & & & -1.400 \\ & & & & & (1.091) \\ [1em] Regressive $\times$ PostAI& & & & & -1.600* \\ & & & & & (0.858) \\ [1em] Ability & & & & & -0.840***\\ & & & & & (0.269) \\ [1em] \multicolumn{5}{l}{Control for individual characteristics} & \checkmark \\ Constant & 4.850***& 11.630***& 5.200***& 7.225***& 26.180***\\ & (0.762) & (1.049) & (0.782) & (0.575) & (7.545) \\ \midrule \textit{Mean of Y} &4.025& 10.100&3.575 &5.900 &5.900\\ Observations & 80 & 80 & 80 & 240 & 240 \\ $R^2$ & 0.033 & 0.046 & 0.138 & 0.045 & 0.372 \\ \bottomrule \end{tabular} \begin{tablenotes} • {\textit{Notes. }This table reports the effect of incentive schemes and AI assistance on medical overuse using random effect models. Unit of observation is individual \emph{frequency of medical overuse out of 20 rounds}, pre- and post-AI. Model 5 additionally controls for other individual characteristics, including gender, age, major of study, cognitive reflection, cognitive ability, risk preference, altruism preference, trust, algorithm literacy, trust in algorithms, awareness of algorithms, and perceptions of algorithm fairness. All standard errors are clustered at the subject level and are reported in parentheses. In this and the following table $* p<0.10, ** p<0.05, *** p<0.010$. } \end{tablenotes} \end{threeparttable} \end{table} \textbf{Causes of overuse. }Recall that the causes of medical overuse are classified into two categories: monetary and non-monetary. Using the \textit{{Flat}} and \textit{{Reg}} treatment, we established a baseline measure of the non-monetary incentives of overtreatment, as in these two treatments, overuse is associated with either no extra or even negative rewards. Indeed, we see a similar rate of overuse in these two treatments before the implementation of AI; the additional overuse in the \textit{{Prog}} treatment shall then be attributed to monetary incentives. Using back-in-the-envelop calculations, we estimate the non-monetary incentive accounts for around 43% of medical overuse, while the monetary incentive accounts for 57%.

The effect of AI on healthcare quality

resultThe use of AI significantly increased the accuracy rate in the medical prescription task, with improvements ranging from 17% to 37%. The effect of AI on minimising unnecessary incorrect choices is strongest in the {Reg} treatment.
proof[Support. ] Support for Result (ref) is provided by Table (ref), Table (ref) and Figure (ref). Table (ref) details the improvement in accuracy rates across treatments. In the {Flat} treatment, the accuracy rate increased from 56.3% pre-AI to 77.3% post-AI, demonstrating a substantial 37% enhancement in decision-making quality. Similarly, in the {Prog} and {Reg} treatments, the accuracy rates increased by 23% and 17%, respectively, after AI assistance. \footnote{Across all three treatments, the accuracy rate of the medical prescription task increased from 12.6% to 15.8% after the introduction of AI advice. The overall accuracy enhancement effect of AI is thus at 25%.} Table (ref) provides a detailed regression analysis showing the statistical impact of AI on accuracy. The coefficients for PostAI across various models consistently indicate significant improvements in accuracy. Model 4 further reveals the nuanced interaction between incentive schemes and the effect of AI; it can be seen that the accuracy improvement effect of AI is more pronounced in the {Flat} treatment compared to the {Reg} treatment ($p < 0.1$). The effects remain robust after controlling for individual characteristics and professional ability (Model 5). \begin{table}[h!] \def\sym#1{\ifmmode^{#1}\else\(^{#1}\)\fi} \caption{The effect of AI on accuracy frequency: Random effect models} \begin{threeparttable} \begin{tabular}{l*{5}{c}} \toprule &\multicolumn{1}{c}{(1)}&\multicolumn{1}{c}{(2)}&\multicolumn{1}{c}{(3)}&\multicolumn{1}{c}{(4)}&\multicolumn{1}{c}{(5)}\\ Dep Var: Accuracy freq &\multicolumn{1}{c}{\textit{{Flat}}}&\multicolumn{1}{c}{\textit{{Prog}}}&\multicolumn{1}{c}{\textit{{Reg}}}&\multicolumn{1}{c}{Pooled}&\multicolumn{1}{c}{Pooled}\\ \hline PostAI & 4.200***& 3.050***& 2.275***& 4.200***& 4.200***\\ (\textit{Baseline: PreAI}) & (0.535) & (0.546) & (0.501) & (0.532) & (0.553) \\ [1em] Progressive & & & & 2.200** & 1.514** \\ (\textit{Baseline: Flat}) & & & & (0.868) & (0.601) \\ [1em] Regressive & & & & 1.975** & 0.924 \\ & & & & (0.818) & (0.635) \\ [1em] Progressive $\times$ PostAI& & & & -1.150 & -1.150 \\ & & & & (0.761) & (0.791) \\ [1em] Regressive $\times$ PostAI& & & & -1.925***& -1.925** \\ & & & & (0.729) & (0.757) \\ [1em] Ability & & & & & 0.239** \\ & & & & & (0.105) \\ [1em] \multicolumn{5}{l}{Control for individual characteristics} & \checkmark \\ Constant & 11.250***& 13.450***& 13.225***& 11.250***& 12.985***\\ & (0.611) & (0.621) & (0.549) & (0.609) & (1.981) \\ \midrule \textit{Mean of Y} &13.350 & 14.975&14.363 &14.229 &14.229\\ Observations & 80 & 80 & 80 & 240 & 240 \\ $R^2$ & 0.361 & 0.219 & 0.167 & 0.294 & 0.569 \\ \bottomrule \end{tabular} \begin{tablenotes} • {\textit{Notes. }This table reports the effect of incentive schemes and AI assistance on accuracy frequency using random effect models. Unit of observation is individual \emph{frequency of being accurate out of 20 rounds}, pre- and post-AI. Model 5 additionally controls for other individual characteristics, including gender, age, major of study, cognitive reflection, cognitive ability, risk preference, altruism preference, trust, algorithm literacy, trust in algorithms, awareness of algorithms, and perceptions of algorithm fairness. All standard errors are clustered at the subject level and are reported in parentheses. } \end{tablenotes} \end{threeparttable} \end{table} \textbf{Mechanisms of Improvement. }Finally, Figure (ref) visually dissects two potential channels of accuracy improvement, elucidating the mechanisms through which AI enhances decision-making: \begin{enumerate} • \textbf{Increase of one exact answer (\ding{51})} - This improvement reflects the adoption of what we term the \textit{Optimal Treatment Adjustment}, where participants aim for the most exact and accurate strategy. • \textbf{Increase of one correct plus one incorrect answer (\ding{51}\ding{55})} - This pattern suggests a \textit{Precision-Compromise Strategy}, where participants make conservative choices, including a mix of correct and incorrect options. \end{enumerate} \begin{figure}[h!] \caption{Accuracy rate in the medical prescription task} \end{figure} The figure highlights that the relative significance of these enhancement channels varies across treatments. We quantify the relative contribution of the two accuracy enhancement mechanisms in each treatment as follows. In the \textit{{Flat}} treatment, the Optimal Treatment Adjustment (Channel 1) accounts for 109.5% of the total enhancement effects, while the Precision-Compromise Strategy (Channel 2) detracts by -9.5%. In the \textit{{Prog}} treatment, Channel 1 contributes 140%, and Channel 2 reduces the enhancement by -40%. Finally, in the \textit{{Prog}} treatment, Channel 1 leads to an enhancement of 184%, with Channel 2 diminishing the effect by -84%. Overall, AI predominantly enhances decision-making through Optimal Treatment Adjustment, slightly offset by a crowing-out effect on the Precision-Compromise Strategy. The influence of AI in reducing the unnecessary incorrect choice is the most pronounced in the \textit{{Reg}} treatment.

Incentive scheme and social welfare

result[Incentive Schemes and Accuracy] Compared to {Flat}, a {Prog} or a {Reg} incentive scheme significantly increases the initial accuracy rate before participants receive AI advice. However, the influence of the incentive schemes on performance becomes less pronounced once AI advice is introduced.
proof[Support. ] Support for this result comes from Table (ref), Table (ref) and Table (ref) (Panel B). Prior studies often advocate that performance-based pay can enhance physician performance more effectively than service-based pay schemes heider2020effects. Our results in the PreAI conditions can test this hypothesis in the context of the medical prescription task context. As shown in Table (ref), in the {PreAI} conditions, the accuracy rate is 56.3% in {Flat} which is lower than in {Prog} (67.3%) and {Reg} (66.1%). Table (ref) (Model 4) complements this comparison by showing that in PreAI conditions, the accuracy rate in {Flat} treatment is significantly lower compared with the {Prog} and the \textit{{Reg}} treatment. These comparisons remain robust in \textit{{Prog}} treatment after controlling for individual characteristics and ability, but not for \textit{{Reg}} treatment. These results suggest the benefit of medical overuse: i.e., the more conservative choices (\ding{51}\ding{55}) under uncertainty could essentially increase the overall accuracy. To delve deeper, in Table (ref) panel B, we break down participants' choice types by treatment. Compared with the baseline (\textit{{Flat}}), the pay-by-performance incentive scheme where participants got punished with incorrect answers (\textit{{Reg}}) significantly increased the exactly correct choice (\ding{51}) from 41% to 48%. Conversely, a pay-by-service incentive scheme where participants were rewarded with the number of medicines prescribed regardless of accuracy (\textit{{Prog}}) significantly boosted the accuracy rate through conservative choices (\ding{51}\ding{55}). Post-AI introduction, the effects of these incentive effects are mitigated, as seen from the significant negative interaction terms in the \textit{{Reg}} treatment in Table (ref), indicating that the quality-enhancing effects of a performance-contingent incentive scheme are equated after the introduction of AI. In other words, the effect of AI on enhancing healthcare quality is as effective as introducing a performance-based incentive scheme. \footnote{Additional evidences can be seen from Table (ref). In \textit{{PreAI}} conditions, the accuracy rate is significantly lower in \textit{{Flat}} than the \textit{{Reg}} treatment ($p=0.014$). \textit{{PostAI}}, the accuracy rate disparity disappeared ($p=0.969$, \textit{{Flat}} vs \textit{{Reg}}, \textit{{PostAI}}).} \textbf{Discussion:} The literature consistently highlights that while performance-based incentives can motivate doctors to improve care quality, they must be carefully designed to avoid unintended consequences such as overtreatment or neglect of non-incentivized activities einav2018provider,petersen2006does,mcdonald2007impact,kovacs2020pay,fainman2020design. \footnote{See van2010systematic for a systematic review.} From a practical perspective, even if altering medical incentive schemes is effective at promoting patient care quality, it may be difficult to implement in hospitals due to system inertia agrawal2022power. Our results indicate that the introduction of AI appears to level the playing field among different incentive schemes in terms of their influence on physician performance (accuracy rate), simplifying incentive structures without compromising care quality.
result[Incentive Schemes and Social Welfare] Patients' payoffs are highest under the {Reg} treatment, followed by {Flat} and {Prog} treatments: $\textit{\scshape{Reg}} > \textit{\scshape{Flat}} = \textit{\scshape{Prog}}$. For physicians, the {Reg} treatment consistently yields the highest payoffs, with the other two treatments being equal to each other: $\textit{\scshape{Reg}}>\textit{\scshape{Prog}}=\textit{\scshape{Flat}}$. Social welfare is maximized under the {Reg} treatment, which is particularly significant when considering different scenarios of benefits and costs of correct/incorrect treatments.
proof[Support.] Support for the result can be seen in Table (ref). Previous results have shown the accuracy enhancement and overuse reduction effects of AI assistance; we, therefore, would reasonably expect welfare enhancement after AI, conditional on the AI advice being costless. In this section, we will quantify the welfare enhancement of AI under each incentive scheme. On the other hand, the effect of incentive schemes on welfare is nuanced. Since overtreatment could potentially increase the accuracy rate, it is uncertain that a {Prog} incentive schemes that promote overtreatment would make the patients better or worse off overall. We, therefore, compare the effects of incentive schemes on patients' welfare before and after AI implementation with different costs of overtreatment. Table (ref) breaks down the welfare analysis under different incentives by various combinations of benefit ($B$) and costs ($C$) of correct(incorrect) treatment. Although we implemented the first scenario in our experiment, the pattern of the flat, progressive and regressive incentive schemes remains consistent across different scenarios. \footnote{We derive the payoffs in the other scenarios in the following method: We first compute the hypothetical patients' payoffs based on the combination of the corresponding combination of $B$ & $C$. Then, we adjust the physicians' payoffs in the {Reg} treatment using the same ratio as the patients' payoffs (patients:physicians = 1.5:1). Next, we adjust the physicians' payoffs in the {Flat} and {Prog} treatments, ensuring that the average payoffs across the three incentive schemes remain constant. Detailed computations are reported in Table (ref) in the Appendix.} \begin{table}[h!] \caption{Welfare analysis and Physicians' choices} \begin{threeparttable} {3mm}{ \begin{tabular}{ccccccc} \toprule & \multicolumn{3}{c}{PreAI} & \multicolumn{3}{c}{PostAI} \\ \cmidrule(r){2-4} \cmidrule(r){5-7} & {Flat} & {Prog} & \textit{{Reg}} & \textit{{Flat}} & \textit{{Prog}} & \textit{{Reg}} \\ \midrule \multicolumn{7}{l}{\textbf{Panel A: Welfare Analysis}} \\ \midrule \textit{Patients' payoffs} \\ B=4 C=2 & 4.89 & 4.87 & \textbf{5.45} & 6.32 & 6.09 & \textbf{6.46} \\ B=2 C=4 & 2.41 & 1.71 & \textbf{2.93} & 4.00 & 3.24 & \textbf{4.26} \\ B=4 C=4 & 3.53 & 3.06 & \textbf{4.25} & 5.54 & 4.89 & \textbf{5.81} \\ B=4 C=0 & 6.25 & \textbf{6.69} & 6.65 & 7.09 & \textbf{7.30} & 7.10 \\ \textit{Physicians' payoff} \\ B=4 C=2 & 3.00 & 3.16 & \textbf{4.21} & 3.00 & 2.86 & \textbf{4.86} \\ B=2 C=4 & 0.75 & 0.79 & \textbf{2.44} & 0.75 & 0.71 & \textbf{3.24} \\ B=4 C=4 & 1.50 & 1.58 & \textbf{3.43} & 1.50 & 1.43 & \textbf{4.40} \\ B=4 C=0 & 4.50 & 4.74 & \textbf{4.98} & 4.50 & 4.29 & \textbf{5.33} \\ \textit{Social welfare} \\ B=4 C=2 & 7.89 & 8.03 & \textbf{9.66} & 9.32 & 8.95 & \textbf{11.32} \\ B=2 C=4 & 3.16 & 2.50 & \textbf{5.37} & 4.75 & 3.95 & \textbf{7.50} \\ B=4 C=4 & 5.03 & 4.64 & \textbf{7.68} & 7.04 & 6.32 & \textbf{10.21} \\ B=4 C=0 & 10.75 & 11.43 & \textbf{11.63} & 11.59 & 11.59 & \textbf{12.43} \\ \midrule \multicolumn{7}{l}{\textbf{Panel B: Choice Analysis}} \\ \midrule \textit{Choice type} \\ \multicolumn{1}{l}{Correct (\ding{51})} & 41.1% & 24.3% & 48.3% & 64.3% & 45.3% & 69.1% \\ \multicolumn{1}{l}{Correct/Incorrect (\ding{51}\ding{55})} & 15.1% & 43.0% & 17.9% & 13.0% & 37.3% & 8.4% \\ \multicolumn{1}{l}{Incorrect (\ding{55})} & 34.6% & 17.6% & 25.8% & 19.8% & 11.8% & 21.1% \\ \multicolumn{1}{l}{Double Incorrect (\ding{55}\ding{55})} & 9.1% & 15.1% & 8.1% & 3.0% & 0.5% & 1.4% \\ \bottomrule \end{tabular} } \begin{tablenotes} • { \textit{Notes:} This table reports (a) Welfare analysis centred around physicians' payoff, patients' benefits, and societal welfare, and (b) Physicians' choice statistics with respect to incentive scheme treatments (\textit{{Flat}}, \textit{{Prog}}, \textit{{Reg}}) before and after AI advice.} \end{tablenotes} \end{threeparttable} \end{table} \begin{enumerate} • \textbf{Scenario 1 ($B=4, C=2$): }In this scenario, the benefit of correct treatment is high, and the cost of incorrect treatment is moderate. The \textit{{Reg}} incentive scheme, characterized by decreasing rewards for each additional incorrect diagnosis, strikes an optimal balance by encouraging optimal treatment (\ding{51}) rather than excessive conservative adjustments (\ding{51}\ding{55}). Despite a higher accuracy rate in \textit{{Prog}} treatment due to an increase in the precision-compromise strategy, patients' welfare is not significantly improved in \textit{{Prog}} due to the moderate cost of overtreatment. • \textbf{Scenario 2 ($B=2, C=4$): }Here, the cost of incorrect treatment is twice as high as the benefit of correct treatment, placing a more significant penalty on overtreatment. In this case, the advantageous position of \textit{{Reg}} compared to \textit{{Prog}} becomes more pronounced in terms of patients' payoffs. As the cost of incorrect treatment increases, the benefit of \textit{{Prog}} treatment encouraging optimal treatment adjustments becomes more appealing. • \textbf{Scenario 3 ($B=4, C=4$): }In this scenario, the benefit of correct treatment equals the cost of incorrect treatment. Similar to the previous scenarios, the \textit{{Prog}} scheme maximizes both physicians' and patients' payoffs. The differences are augmented after AI assistance. • \textbf{Scenario 4 ($B=4, C=0$): }Here, we assume there is no cost for incorrect medicine, making overtreatment harmless. The \textit{{Prog}} schemes encouraging more medical prescriptions appeared to be optimal for the patients' welfare. However, physician's payoffs are still maximized under the \textit{{Reg}} scheme due to the largest proportion of optimal treatment (\ding{51}), and the total social welfare is still maximized in the \textit{{Reg}} treatment. \end{enumerate} Taken together, the \textit{{Reg}} treatment maximizes patients' payoffs (except when $C=0$), physicians' payoffs, and social welfare, particularly when augmented by AI assistance. The robust performance of the \textit{{Reg}} scheme across different scenarios of $B$ and $C$ values suggests that it may be a preferred approach in various contexts, offering a balance that optimizes welfare while accounting for the complexities of medical decision-making.

Further analysis: AI adoption and individual heterogeneity

result[Determinants of AI-adoption] Participants with higher scores in algorithm trust are more likely to adopt AI advice. Conversely, participants with higher confidence, higher ability, or those more senior in medical school are less likely to follow AI advice, ceteris paribus.
proof[Support.] Support for this result is provided by Table (ref).Previous results demonstrated the substantial effect of AI on reducing overuse and improving accuracy. Here, we delve into the behavioral question of who is more likely to adopt AI advice in their decision-making. We explore treatment differences and individual heterogeneities in AI adoption. Table (ref) shows that AI adoption is marginally higher in the {Prog} treatment using the main measures due to higher levels of affirmations (e.g., initial choice-AI advice-second choice follows A-A-A). However, the treatment differences disappear when we use the adjusted measure of AI adoption, which omits the affirmation of the adoption of AI. \begin{table}[h!] \def\sym#1{\ifmmode^{#1}\else\(^{#1}\)\fi} \caption{Determinants of AI adoption: Random effect models} \begin{threeparttable} {1mm}{ \begin{tabular}{l*{6}{c}} \toprule &\multicolumn{1}{c}{(1)}&\multicolumn{1}{c}{(2)}&\multicolumn{1}{c}{(3)}&\multicolumn{1}{c}{(4)}&\multicolumn{1}{c}{(5)}&\multicolumn{1}{c}{(6)}\\ Dep Var: AI adoption &\multicolumn{1}{c}{Main}&\multicolumn{1}{c}{Main}&\multicolumn{1}{c}{Main}&\multicolumn{1}{c}{Adjust}&\multicolumn{1}{c}{Adjust}&\multicolumn{1}{c}{Adjust}\\ \midrule {Prog} & 0.024* & 0.026** & 0.032***& -0.012 & 0.029 & 0.046 \\ (Baseline: {Flat}) & (0.012) & (0.012) & (0.011) & (0.052) & (0.043) & (0.039) \\ [1em] {Reg} & 0.015 & 0.016 & 0.012 & -0.008 & 0.017 & -0.019 \\ & (0.012) & (0.012) & (0.014) & (0.044) & (0.037) & (0.051) \\ [1em] Ability & & -0.034 & -0.031 & & -0.643***& -0.469***\\ & & (0.024) & (0.032) & & (0.071) & (0.094) \\ [1em] Belief & & & -0.036 & & & -0.229** \\ & & & (0.035) & & & (0.109) \\ [1em] Surgery & & & -0.036** & & & -0.083* \\ (Baseline: Internal Medicine) & & & (0.015) & & & (0.044) \\ [1em] Grade & & & -0.005 & & & -0.062***\\ & & &(0.005) & & & (0.017) \\ [1em] AI trust & & & 0.005 & & & 0.086** \\ & & & (0.011) & & & (0.035) \\ [1em] \multicolumn{3}{l}{Control for other individual characteristics} & \checkmark & & & \checkmark \\ \\ Constant & 0.952***& 0.969***& 0.990***& 0.489***& 0.810***& 0.827***\\ & (0.010) & (0.016) & (0.074) & (0.034) & (0.043) & (0.227) \\ \hline Observations & 2400 & 2400 & 2400 & 2400 & 2400 & 2400 \\ \bottomrule \end{tabular} } \begin{tablenotes} • {\textit{Notes. }This table reports the factors that influence AI adoption rate using random effect models. The unit of observation is at the individual-round level, totalling 120 subjects $\times$ 20 rounds of observations. Robust standard errors are clustered at individual levels and are reported in parentheses. Model 3 & 5 control for other individual characteristics, including gender, age, risk attitudes, altruistic preferences, trust, cognitive ability (measured by the CRT & Raven tests, and other algorithm preferences (awareness, literacy and fairness perceptions). } \end{tablenotes} \end{threeparttable} \end{table} We observe notable individual heterogeneity in the tendency to adopt AI advice: \footnote{We also have additional results regarding individual heterogeneity on the effect of AI on medical overuse propensity and accuracy rate, considering space and relevance, these results are reported in the Appendix (Result (ref)).} \begin{itemize} • \textbf{Ability}: Measured by the accuracy rate in incentivized practice questions at the beginning of the experiment, ability is negatively correlated with adjusted measures of AI adoption (Models 5 and 6). This is consistent with existing literature indicating that experts exhibit higher algorithm aversion than laypeople kawaguchi2021will. • \textbf{Belief: }Measured by self-reported belief in the accuracy rate in practice questions, which can be seen as a measure of self-confidence, is also negatively related to AI adoption (Model 6). This aligns with the general finding that overconfidence is associated with greater algorithm aversion burton2020systematic. • \textbf{AI trust: }Measured by self-reported trust in AI in everyday life using post-survey questionnaires, AI trust is positively correlated with AI adoption (Model 6), suggesting that our AI trust measures are good indicators of preferences towards AI advice. • \textbf{Grade: }We found that higher-grade medical students are less likely to adopt AI advice, controlling for ability and self-reported beliefs in accuracy (Model 6). \end{itemize} Most interestingly, we found that students who majored in Surgery, compared with internal medicine, were significantly less likely to change their choices after seeing AI advice in both main and adjusted measures of AI adoption (Model 3 & 5). We believe it is the first evidence of heterogeneity in algorithm appreciation propensity among medical specialities. Although these results should be further tested for their robustness and generalizability, we hypothesize that this may be due to the selection effect of the surgery speciality. Individuals who are more likely to trust their own intuition and practice rather than outsource information may be more inclined to choose a surgical discipline compared with Internal medicine. \footnote{For example, existing studies found that surgical residents tend to exhibit higher levels of assertiveness and decisiveness, traits associated with a lower likelihood of relying on external advice; In contrast, internal medicine residents, who typically engage in more collaborative and deliberative decision-making processes, may be more open to incorporating advice into their clinical judgments shanafelt2010burnout,hussenoeder2021comparing. This evidence suggests that the disparity in AI advice adoption between surgery and internal medicine major students might stem from disparity in cognitive reasoning and judgements. }

Discussions and Conclusion

Our findings underscore that integrating AI significantly curbs medical overtreatment, achieving reductions of up to 62% under aligned incentive structures ({Reg}). This phenomenon termed the Synergistic AI Effect, highlights the critical importance of incentive alignment in maximizing AI's efficacy. In terms of quality, AI intervention not only reduced overtreatment but also enhanced diagnostic accuracy across all incentive schemes. The implementation of AI led to more precise medical prescriptions, particularly evident in the shift towards the “Optimal Treatment Adjustment” strategy in {Reg} treatment, which optimized the quality of care without increasing overtreatment.

The study identifies two pivotal causes of overtreatment: monetary incentives and non-monetary factors such as defensive medicine and knowledge gaps. Our results indicate that around 43% of medical overuse is due to non-monetary incentives—participants exhibit overuse even if the extra choice is not rewarded ({Flat}) or even negatively rewarded ({Reg}). Conversely, monetary incentives account for 57% of total medical overuse in the {Prog} treatment, where overtreatment is financially incentivized.

Based on this dichotomy, we propose two potential interventions to curb medical overuse and improve healthcare quality: 1) Implement a pay-by-performance incentive scheme where overtreatment is penalized, such as the {Reg} incentive schemes in our experiment, to align the interests of physicians and patients; 2) Introduce human-AI collaboration by incorporation AI advice in decision making. From a policy perspective, we show that the effects of these two interventions at ensuring healthcare quality can be comparable within existing incentive structures environment, such as pay-by-visit ({Flat}) and pay-by-service ({Prog}), pointing to a promising avenue for integrating AI into the healthcare system.

Theoretically, our findings can be interpreted through the lens of principal-agent theory, where AI acts as a mechanism to align the interests of healthcare providers (agents) with those of patients and payers (principals) by reducing non-monetary incentives of overtreatment. The effectiveness of AI in reducing overtreatment could be modelled as a function of its ability to provide transparent, unbiased, and evidence-based recommendations, which counteract the misaligned incentives inherent in many healthcare systems.

\noindentConclusion. This study demonstrates how artificial intelligence (AI) can be pivotal in reducing medical overtreatment and enhancing diagnostic precision. Reflecting the quote at the beginning of this paper, Taleb advocates querying a doctor not just on general advice but on personal choices if placed in similar circumstances. Our research supports this through the results observed in the {Reg} treatment, where AI's impact was most profound when physician and patient interests were closely aligned -- seeking optimal exact treatment, thus maximizing welfare.

Drawing from “Power and Prediction”, the introduction of AI in healthcare signifies a shift from conservative protocols towards a system driven by predictive analytics and tailored treatments agrawal2022power. This evolution in healthcare suggests that AI's role can extend beyond mere assistance to being a fundamental part of strategic health management, focusing on prevention and precision. Our study takes the first step in mimicking the results and consequences of AI-physician synergy in a harmless, risk-free environment.

Future research should extend these findings by exploring the long-term impact of AI across various healthcare systems while rigorously evaluating ethical dimensions to ensure that the integration of AI upholds principles of fairness and equity in patient care globally.

{ \onehalfspacing }