DESTIGMA: DEtecting STIgma using Generative AI among Multisite pAtients with substance use disorders - PROJECT SUMMARY The United States (US) has experienced a growing substance overdose crisis over the past two decades, resulting in over 105,000 drug overdose deaths in 2023. Despite 48.7 million Americans over age 12 having substance use disorders (SUD) in the past year, only 1 in 4 received treatment. Stigma is a significant barrier to treatment and is more severe for SUDs than other mental health conditions. In healthcare settings, stigma, characterized by implicit bias and negative attitudes from providers, critically impedes treatment engagement. SUD-related provider stigma may lead to a range of negative health outcomes for people with SUD, including reluctance to seek care, not disclosing drug use, delaying care, opting for self-directed discharge, and discontinuing treatment. While some research has examined individual-level impacts of exposure to provider stigma, the variations in experiencing different types of provider stigma due to contextual-level non-clinical health determinants, such as geolocation (e.g., rural vs. urban), income, education, or access to healthcare, and their compounding impact, remain largely unknown. Provider stigma/implicit bias is malleable and can be rectified via interventions, which in turn can be crucial to improving treatment outcomes for SUD patients. Investigations of linguistic bias and stigma in patient notes have identified salient features, such as the expression of doubt in patient testimony via quotations or evidential use, and use of stigmatizing and negative descriptors. Recent studies employing natural language processing (NLP) to detect stigmatizing expressions via lexical matching have limitations in capturing the full spectrum of stigmatizing language and its contextual variations. Advances in NLP, including the emergence of large language models (LLMs), present an opportunity to address this gap by comprehensively detecting stigmatizing language, enabling large-scale analyses of its impact, and mitigating stigmatizing language. Our overarching objective is to employ NLP, machine learning, and LLMs, in a multi-agent AI architecture to systematically detect and de-stigmatize provider language in clinical notes of SUD patients, examining associations between stigmatizing language and patient outcomes across data from four sites (Emory, Cornell, UPenn, and Mt. Sinai). We propose three specific aims: (aim 1) to extend our pilot supervised machine learning approach for detecting stigmatizing language to achieve human-level agreement; (aim 2) to examine the prevalence of stigmatizing language across patient groups, and their associations with patient outcomes; and (aim 3) to develop and validate a real- time co-pilot system, powered by an open-source, instruction-tuned generative LLM, to assist providers in using non-stigmatizing language in SUD patient notes. Successful execution of this project will create the first EHR integration-ready co-pilot system to substantially reduce stigmatizing language in SUD patient notes, with applicability beyond SUD treatment. It will also provide unprecedented insights into the impact of stigmatizing language on treatment outcomes across patient populations, informing targeted future interventions.