OpenCARE: Open-Source Infrastructure for Evaluating Conversational AI in Dementia Care - PROJECT SUMMARY/ABSTRACT The United States faces a profound public health crisis in Alzheimer's Disease and Related Dementias (ADRD), with a rapidly escalating demand-supply imbalance between rising incidence and declining caregiving workforces. Conversational AI, such as chatbots and virtual assistants, offers a promising, scalable solution. However, the FDA may consider these AIs medical devices if used for diagnosis or management, which would require them to pass technical evaluation tests. Currently, no such technical evaluation standard exists for conversational AI in ADRD care. Since Large Language Models (LLMs) are the backbone of this technology, the first step is to establish LLM benchmarking specifically for ADRD care. This benchmarking process is complex, requiring collaboration among diverse stakeholders like AI developers, clinicians, designers, patients, and care partners. The current discovery-to-delivery process for this work is too slow, as its traditional waterfall design cannot keep up with the fast evolution of AI. To address the above issues, our interdisciplinary team from the University of Notre Dame and Indiana University proposes to develop and diffuse two open-source research infrastructures. The first is a benchmarking dataset for LLMs focused on ADRD care, and the second is a collaboration software platform called AICareNexus, which is designed to support agile, closed-loop communication among stakeholders. The project's first phase (R61, Years 1-2) focuses on building these initial OpenCARE infrastructures. Specific Aim 1 (R61) will establish the LLM benchmarking dataset as a technical checkpoint, delivering a validated ADRD benchmark and scoring protocol. Specific Aim 2 (R61) will build the AICareNexus software, delivering a functional prototype that passes technical tests. After the initial development, the project moves to Phase 2 (R33, Years 3-5) for community- based dissemination and evaluation, guided by the Consolidated Framework for Implementation Research (CFIR). Specific Aim 3 (R33) will evaluate and improve the LLM benchmark through national competitions, resulting in an expanded benchmark, public leaderboards, and improved LLM models. Specific Aim 4 (R33) will evaluate the acceptability, feasibility, and appropriateness of AICareNexus by disseminating the tool for research, resulting in comprehensive user evaluations. This project is planned with detailed milestones and measurements for aim in each phase. Overall impact: The proposed project will provide an agile and shared infrastructure solution to enable more effective, scalable, safe, and personalized conversational AI for ADRD care, aiming to address the bottlenecks in the discovery-to-delivery AI translational cycle.