Using Machine Learning to Understand and Predict the Difficulty of Learning to Read Individual Words from Their Features - PROJECT SUMMARY (ABSTRACT) Underachievement in reading in the United States continues to be a critical public health concern. Literacy skills are essential for being a functioning member of today’s society, from filling out a job application to using a credit card reader at the grocery store; our society presupposes that its members are literate. Individuals with low literacy levels on average experience lower income, lower self-esteem, higher incarceration rates, and many other outcomes that affect their ability to thrive (Bennett et al., 2013; Leone et al., 2005; Morgan et al., 2012; NICHD, 2000; Snow et al., 1998). Given the poor outcomes associated with low literacy skills, there is a critical need to understand how to enhance reading development. Developing robust word reading skills requires early and impactful exposure to printed words, but little is known about the specific word features that impact learning. This research aims to fill that gap by exploring how word features influence their reading difficulty in elementary aged students using a recent NICHD-funded database, the Developmental English Lexicon Project (d-ELP; Compton et al., 2023), which provides difficulty scores for 10,000 words in English derived from word reading data in children grades 1-5. This research will be conducted by creating 12 new word feature variables and exploring the complex relations between these and existing word feature metrics in the prediction of word difficulty. Knowledge of the difficulty of reading specific words can help teachers, curriculum developers, and test-makers make informed decisions about instruction and assessment, and assist researchers in studying reading development, however, the d-ELP data covers only a fraction of the words in English. Collecting additional child data would be costly and time-consuming, therefore, this project will use machine learning to expand the database. The approach will leverage word features as well as four performance predictors, including data from adult naming tasks and base word difficulty for derivatives. By applying and validating machine learning techniques on a hold-out sample of 1,000 words from the larger corpus of 10,000, the project aims to provide difficulty estimates for the remaining words in the English Lexicon Project (ELP; Balota et al., 2007), a similar database of word reading data in adults. This would significantly expand the d-ELP database without the need for additional child testing.