Skip to main content Skip to secondary navigation

AI Safety: Defining and Measuring Potential Harms of Chatbots

Main content start

Individuals are increasingly using AI chatbots in everyday life, whether for work, learning, entertainment, or routine tasks. As these systems become more sophisticated and personalized, questions about the potential risks they pose to users become increasingly urgent, and may not be adequately addressed by existing frameworks for assessing digital harms.

The Stanford Center for Digital Health and Tech Impact Policy Center convened a group of researchers, industry representatives, policymakers, and public health experts for a multidisciplinary workshop focused on three key questions: (1) How should potential harms of AI chatbots be defined and categorized? (2) How can these harms be measured in real-world settings? (3) How can measurement better support mitigation and safety improvement efforts?

This report summarizes the key themes and recommendations emerging from the workshop, including the need for clearer taxonomies of harm, more standardized approaches to measurement, and greater coordination across industry, academia, and regulatory stakeholders. Participants also highlighted the importance of distinguishing between harmful outcomes, the mechanisms through which harms occur, and the broader social and developmental contexts that shape user experiences.

View the report here