Sexism in Focus: An Annotated Dataset of YouTube Comments for Gender Bias Research

Thales Bertaglia*, Katarina Bartekova, Rinske Jongma, Stephen Mccarthy, Adriana Iamnitchi

*Corresponding author for this work

Research output: Chapter in Book/Report/Conference proceedingConference article in proceedingAcademicpeer-review

Abstract

This paper presents a novel dataset of 200k YouTube comments from 468 videos across 109 channels in four content categories: Entertainment, Gaming, People & Blogs, and Science & Technology. We applied state-of-the-art NLP methods to augment the dataset with sexism-related features such as sentiment, toxicity, offensiveness, and hate speech. These features can assist manual content analyses and enable automated analysis of sexism in online platforms. Furthermore, we develop an annotation framework inspired by the Ambivalent Sexism Theory to promote a nuanced understanding of how comments relate to the gender of content creators. We release a small sample of comments annotated using this framework. Our dataset analysis confirms that female content creators receive more sexist and hateful comments than their male counterparts, underscoring the need for further research and intervention in addressing online sexism.
Original languageEnglish
Title of host publicationOASIS '23: Proceedings of the 3rd International Workshop on Open Challenges in Online Social Networks
Pages22-28
Number of pages7
ISBN (Electronic)979-8-4007-0225-9
DOIs
Publication statusPublished - 4 Sept 2023

Fingerprint

Dive into the research topics of 'Sexism in Focus: An Annotated Dataset of YouTube Comments for Gender Bias Research'. Together they form a unique fingerprint.

Cite this