Sexism in Focus: An Annotated Dataset of YouTube Comments for Gender Bias Research

Thales Bertaglia; Katarina Bartekova; Rinske Jongma; Stephen Mccarthy; Adriana Iamnitchi

doi:10.1145/3599696.3612900

Sexism in Focus: An Annotated Dataset of YouTube Comments for Gender Bias Research

Thales Bertaglia, Katarina Bartekova, Rinske Jongma, Stephen Mccarthy, Adriana Iamnitchi

Research output: Chapter in Book/Report/Conference proceeding › Conference article in proceeding › Academic › peer-review

Abstract

This paper presents a novel dataset of 200k YouTube comments from 468 videos across 109 channels in four content categories: Entertainment, Gaming, People & Blogs, and Science & Technology. We applied state-of-the-art NLP methods to augment the dataset with sexism-related features such as sentiment, toxicity, offensiveness, and hate speech. These features can assist manual content analyses and enable automated analysis of sexism in online platforms. Furthermore, we develop an annotation framework inspired by the Ambivalent Sexism Theory to promote a nuanced understanding of how comments relate to the gender of content creators. We release a small sample of comments annotated using this framework. Our dataset analysis confirms that female content creators receive more sexist and hateful comments than their male counterparts, underscoring the need for further research and intervention in addressing online sexism.

Original language	English
Title of host publication	OASIS '23: Proceedings of the 3rd International Workshop on Open Challenges in Online Social Networks
Pages	22-28
Number of pages	7
ISBN (Electronic)	979-8-4007-0225-9
DOIs	https://doi.org/10.1145/3599696.3612900
Publication status	Published - 4 Sept 2023

Access to Document

10.1145/3599696.3612900

https://dblp.org/db/conf/ht/oasis2023.html#BertagliaBJMI23

Cite this

@inproceedings{ee0e741a77154f5390d951d93d928779,

title = "Sexism in Focus: An Annotated Dataset of YouTube Comments for Gender Bias Research",

abstract = "This paper presents a novel dataset of 200k YouTube comments from 468 videos across 109 channels in four content categories: Entertainment, Gaming, People & Blogs, and Science & Technology. We applied state-of-the-art NLP methods to augment the dataset with sexism-related features such as sentiment, toxicity, offensiveness, and hate speech. These features can assist manual content analyses and enable automated analysis of sexism in online platforms. Furthermore, we develop an annotation framework inspired by the Ambivalent Sexism Theory to promote a nuanced understanding of how comments relate to the gender of content creators. We release a small sample of comments annotated using this framework. Our dataset analysis confirms that female content creators receive more sexist and hateful comments than their male counterparts, underscoring the need for further research and intervention in addressing online sexism.",

author = "Thales Bertaglia and Katarina Bartekova and Rinske Jongma and Stephen Mccarthy and Adriana Iamnitchi",

note = "DBLP License: DBLP's bibliographic metadata records provided through http://dblp.org/ are distributed under a Creative Commons CC0 1.0 Universal Public Domain Dedication. Although the bibliographic metadata records are provided consistent with CC0 1.0 Dedication, the content described by the metadata records is not. Content may be subject to copyright, rights of privacy, rights of publicity and other restrictions.",

year = "2023",

month = sep,

day = "4",

doi = "10.1145/3599696.3612900",

language = "English",

pages = "22--28",

booktitle = "OASIS '23: Proceedings of the 3rd International Workshop on Open Challenges in Online Social Networks",

}

Sexism in Focus: An Annotated Dataset of YouTube Comments for Gender Bias Research. / Bertaglia, Thales; Bartekova, Katarina; Jongma, Rinske et al.
OASIS '23: Proceedings of the 3rd International Workshop on Open Challenges in Online Social Networks. 2023. p. 22-28.

Research output: Chapter in Book/Report/Conference proceeding › Conference article in proceeding › Academic › peer-review

TY - GEN

T1 - Sexism in Focus: An Annotated Dataset of YouTube Comments for Gender Bias Research

AU - Bertaglia, Thales

AU - Bartekova, Katarina

AU - Jongma, Rinske

AU - Mccarthy, Stephen

AU - Iamnitchi, Adriana

N1 - DBLP License: DBLP's bibliographic metadata records provided through http://dblp.org/ are distributed under a Creative Commons CC0 1.0 Universal Public Domain Dedication. Although the bibliographic metadata records are provided consistent with CC0 1.0 Dedication, the content described by the metadata records is not. Content may be subject to copyright, rights of privacy, rights of publicity and other restrictions.

PY - 2023/9/4

Y1 - 2023/9/4

N2 - This paper presents a novel dataset of 200k YouTube comments from 468 videos across 109 channels in four content categories: Entertainment, Gaming, People & Blogs, and Science & Technology. We applied state-of-the-art NLP methods to augment the dataset with sexism-related features such as sentiment, toxicity, offensiveness, and hate speech. These features can assist manual content analyses and enable automated analysis of sexism in online platforms. Furthermore, we develop an annotation framework inspired by the Ambivalent Sexism Theory to promote a nuanced understanding of how comments relate to the gender of content creators. We release a small sample of comments annotated using this framework. Our dataset analysis confirms that female content creators receive more sexist and hateful comments than their male counterparts, underscoring the need for further research and intervention in addressing online sexism.

AB - This paper presents a novel dataset of 200k YouTube comments from 468 videos across 109 channels in four content categories: Entertainment, Gaming, People & Blogs, and Science & Technology. We applied state-of-the-art NLP methods to augment the dataset with sexism-related features such as sentiment, toxicity, offensiveness, and hate speech. These features can assist manual content analyses and enable automated analysis of sexism in online platforms. Furthermore, we develop an annotation framework inspired by the Ambivalent Sexism Theory to promote a nuanced understanding of how comments relate to the gender of content creators. We release a small sample of comments annotated using this framework. Our dataset analysis confirms that female content creators receive more sexist and hateful comments than their male counterparts, underscoring the need for further research and intervention in addressing online sexism.

U2 - 10.1145/3599696.3612900

DO - 10.1145/3599696.3612900

M3 - Conference article in proceeding

SP - 22

EP - 28

BT - OASIS '23: Proceedings of the 3rd International Workshop on Open Challenges in Online Social Networks

ER -