Cargando…

CELER: A 365-Participant Corpus of Eye Movements in L1 and L2 English Reading

We present CELER (Corpus of Eye Movements in L1 and L2 English Reading), a broad coverage eye-tracking corpus for English. CELER comprises over 320,000 words, and eye-tracking data from 365 participants. Sixty-nine participants are L1 (first language) speakers, and 296 are L2 (second language) speak...

Descripción completa

Detalles Bibliográficos
Autores principales:	Berzak, Yevgeni, Nakamura, Chie, Smith, Amelia, Weng, Emily, Katz, Boris, Flynn, Suzanne, Levy, Roger
Formato:	Online Artículo Texto
Lenguaje:	English
Publicado:	MIT Press 2022
Materias:	Research Article
Acceso en línea:	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9692049/ https://www.ncbi.nlm.nih.gov/pubmed/36439073 http://dx.doi.org/10.1162/opmi_a_00054

_version_	1784837173577515008
author	Berzak, Yevgeni Nakamura, Chie Smith, Amelia Weng, Emily Katz, Boris Flynn, Suzanne Levy, Roger
author_facet	Berzak, Yevgeni Nakamura, Chie Smith, Amelia Weng, Emily Katz, Boris Flynn, Suzanne Levy, Roger
author_sort	Berzak, Yevgeni
collection	PubMed
description	We present CELER (Corpus of Eye Movements in L1 and L2 English Reading), a broad coverage eye-tracking corpus for English. CELER comprises over 320,000 words, and eye-tracking data from 365 participants. Sixty-nine participants are L1 (first language) speakers, and 296 are L2 (second language) speakers from a wide range of English proficiency levels and five different native language backgrounds. As such, CELER has an order of magnitude more L2 participants than any currently available eye movements dataset with L2 readers. Each participant in CELER reads 156 newswire sentences from the Wall Street Journal (WSJ), in a new experimental design where half of the sentences are shared across participants and half are unique to each participant. We provide analyses that compare L1 and L2 participants with respect to standard reading time measures, as well as the effects of frequency, surprisal, and word length on reading times. These analyses validate the corpus and demonstrate some of its strengths. We envision CELER to enable new types of research on language processing and acquisition, and to facilitate interactions between psycholinguistics and natural language processing (NLP).
format	Online Article Text
id	pubmed-9692049
institution	National Center for Biotechnology Information
language	English
publishDate	2022
publisher	MIT Press
record_format	MEDLINE/PubMed
spelling	pubmed-96920492022-11-25 CELER: A 365-Participant Corpus of Eye Movements in L1 and L2 English Reading Berzak, Yevgeni Nakamura, Chie Smith, Amelia Weng, Emily Katz, Boris Flynn, Suzanne Levy, Roger Open Mind (Camb) Research Article We present CELER (Corpus of Eye Movements in L1 and L2 English Reading), a broad coverage eye-tracking corpus for English. CELER comprises over 320,000 words, and eye-tracking data from 365 participants. Sixty-nine participants are L1 (first language) speakers, and 296 are L2 (second language) speakers from a wide range of English proficiency levels and five different native language backgrounds. As such, CELER has an order of magnitude more L2 participants than any currently available eye movements dataset with L2 readers. Each participant in CELER reads 156 newswire sentences from the Wall Street Journal (WSJ), in a new experimental design where half of the sentences are shared across participants and half are unique to each participant. We provide analyses that compare L1 and L2 participants with respect to standard reading time measures, as well as the effects of frequency, surprisal, and word length on reading times. These analyses validate the corpus and demonstrate some of its strengths. We envision CELER to enable new types of research on language processing and acquisition, and to facilitate interactions between psycholinguistics and natural language processing (NLP). MIT Press 2022-07-01 /pmc/articles/PMC9692049/ /pubmed/36439073 http://dx.doi.org/10.1162/opmi_a_00054 Text en © 2022 Massachusetts Institute of Technology https://creativecommons.org/licenses/by/4.0/This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International License (https://creativecommons.org/licenses/by/4.0/) , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. For a full description of the license, please visit https://creativecommons.org/licenses/by/4.0/.
spellingShingle	Research Article Berzak, Yevgeni Nakamura, Chie Smith, Amelia Weng, Emily Katz, Boris Flynn, Suzanne Levy, Roger CELER: A 365-Participant Corpus of Eye Movements in L1 and L2 English Reading
title	CELER: A 365-Participant Corpus of Eye Movements in L1 and L2 English Reading
title_full	CELER: A 365-Participant Corpus of Eye Movements in L1 and L2 English Reading
title_fullStr	CELER: A 365-Participant Corpus of Eye Movements in L1 and L2 English Reading
title_full_unstemmed	CELER: A 365-Participant Corpus of Eye Movements in L1 and L2 English Reading
title_short	CELER: A 365-Participant Corpus of Eye Movements in L1 and L2 English Reading
title_sort	celer: a 365-participant corpus of eye movements in l1 and l2 english reading
topic	Research Article
url	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9692049/ https://www.ncbi.nlm.nih.gov/pubmed/36439073 http://dx.doi.org/10.1162/opmi_a_00054
work_keys_str_mv	AT berzakyevgeni celera365participantcorpusofeyemovementsinl1andl2englishreading AT nakamurachie celera365participantcorpusofeyemovementsinl1andl2englishreading AT smithamelia celera365participantcorpusofeyemovementsinl1andl2englishreading AT wengemily celera365participantcorpusofeyemovementsinl1andl2englishreading AT katzboris celera365participantcorpusofeyemovementsinl1andl2englishreading AT flynnsuzanne celera365participantcorpusofeyemovementsinl1andl2englishreading AT levyroger celera365participantcorpusofeyemovementsinl1andl2englishreading

CELER: A 365-Participant Corpus of Eye Movements in L1 and L2 English Reading

Ejemplares similares