Subtlex-pl

subtitle-based word frequency estimates for Polish

Pawel Mandera*, Emmanuel Keuleers, Zofia Wodniecka, Marc Brysbaert

*Corresponding author for this work

Research output: Contribution to journalArticleScientificpeer-review

Abstract

We present SUBTLEX-PL, Polish word frequencies based on movie subtitles. In two lexical decision experiments, we compare the new measures with frequency estimates derived from another Polish text corpus that includes predominantly written materials. We show that the frequencies derived from the two corpora perform best in predicting human performance in a lexical decision task if used in a complementary way. Our results suggest that the two corpora may have unequal potential for explaining human performance for words in different frequency ranges and that corpora based on written materials severely overestimate frequencies for formal words. We discuss some of the implications of these findings for future studies comparing different frequency estimates. In addition to frequencies for word forms, SUBTLEX-PL includes measures of contextual diversity, part-of-speech-specific word frequencies, frequencies of associated lemmas, and word bigrams, providing researchers with necessary tools for conducting psycholinguistic research in Polish. The database is freely available for research purposes and may be downloaded from the authors' university Web site at http://crr.ugent.be/subtlex-pl.

Original languageEnglish
Pages (from-to)471-483
Number of pages13
JournalBehavior Research Methods
Volume47
Issue number2
DOIs
Publication statusPublished - Jun 2015
Externally publishedYes

Keywords

  • Word frequencies
  • Polish language
  • Lexical decision
  • Visual word recognition
  • ENGLISH
  • CHOICE

Cite this

Mandera, Pawel ; Keuleers, Emmanuel ; Wodniecka, Zofia ; Brysbaert, Marc. / Subtlex-pl : subtitle-based word frequency estimates for Polish. In: Behavior Research Methods. 2015 ; Vol. 47, No. 2. pp. 471-483.
@article{cdd52760c1684e88b3c4845ffda108f8,
title = "Subtlex-pl: subtitle-based word frequency estimates for Polish",
abstract = "We present SUBTLEX-PL, Polish word frequencies based on movie subtitles. In two lexical decision experiments, we compare the new measures with frequency estimates derived from another Polish text corpus that includes predominantly written materials. We show that the frequencies derived from the two corpora perform best in predicting human performance in a lexical decision task if used in a complementary way. Our results suggest that the two corpora may have unequal potential for explaining human performance for words in different frequency ranges and that corpora based on written materials severely overestimate frequencies for formal words. We discuss some of the implications of these findings for future studies comparing different frequency estimates. In addition to frequencies for word forms, SUBTLEX-PL includes measures of contextual diversity, part-of-speech-specific word frequencies, frequencies of associated lemmas, and word bigrams, providing researchers with necessary tools for conducting psycholinguistic research in Polish. The database is freely available for research purposes and may be downloaded from the authors' university Web site at http://crr.ugent.be/subtlex-pl.",
keywords = "Word frequencies, Polish language, Lexical decision, Visual word recognition, ENGLISH, CHOICE",
author = "Pawel Mandera and Emmanuel Keuleers and Zofia Wodniecka and Marc Brysbaert",
year = "2015",
month = "6",
doi = "10.3758/s13428-014-0489-4",
language = "English",
volume = "47",
pages = "471--483",
journal = "Behavior Research Methods",
issn = "1554-351X",
publisher = "Springer",
number = "2",

}

Subtlex-pl : subtitle-based word frequency estimates for Polish. / Mandera, Pawel; Keuleers, Emmanuel; Wodniecka, Zofia; Brysbaert, Marc.

In: Behavior Research Methods, Vol. 47, No. 2, 06.2015, p. 471-483.

Research output: Contribution to journalArticleScientificpeer-review

TY - JOUR

T1 - Subtlex-pl

T2 - subtitle-based word frequency estimates for Polish

AU - Mandera, Pawel

AU - Keuleers, Emmanuel

AU - Wodniecka, Zofia

AU - Brysbaert, Marc

PY - 2015/6

Y1 - 2015/6

N2 - We present SUBTLEX-PL, Polish word frequencies based on movie subtitles. In two lexical decision experiments, we compare the new measures with frequency estimates derived from another Polish text corpus that includes predominantly written materials. We show that the frequencies derived from the two corpora perform best in predicting human performance in a lexical decision task if used in a complementary way. Our results suggest that the two corpora may have unequal potential for explaining human performance for words in different frequency ranges and that corpora based on written materials severely overestimate frequencies for formal words. We discuss some of the implications of these findings for future studies comparing different frequency estimates. In addition to frequencies for word forms, SUBTLEX-PL includes measures of contextual diversity, part-of-speech-specific word frequencies, frequencies of associated lemmas, and word bigrams, providing researchers with necessary tools for conducting psycholinguistic research in Polish. The database is freely available for research purposes and may be downloaded from the authors' university Web site at http://crr.ugent.be/subtlex-pl.

AB - We present SUBTLEX-PL, Polish word frequencies based on movie subtitles. In two lexical decision experiments, we compare the new measures with frequency estimates derived from another Polish text corpus that includes predominantly written materials. We show that the frequencies derived from the two corpora perform best in predicting human performance in a lexical decision task if used in a complementary way. Our results suggest that the two corpora may have unequal potential for explaining human performance for words in different frequency ranges and that corpora based on written materials severely overestimate frequencies for formal words. We discuss some of the implications of these findings for future studies comparing different frequency estimates. In addition to frequencies for word forms, SUBTLEX-PL includes measures of contextual diversity, part-of-speech-specific word frequencies, frequencies of associated lemmas, and word bigrams, providing researchers with necessary tools for conducting psycholinguistic research in Polish. The database is freely available for research purposes and may be downloaded from the authors' university Web site at http://crr.ugent.be/subtlex-pl.

KW - Word frequencies

KW - Polish language

KW - Lexical decision

KW - Visual word recognition

KW - ENGLISH

KW - CHOICE

U2 - 10.3758/s13428-014-0489-4

DO - 10.3758/s13428-014-0489-4

M3 - Article

VL - 47

SP - 471

EP - 483

JO - Behavior Research Methods

JF - Behavior Research Methods

SN - 1554-351X

IS - 2

ER -