Szczegóły grantu

GO-076

2026-07-24

2028-06-30

Corpus-Based and AI-Generated Dictionary Examples: Statistical Modelling of Differences and Learner Preferences

dr Tomasz Michta

Corpus-Based and AI-Generated Dictionary Examples: Statistical Modelling of Differences and Learner Preferences A distinguishing feature of English dictionaries for learners is that most entries include example sentences or phrases copied or adapted from corpora. These examples reinforce meaning by showing how words are used in context, with attention to typical grammar patterns and lexical collocations (Fox 1987; Frankenberg-Garcia 2014; Frankenberg-Garcia, Rees and Lew 2021). Earlier corpus-based lexicographers had to scan concordances line by line to identify suitable examples (Krishnamurthy 1987), while tools such as Word Sketches (Kilgarriff et al. 2014) and GDEX (Kilgarriff et al. 2008) have since made this process more efficient. The rise of large language models changes this situation again: dictionary examples can now be generated quickly, although early evaluations have described such examples as redundant, unimaginative or inauthentic (Lew 2023; Jakubíček & Rundell 2023). This project investigates how corpus-based and AI-generated dictionary examples differ, and how language learners evaluate them. The study is based on an online task completed by over 200 university students from language and non-language degree programmes. For each of 15 monosemous lexical items likely to be unknown, participants were shown definitions from the Reverso English Dictionary, representing AI-generated content, and the Oxford Advanced Learner’s Dictionary, representing corpus-based lexicography. Each definition was followed by one example from the relevant dictionary. Since most entries contained more than one example, the examples shown to each participant were randomly selected from the full set available. Participants rated 15 AI-generated and 15 corpus-based examples on a five-point scale, explained what made them give examples high or low ratings, provided demographic information and completed LexTALE as a standardised measure of vocabulary knowledge (Lemhöfer & Broersma 2012). The computational and statistical part of the project builds on a data-driven appraisal of all 63 examples in the dataset. The examples were coded using a systematic framework developed through comparison of interpretive coding criteria and discussion of divergences between the researchers. This framework makes it possible to compare the two sources in terms of features such as example length, syntactic form, sentence completeness, contextual support for meaning and the lexical environment of the target word. During the statistical modelling stage, these variables will be linked to learner ratings. Because the data include ordered ratings and repeated observations from the same participants, lexical items and examples, the main statistical analysis will use mixed-effects regression models, with ordinal mixed-effects models used for the learner ratings. This stage is computationally demanding, as the models combine ordered-response thresholds with fixed effects and random effects for the main sources of repeated measurement in the data. Fitting and comparing complex random-effects structures is likely to be time-consuming, as the analysis will involve model selection and repeated fitting of computationally demanding ordinal mixed-effects models. The models will estimate the effects of source and example-level features on learner ratings. Participant-level predictors, especially vocabulary knowledge, will be included to test whether learners with different proficiency profiles respond differently to corpus-based and AI-generated examples. As an exploratory extension, model-based clustering will be used to examine learner-response profiles. In particular, participant-level effects from the mixed models can be analysed with Gaussian mixture modelling to identify possible subgroups of learners who differ in their sensitivity to source, sentence completeness or contextual support. The results will be presented at two international conferences and developed into an article for submission to an international peer-reviewed journal. Frankenberg-Garcia, A. (2014). The use of corpus examples for language comprehension and production. ReCALL, 26(2), 128–146. https://doi.org/10.1017/S0958344014000093 Frankenberg-Garcia, A., Rees, G. & Lew, R. (2021). Slipping Through the Cracks in e- Lexicography. International Journal of Lexicography, 34(2), 206–234. https://doi.org/10.1093/ijl/ecaa022 Jakubíček, M. & Rundell, M. (2023). The end of lexicography? Can ChatGPT outperform current tools for post-editing lexicography? In M. Medveď, M. Měchura, I. Kosem, J. Kallas, C. Tiberius & M. Jakubíček (Eds.), Electronic lexicography in the 21st century (eLex 2023): Invisible Lexicography. Proceedings of the eLex 2023 conference (pp. 518–533). Lexical Computing CZ s.r.o. Kilgarriff, A., Baisa, V., Bušta, J., Jakubíček, M., Kovář, V., Michelfeit, J., Rychlý, P. & Suchomel, V. (2014). The Sketch Engine: Ten years on. Lexicography, 1(1), 7–36. https://doi.org/10.1007/s40607-014-0009-9 Kilgarriff, A., Husák, M., McAdam, K., Rundell, M. & Rychlý, P. (2008). GDEX: Automatically Finding Good Dictionary Examples in a Corpus. In E. Bernal & J. DeCesaris (Eds.), Proceedings of the XIII EURALEX International Congress (pp. 425– 432). Institut Universitari de Linguistica Aplicada, Universitat Pompeu Fabra. Krishnamurthy, R. (1987). The process of compilation. In J. Sinclair (Ed.), Looking Up: An Account of the COBUILD Project in Lexical Computing (pp. 62–85). Collins ELT. Lemhöfer, K. & Broersma, M. (2012). Introducing LexTALE: A quick and valid lexical test for advanced learners of English. Behavior Research Methods, 44, 325–343. https://doi.org/10.3758/s13428-011-0146-0 Lew, R. (2023). ChatGPT as a COBUILD lexicographer. Humanities and Social Sciences Communications, 10(704), 1–10. https://doi.org/10.1057/s41599-023-02119-6