The Indo-European Cognate Relationships dataset

Anderson, C., Scarborough, M., Jocz, L., Kümmel, M.J., Jügel, T., Irslinger, B., Pooth, R., Liljegren, H., Strand, R.F., Haig, G., Geupel, U., Macak, M., Kim, R.I., Anonby, E., Pronk, T., Belyaev, O., Dewey-Findell, T.K., Boutilier, M., Freiberg, C., Tegethoff, R., Serangeli, M., Stroński,, Falileyev, A., Liosis, N., Schulte, K., Gupta, G.K., Izadifar, R., Markus, P., Williams, N., Loi, S., Sims-Williams, N., Findell, M., Adibifar, S., Abete, G., Atanasov, P., Baiwir, E., Bastardas, M-R., Adam Benkato, A., Bevevino, L.S., Buchi, V., Cadorini, G., Cathcart, C., Cheveau, L., Christodoulou, C., Delorme, J., Dworkin, S.N., Ekici, D., Farridnejad, S., Gheitasi, M., Hammarström, H., Hewitt, S., Khan, A.A., Khan, M.K., Khokhlova, L., Kim, D., Lewin, C., Lushaj, B., Mahmoudveysi, P., Mahommadirad, M., Mersch, S., Mustafa, B., Nemati, F., Nourzaei, M., Ó Muircheartaigh, P., Oogjen, V., Ourang, M., Pagan, H., Palmer, T.S., Pepper, S., Purandare, M., Rehman, K., Rhys, G., Røyneland, U., Sagar, M.Z., Sandstedt, J.J., Steensland, L., Taheri-Ardali, M., Talebi-Dastena, M., Tittel, S., Tresoldi, T., de Vaan, M., Verkerk, A., Versloot, A., Videsott, P., Vuletić, N., Widmer, M., Zeini, A., Bibiko, H-J., Runge, F., Gray, R.D. and Heggarty, P. 2025. The Indo-European Cognate Relationships dataset. Scientific Data. 12 1541. https://doi.org/10.1038/s41597-025-05445-3

TitleThe Indo-European Cognate Relationships dataset
TypeJournal article
AuthorsAnderson, C., Scarborough, M., Jocz, L., Kümmel, M.J., Jügel, T., Irslinger, B., Pooth, R., Liljegren, H., Strand, R.F., Haig, G., Geupel, U., Macak, M., Kim, R.I., Anonby, E., Pronk, T., Belyaev, O., Dewey-Findell, T.K., Boutilier, M., Freiberg, C., Tegethoff, R., Serangeli, M., Stroński,, Falileyev, A., Liosis, N., Schulte, K., Gupta, G.K., Izadifar, R., Markus, P., Williams, N., Loi, S., Sims-Williams, N., Findell, M., Adibifar, S., Abete, G., Atanasov, P., Baiwir, E., Bastardas, M-R., Adam Benkato, A., Bevevino, L.S., Buchi, V., Cadorini, G., Cathcart, C., Cheveau, L., Christodoulou, C., Delorme, J., Dworkin, S.N., Ekici, D., Farridnejad, S., Gheitasi, M., Hammarström, H., Hewitt, S., Khan, A.A., Khan, M.K., Khokhlova, L., Kim, D., Lewin, C., Lushaj, B., Mahmoudveysi, P., Mahommadirad, M., Mersch, S., Mustafa, B., Nemati, F., Nourzaei, M., Ó Muircheartaigh, P., Oogjen, V., Ourang, M., Pagan, H., Palmer, T.S., Pepper, S., Purandare, M., Rehman, K., Rhys, G., Røyneland, U., Sagar, M.Z., Sandstedt, J.J., Steensland, L., Taheri-Ardali, M., Talebi-Dastena, M., Tittel, S., Tresoldi, T., de Vaan, M., Verkerk, A., Versloot, A., Videsott, P., Vuletić, N., Widmer, M., Zeini, A., Bibiko, H-J., Runge, F., Gray, R.D. and Heggarty, P.
Abstract

The Indo-European Cognate Relationships (IE-CoR) dataset is an open-access relational dataset showing how related, inherited words (‘cognates’) pattern across 160 languages of the Indo-European family. IE-CoR is intended as a benchmark dataset for computational research into the evolution of the Indo-European languages. It is structured around 170 reference meanings in core lexicon, and contains 25731 lexeme entries, analysed into 4981 cognate sets. Novel, dedicated structures are used to code all known cases of horizontal transfer. All 13 main documented clades of Indo-European, and their main subclades, are well represented. Time calibration data for each language are also included, as are relevant geographical and social metadata. Data collection was performed by an expert consortium of 89 linguists drawing on 355 cited sources. The dataset is extendable to further languages and meanings and follows the Cross-Linguistic Data Format (CLDF) protocols for linguistic data. It is designed to be interoperable with other cross-linguistic datasets and catalogues, and provides a reference framework for similar initiatives for other language families.

Article number1541
JournalScientific Data
Journal citation12
ISSN2052-4463
Year2025
PublisherNature Research
Publisher's version
License
CC BY 4.0
File Access Level
Open (open metadata and files)
Digital Object Identifier (DOI)https://doi.org/10.1038/s41597-025-05445-3
Publication dates
Published02 Sep 2025

Related outputs

The Anglo-Norman Bible's Books of Kings, A Critical Edition
Pagan, H. Pitts, B.A. (ed.) 2026. The Anglo-Norman Bible's Books of Kings, A Critical Edition. Brepols.

Multilingual glossing and translanguaging in John of Garland’s Dictionarius: The case of Bruges, Public Library, MS 536
Wallis, C., Seiler, A. and Pagan, H. 2024. Multilingual glossing and translanguaging in John of Garland’s Dictionarius: The case of Bruges, Public Library, MS 536. Lexis: Journal of English Lexicology. 3. https://doi.org/10.4000/12izg

Review: Les Proverbes del vilain (MS Oxford, Bodleian Library, Digby 86)
Pagan, H. 2024. Review: Les Proverbes del vilain (MS Oxford, Bodleian Library, Digby 86) . French Studies. 78 (2), p. 322. https://doi.org/10.1093/fs/knad253

Linguistic Layers in John of Garland's Dictionarius
Pagan, H., Seiler, A. and Wallis, C. 2023. Linguistic Layers in John of Garland's Dictionarius. Études Médiévales Anglaises: A French Journal of English Medieval Studies. 102, pp. 63-108.

Contact-Induced Lexical Effects in Medieval English
Dance, R., Durkin, P., Hough, C. and Pagan, H. 2023. Contact-Induced Lexical Effects in Medieval English. in: Sylvester, L.M. and Pons-Sanz, S.M. (ed.) Medieval English in a Multilingual Context: Current Methodologies and Approaches Palgrave Macmillan. pp. 95-121

Anglo-Norman Glossaries
Pagan, H. 2023. Anglo-Norman Glossaries. in: Seiler, A., Benati, C. and Pons-Sanz, S.M. (ed.) Medieval Glossaries from North-Western Europe: Tradition and Innovation Turnhout, Belgium Brepols. pp. 333-341

Multilingual Annotations in Ælfric’s Glossary in London, British Library, Cotton Faustina A. x: A commented edition
Pagan, H. and Seiler, A 2019. Multilingual Annotations in Ælfric’s Glossary in London, British Library, Cotton Faustina A. x: A commented edition. Early Middle English. 1 (2), pp. 13-64.

The Anglo-Norman Prose Chronicle of Early British Kings or the Abbreviated Prose Brut: Text and Translation
Pagan, H. and De Wilde, Geert 2016. The Anglo-Norman Prose Chronicle of Early British Kings or the Abbreviated Prose Brut: Text and Translation. in: Afanasyev, I., Dresvina, J. and Kooper, E. (ed.) The Medieval Chronicle X Brill. pp. 225-319

L’édition de texte et l’Anglo-Norman Dictionary
Pagan, H. and De Wilde, G. 2016. L’édition de texte et l’Anglo-Norman Dictionary. in: Dorr, S. and Greub, Y. (ed.) Quelle philologie pour quelle lexicographie?: Actes de la section 17 du XXVIIème Congrès International de Linguistique et de Philologie Romanes Heidelberg Universitatsverlag Winter. pp. 107-116

Trevet’s Les Cronicles: Manuscripts, Owners and Readers
Pagan, H. 2016. Trevet’s Les Cronicles: Manuscripts, Owners and Readers. in: Rajsic, J., Kooper, E. and Hoche, D. (ed.) The Prose Brut and Other Late Medieval Chronicles: Books have their Histories. Essays in Honour of Lister M. Matheson York York Medieval Press. pp. 149-164

When is a Brut no longer a Brut?
Pagan, H. 2015. When is a Brut no longer a Brut? in: Tétrel, H. and Veysseyre, G. (ed.) L’Historia regum Britannie et les ‘Bruts’ en Europe Classiques Garnier.

The Anglo-Norman Prose Brut and the Political Climate under Edward I
Pagan, H. 2011. The Anglo-Norman Prose Brut and the Political Climate under Edward I . in: Bouget, H. (ed.) Itinéraires Et Confins Éditions du CRBC. pp. 91-105

Permalink - https://westminsterresearch.westminster.ac.uk/item/x9449/the-indo-european-cognate-relationships-dataset


Share this

Usage statistics

1 total views
2 total downloads
These values cover views and downloads from WestminsterResearch and are for the period from September 2nd 2018, when this repository was created.