返回导师列表
NA

Nikolay Arefev

Research Fellow · Department of Informatics

University of Oslo · Norway

简介

We present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely the largest generally available multilingual collection of LLM pre-training data. These datasets are derived from web crawls from different sources and accompanied with a complete, open-source pipeline for document selection from web archives, text extraction from HTML, language identification for noisy texts, exact and near-deduplication, annotation with, among others, register labels, text quality estimates, and personally identifiable information; and final selection and filtering. We report on data quality probes through contrastive and analytical statistics, through manual inspection of samples for some 20 lan

代表成果

  • Scientific articles and book chapters
  • Oepen, Stephan; Arefyev, Nikolay; Aulamo, Mikko; Bañón, Marta; Buljan, Maja & Burchell, Laurie V. [Show all 29 contributors for this article] (2026). HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models. In Piperidis,, Stelios; Bel, Núria; Heuvel, Henk van den; Ide, Nancy; Krek, Simon & Toral, Antonio (Ed.), Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026). European Language Resources Association. ISSN 9782493814494. p. 1409–1434. doi: 10.63317/25xbdofco9od. Full text in Research Archive Show summary We present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely the largest generally available multilingual collection of LLM pre-training data. These datasets are derived from web crawls from different sources and accompanied with a complete, open-source pipeline for document selection from web archives, text extraction from HTML, language identification for noisy texts, exact and near-deduplication, annotation with, among others, register labels, text quality estimates, and personally identifiable information; and final selection and filtering. We report on data quality probes through contrastive and analytical statistics, through manual inspection of samples for some 20 languages, and through end-to-end evaluation of various language model architectures trained on this data. For multilingual LLM evaluation, we provide a comprehensive collection of benchmarks for nine European languages, with special emphasis on natively created tasks, mechanisms to mitigate prompt sensitivity, and refined normalization and aggregation of scores. Additionally, we train and evaluate a family of 57 monolingual encoder–decoder models, as well as about 30 “smallish” monolingual GPT-like reference models. Besides the monolingual data and models, we also present a very large collection of parallel texts automatically mined from this data, together with a novel parallel corpus synthesized via machine translation.
  • Fedorova, Mariia; Arefyev, Nikolay; Buljan, Maja; Helcl, Jindřich; Oepen, Stephan & Rønningstad, Egil [Show all 7 contributors for this article] (2026). OpenLID-v3: Improving the Precision of Closely Related Language Identification – An Experience Report. In Scherrer, Yves; Aepli, Noëmi; Blaschke, Verena; Jauhiainen, Tommi; Ljubešić, Nikola; Nakov, Preslav; Tiedemann, Jörg & Zampieri, Marcos (Ed.), Proceedings of the 13th Workshop on NLP for Similar Languages, Varieties and Dialects. Association for Computational Linguistics (ACL). ISSN 9798891763722. p. 275–292. doi: 10.18653/v1/2026.vardial-1.23. Full text in Research Archive Show summary Language identification (LID) is an essential step in building high-quality multilingual datasets from web data. Existing LID tools (such as OpenLID or GlotLID) often struggle to identify closely related languages and to distinguish valid natural language from noise, which contaminates language-specific subsets, especially for low-resource languages. In this work we extend the OpenLID classifier by adding more training data, merging problematic language variant clusters, and introducing a special label for marking noise. We call this extended system OpenLID-v3 and evaluate it against GlotLID on multiple benchmarks. During the development we focus on three groups of closely related languages (Bosnian, Croatian, and Serbian; Romance varieties of Northern Italy and Southern France; and Scandinavian languages) and contribute new evaluation datasets where existing ones are inadequate. We find that ensemble approaches improve precision but also substantially reduce coverage for low-resource languages.
  • Bian, Jie; Arefyev, Nikolay; Mühlhäuser, Max & Welzl, Michael (2025). Automated Insights Into GitHub Collaboration Dynamics. IEEE Access. 13, p. 85526–85542. doi: 10.1109/access.2025.3566309. Full text in Research Archive Show summary Today, GitHub (Trademark) is the most widely used platform for open-source software development. Large projects may comprise hundreds of distributed collaborators and thousands of GitHub “issues” (structured discussions). It includes basic support for dealing with issues, via “Pull Requests” (PRs)—document changes that can manually be defined to “close” them (i.e., they address and thereby conclude the issue discussion). Unresolved issues can pile up. For example, at the time of writing, the Kubernetes repository has almost 2000 open issues; finding which ones a PR might close is a hard task by itself. We address this by automatically identifying issue-PR relationships using language models (LMs) and leveraging Information Retrieval (IR) techniques. To foster further research, we contribute a carefully curated novel dataset called CodeConvo, reflecting the most influential open-source repositories for code development as well as some technical document repositories. We use this dataset to benchmark several state-of-the-art (SOTA) non-proprietary models that show exceptional performance on the MTEB benchmark, as well as to train and evaluate the performance of Smart Insights into GitHub Issue-PR Relations (SIGIR), our tailored model. The best SIGIR model/data combination yields an average Mean Reciprocal Rank (MRR) above 0.7, around 20% higher than the best baseline performance. Notably, ablation studies revealed that knowledge transfer occurs not only between different programming languages but also between code and technical documents, albeit to a lesser extent. We believe these results are encouraging and can stimulate the practical application of LMs for taming the complexity of very large projects, in GitHub and beyond.
  • Simon, Étienne; Olsen, Helene Bøsei; Villar, Ramón Carreño; Mishra, Rahul; Arefyev, Nikolay & Yilmaz, Mert Can [Show all 8 contributors for this article] (2025). Abstractive Event Analysis of Armed Conflicts: Introducing the UCDP-AEC Dataset. In Wartena, Christian & Heid, Ulrich (Ed.), Proceedings of the 21st Conference on Natural Language Processing (KONVENS 2025): Workshops. HsH Applied Academics. ISSN 9783690180177. p. 104–119. doi: https:/aclanthology.org/2025.konvens-2.8.pdf. Full text in Research Archive Show summary This paper introduces a new dataset of document-level event annotations in the domain of armed conflict. By augmenting the event database from the Uppsala Conflict Data Program (UCDP) with source documents identified in public web archives, we create the UCDP Abstractive Event analysis Corpus (UCDP-AEC). While a large part of research on information extraction is focused on extracting text spans, realworld use cases often require inferring more high-level information that is not necessarily explicitly mentioned in texts. UCDP-AEC differs from traditional event extraction datasets in that the document-level annotations do not correspond to mere text spans of the input, but capture expert-interpreted and often implicit information. With more than 10 000 documents, UCDP-AEC is of comparable size to the largest human-annotated traditional event extraction datasets. We also report preliminary experimental results for various generative approaches, by fine-tuning both decoder models and existing event argument extraction models that require minimal adaptation to our abstractive formulation of the task
  • Burchell, Laurie; Gibert, Ona de; Arefev, Nikolay; Aulamo, Mikko; Bañón, Marta & Chen, Pinzhen [Show all 35 contributors for this article] (2025). An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT). In Che, Wanxiang; Nabende, Joyce; Shutova, Ekaterina & Pilehvar, Mohammad (Ed.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics (ACL). ISSN 9798891762510. p. 17452–17485. doi: 10.18653/v1/2025.acl-long.854. Full text in Research Archive Show summary Training state-of-the-art large language models requires vast amounts of clean and diverse textual data. However, building suitable multilingual datasets remains a challenge. In this work, we present HPLT v2, a collection of high-quality multilingual monolingual and parallel corpora, extending prior work of the HPLT project. The monolingual portion of the data contains 8T tokens covering 193 languages, while the parallel data contains 380M sentence pairs covering 51 languages. We document the entire data pipeline and release the code to reproduce it. We provide extensive analysis of the quality and characteristics of our data. Finally, we evaluate the performance of language models and machine translation systems trained on HPLT v2, demonstrating its value.
  • Fedorova, Mariia; Kutuzov, Andrei; Arefev, Nikolay & Schlechtweg, Dominik (2024). Enriching Word Usage Graphs with Cluster Definitions. In Calzolari, Nicoletta; Kan, Min-Yen; Hoste, Veronique; Lenci, Alessandro; Sakti, Sakriani & Xue, Nianwen (Ed.), Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). European Language Resources Association. ISSN 9782493814104. p. 6189–6198. Full text in Research Archive Show summary We present a dataset of word usage graphs (WUGs), where the existing WUGs for multiple languages are enriched with cluster labels functioning as sense definitions. They are generated from scratch by fine-tuned encoder-decoder language models. The conducted human evaluation has shown that these definitions match the existing clusters in WUGs better than the definitions chosen from WordNet by two baseline systems. At the same time, the method is straightforward to use and easy to extend to new languages. The resulting enriched datasets can be extremely helpful for moving on to explainable semantic change modeling.
  • Gibert, Ona de; Nail, Graeme; Arefev, Nikolay; Bañón, Marta; Linde, Jelmer van der & Ji, Shaoxiong [Show all 13 contributors for this article] (2024). A New Massive Multilingual Dataset for High-Performance Language Technologies. In Calzolari, Nicoletta; Kan, Min-Yen; Hoste, Veronique; Lenci, Alessandro; Sakti, Sakriani & Xue, Nianwen (Ed.), Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). European Language Resources Association. ISSN 9782493814104. p. 1116–1128. Full text in Research Archive Show summary We present the HPLT (High Performance Language Technologies) language resources, a new massive multilingual dataset including both monolingual and bilingual corpora extracted from CommonCrawl and previously unused web crawls from the Internet Archive. We describe our methods for data acquisition, management and processing of large corpora, which rely on open-source software tools and high-performance computing. Our monolingual collection focuses on low- to medium-resourced languages and covers 75 languages and a total of ≈ 5.6 trillion word tokens de-duplicated on the document level. Our English-centric parallel corpus is derived from its monolingual counterpart and covers 18 language pairs and more than 96 million aligned sentence pairs with roughly 1.4 billion English tokens. The HPLT language resources are one of the largest open text corpora ever released, providing a great resource for language modeling and machine translation training. We publicly release the corpora, the software, and the tools used in this work.
  • Bian, Jie; Welzl, Michael; Kutuzov, Andrei & Arefyev, Nikolay (2024). Tell Me Why: Language Models Help Explain the Rationale Behind Internet Protocol Design. In LI, Geoffrey Ye; Liang, Le; Gündüz, Deniz & Antón-Haro, Carles (Ed.), 2024 IEEE International Conference on Machine Learning for Communication and Networking (ICMLCN). IEEE (Institute of Electrical and Electronics Engineers). ISSN 9798350343199. p. 447–453. doi: 10.1109/ICMLCN59089.2024.10624781. Full text in Research Archive Show summary Request for Comments (RFCs) serve as guidebooks for the implementation of Internet protocols or network mechanisms. They reveal how these protocols and mechanisms work, but the underlying reasons for their operation are not always available in RFCs. We present an attempt to automate the discovery of these reasons in mailing list archives via natural language processing methods. Our approach leverages the known relationship between text changes and the discussions leading to these changes in the Working Group (WG) GitHub repositories of the Internet Engineering Task Force (IETF) to obtain labeled training data. We find that our model is able to generalize, and it can indeed discover emails that have led to formulations in RFCs. This is a first step towards facilitating a deeper understanding of these often complex documents, which can be helpful for developers, protocol designers, and educators alike.

数据校验于 9/6/2026数据来源

学生评价

还没有评价。成为第一位分享经验的学生吧。