返回导师列表
LG

Lucas Georges Gabriel Charpentier

Department of Informatics

University of Oslo · Norway

简介

We present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely the largest generally available multilingual collection of LLM pre-training data. These datasets are derived from web crawls from different sources and accompanied with a complete, open-source pipeline for document selection from web archives, text extraction from HTML, language identification for noisy texts, exact and near-deduplication, annotation with, among others, register labels, text quality estimates, and personally identifiable information; and final selection and filtering. We report on data quality probes through contrastive and analytical statistics, through manual inspection of samples for some 20 lan

代表成果

  • Scientific articles and book chapters
  • Oepen, Stephan; Arefyev, Nikolay; Aulamo, Mikko; Bañón, Marta; Buljan, Maja & Burchell, Laurie V. [Show all 29 contributors for this article] (2026). HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models. In Piperidis,, Stelios; Bel, Núria; Heuvel, Henk van den; Ide, Nancy; Krek, Simon & Toral, Antonio (Ed.), Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026). European Language Resources Association. ISSN 9782493814494. p. 1409–1434. doi: 10.63317/25xbdofco9od. Full text in Research Archive Show summary We present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely the largest generally available multilingual collection of LLM pre-training data. These datasets are derived from web crawls from different sources and accompanied with a complete, open-source pipeline for document selection from web archives, text extraction from HTML, language identification for noisy texts, exact and near-deduplication, annotation with, among others, register labels, text quality estimates, and personally identifiable information; and final selection and filtering. We report on data quality probes through contrastive and analytical statistics, through manual inspection of samples for some 20 languages, and through end-to-end evaluation of various language model architectures trained on this data. For multilingual LLM evaluation, we provide a comprehensive collection of benchmarks for nine European languages, with special emphasis on natively created tasks, mechanisms to mitigate prompt sensitivity, and refined normalization and aggregation of scores. Additionally, we train and evaluate a family of 57 monolingual encoder–decoder models, as well as about 30 “smallish” monolingual GPT-like reference models. Besides the monolingual data and models, we also present a very large collection of parallel texts automatically mined from this data, together with a novel parallel corpus synthesized via machine translation.
  • Samuel, David; Mikhailov, Vladislav; Velldal, Erik; Øvrelid, Lilja; Charpentier, Lucas Georges Gabriel & Kutuzov, Andrey [Show all 7 contributors for this article] (2025). Small Languages, Big Models: A Study of Continual Training on Languages of Norway. In Johansson, Richard & Stymne, Sara (Ed.), Proceedings of the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies (NoDaLiDa/Baltic-HLT 2025). University of Tartu Library. ISSN 9789908531090. p. 573–608. doi: https:/aclanthology.org/2025.nodalida-1.61/. Full text in Research Archive Show summary Training large language models requires vast amounts of data, posing a challenge for less widely spoken languages like Norwegian and even more so for truly low-resource languages like Northern S{\'a}mi. To address this issue, we present a novel three-stage continual training approach that substantially improves the downstream performance together with the inference efficiency for the target languages. Based on our findings, we train, evaluate, and openly release a new generative language model for Norwegian Bokm{\r{a}}l, Nynorsk, and Northern S{\'a}mi with 11.4 billion parameters: NorMistral-11B.
  • Wold, Sondre; Charpentier, Lucas Georges Gabriel & Simon, Étienne (2025). Systematic Generalization in Language Models Scales with Information Entropy. In Che, Wanxiang; Nabende, Joyce; Shutova, Ekaterina & Pilehvar, Mohammad (Ed.), Findings of the Association for Computational Linguistics: ACL 2025. Association for Computational Linguistics (ACL). ISSN 9798891762565. p. 1807–1819. doi: 10.18653/v1/2025.findings-acl.90. Full text in Research Archive Show summary Systematic generalization remains challenging for current language models, which are known to be both sensitive to semantically similar permutations of the input and to struggle with known concepts presented in novel contexts. Although benchmarks exist for assessing compositional behavior, it is unclear how to measure the difficulty of a systematic generalization problem. In this work, we show how one aspect of systematic generalization can be described by the entropy of the distribution of component parts in the training data. We formalize a framework for measuring entropy in a sequence-to-sequence task and find that the performance of popular model architectures scales with the entropy. Our work connects systematic generalization to information efficiency, and our results indicate that success at high entropy can be achieved even without built-in priors, and that success at low entropy can serve as a target for assessing progress towards robust systematic generalization.
  • Charpentier, Lucas Georges Gabriel & Lison, Pierre (2025). Re-identification of De-identified Documents with Autoregressive Infilling. In Che, Wanxiang; Nabende, Joyce; Shutova, Ekaterina & Pilehvar, Mohammad (Ed.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics (ACL). ISSN 9798891762510. p. 1192–1209. doi: 10.18653/v1/2025.acl-long.60. Full text in Research Archive Show summary Documents revealing sensitive information about individuals must typically be de-identified. This de-identification is often done by masking all mentions of personally identifiable information (PII), thereby making it more difficult to uncover the identity of the person(s) in question. To investigate the robustness of de-identification methods, we present a novel, RAG-inspired approach that attempts the reverse process of re-identification based on a database of documents representing background knowledge. Given a text in which personal identifiers have been masked, the re-identification proceeds in two steps. A retriever first selects from the background knowledge passages deemed relevant for the re-identification. Those passages are then provided to an infilling model which seeks to infer the original content of each text span. This process is repeated until all masked spans are replaced. We evaluate the re-identification on three datasets (Wikipedia biographies, court rulings and clinical notes). Results show that (1) as many as 80% of de-identified text spans can be successfully recovered and (2) the re-identification accuracy increases along with the level of background knowledge.
  • Wold, Sondre; Simon, Étienne; Charpentier, Lucas Georges Gabriel; Kostylev, Egor V.; Velldal, Erik & Øvrelid, Lilja (2024). Compositional Generalization with Grounded Language Models, Findings of the Association for Computational Linguistics ACL 2024. Association for Computational Linguistics (ACL). ISSN 9798891760998. p. 3447–3460. doi: 10.18653/v1/2024.findings-acl.205. Full text in Research Archive Show summary Grounded language models use external sources of information, such as knowledge graphs, to meet some of the general challenges associated with pre-training. By extending previous work on compositional generalization in semantic parsing, we allow for a controlled evaluation of the degree to which these models learn and generalize from patterns in knowledge graphs. We develop a procedure for generating natural language questions paired with knowledge graphs that targets different aspects of compositionality and further avoids grounding the language models in information already encoded implicitly in their weights. We evaluate existing methods for combining language models with knowledge graphs and find them to struggle with generalization to sequences of unseen lengths and to novel combinations of seen base components. While our experimental results provide some insight into the expressive power of these models, we hope our work and released datasets motivate future research on how to better combine language models with structured knowledge representations.
  • Charpentier, Lucas Georges Gabriel & Samuel, David (2024). GPT or BERT: why not both? In Hu, Michael Y.; Mueller, Aaron; Ross, Candace; Williams, Adina; Linzen, Tal; Zhuang, Chengxu; Choshen, Leshem; Cotterell, Ryan; Warstadt, Alex & Wilcox, Ethan Gotlieb (Ed.), The 2nd BabyLM Challenge at the 28th Conference on Computational Natural Language Learning. Association for Computational Linguistics (ACL). ISSN 9798891762220. p. 262–283. Full text in Research Archive
  • Samuel, David; Charpentier, Lucas Georges Gabriel & Wold, Sondre (2024). More room for language: Investigating the effect of retrieval on language models. In Duh, Kevin; Gomez, Helena & Bethard, Steven (Ed.), Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers). Association for Computational Linguistics (ACL). ISSN 9798891761155. p. 282–305. Full text in Research Archive
  • Charpentier, Lucas Georges Gabriel; Wold, Sondre; Samuel, David & Rønningstad, Egil (2023). BRENT: Bidirectional Retrieval Enhanced Norwegian Transformer. In Alumäe, Tanel & Fishel, Mark (Ed.), Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). University of Tartu. ISSN 9789916219997. Full text in Research Archive
  • Charpentier, Lucas Georges Gabriel & Samuel, David (2023). Not all layers are equally as important: Every Layer Counts BERT. In Warstadt, Alex; Mueller, Aaron; Choshen, Leshem; Wilcox, Ethan; Zhuang, Chengxu; Ciro, Juan; Mosquera, Rafael; Paranjabe, Bhargavi; Williams, Adina; Linzen, Tal & Cotterell, Ryan (Ed.), Proceedings of the BabyLM Challenge at the 27th Conference on Computational Natural Language Learning. Association for Computational Linguistics (ACL). ISSN 9781952148026. p. 238–252. doi: 10.18653/v1/2023.conll-babylm.20. Full text in Research Archive

数据校验于 9/6/2026数据来源

学生评价

还没有评价。成为第一位分享经验的学生吧。