Skip to main navigation Skip to search Skip to main content

RecombiText: Compositional Data Augmentation for Enhancing LLM Pre-Training Datasets in Low-Resource Scenarios

Publications: Contribution to conferencePaperPeer Reviewed

Abstract

We introduce RecombiText Augmentation (RTA), a novel purely statistical NLP method for compositional data augmentation for data-efficient LLM pre-training in low-resource scenarios. RTA identifies lexically and semantically similar sentences within the corpus and generates synthetic sentence pairs from them while preserving underlying patterns from the corpus. We pre-train GPT-2 and RoBERTa language models on a domain-specific, low-resource corpus of 10 million words, with different proportions of augmented data. We compare our RTA-augmented model variants to a baseline model trained on the full original dataset. Zero-shot results show that the language models pre-trained on synthetic data improve in entity tracking, self-paced reading, and morphological generalization benchmarks. In other tasks, the performance is comparable to the baseline model. We demonstrate that it is possible to expand low-resource datasets by two- to four-fold without compromising benchmark performance, solely through statistical processing of the available data.
Original languageGerman
DOIs
Publication statusPublished - 2025
EventThe First BabyLM Workshop at the Conference
on Empirical Methods in Natural Language Processing 2025
- Suzhou, China
Duration: 8 Nov 20258 Nov 2025

Seminar/Workshop

Seminar/WorkshopThe First BabyLM Workshop at the Conference
on Empirical Methods in Natural Language Processing 2025
Country/TerritoryChina
CitySuzhou
Period8/11/258/11/25

Austrian Fields of Science 2012

  • 102033 Data mining

Cite this