Abstract
We introduce RecombiText Augmentation (RTA), a novel purely statistical NLP method for compositional data augmentation for data-efficient LLM pre-training in low-resource scenarios. RTA identifies lexically and semantically similar sentences within the corpus and generates synthetic sentence pairs from them while preserving underlying patterns from the corpus. We pre-train GPT-2 and RoBERTa language models on a domain-specific, low-resource corpus of 10 million words, with different proportions of augmented data. We compare our RTA-augmented model variants to a baseline model trained on the full original dataset. Zero-shot results show that the language models pre-trained on synthetic data improve in entity tracking, self-paced reading, and morphological generalization benchmarks. In other tasks, the performance is comparable to the baseline model. We demonstrate that it is possible to expand low-resource datasets by two- to four-fold without compromising benchmark performance, solely through statistical processing of the available data.
| Original language | German |
|---|---|
| DOIs | |
| Publication status | Published - 2025 |
| Event | The First BabyLM Workshop at the Conference on Empirical Methods in Natural Language Processing 2025 - Suzhou, China Duration: 8 Nov 2025 → 8 Nov 2025 |
Seminar/Workshop
| Seminar/Workshop | The First BabyLM Workshop at the Conference on Empirical Methods in Natural Language Processing 2025 |
|---|---|
| Country/Territory | China |
| City | Suzhou |
| Period | 8/11/25 → 8/11/25 |
Austrian Fields of Science 2012
- 102033 Data mining
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver