MANTA: A Scalable Pipeline for Transmuting Massive Web Corpora into Instruction Datasets

EMNLP Conference (2025)

Heuiyeen Yeen, Seokee Hong, Hyeongu Yun, Jinsik Lee

Abstract

We introduce MANTA, an automated pipeline that generates high-quality large-scale instruction fine-tuning datasets from massive web corpora while preserving their diversity and scalability. By extracting structured syllabi from web documents and leveraging highperformance LLMs, our approach enables highly effective query-response generation with minimal human intervention. Extensive experiments on 8B-scale LLMs demonstrate that fine-tuning on the MANTA-1M dataset significantly outperforms other massive dataset generation methodologies, particularly in knowledgeintensive tasks such as MMLU and MMLUPro. Our approach also delivers superior performance across a broad spectrum of other tasks, such as Math and Coding. Moreover, MANTA supports seamless scalability by allowing the continuous integration of web corpus data, enabling expansion into domains requiring intensive knowledge.