DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
Hao Liang, Xiaochen Ma, Zhou Liu, Zhen Hao Wong, Zhengyang Zhao, Zimo Meng, Runming He, Chengyu Shen, Qifeng Cai, Zhaoyang Han, Meiyi Qiang, Yalin Feng, Tianyi Bai, Zewei Pan, Ziyi Guo, Yizhen Jiang, Jingwen Deng, Qijie You, Peichao Lai, Tianyu Guo, Chi Hsu Tsai, Hengyi Feng, Rui Hu, Wenkai Yu, Junbo Niu, Bohan Zeng, Ruichuan An, Lu Ma, Jihao Huang, Yaowei Zheng, Conghui He, Linpeng Tang, Bin Cui, Weinan E, Wentao Zhang
December, 2025
Abstract
DataFlow is a unified and extensible framework for reliable, reproducible, and scalable LLM data preparation. It provides system-level abstractions, a PyTorch-style pipeline API, nearly 200 reusable operators, six domain-general pipelines, and an agent that translates natural-language specifications into executable data workflows.
Publication
arXiv preprint arXiv:2512.16676

Ph.D. Student in Computer Science and Engineering
I work on data infrastructure and high-performance distributed systems for large-scale LLM data preparation, with a focus on Ray and Apache Spark.