paper-with-me

Papers

SyGra: A Unified Graph-Based Framework for Scalable Generation, Quality Tagging, and Management of Synthetic Data

2025-08-21 · Bidyapati Pradhan, Surajit Dasgupta, Amit Kumar Saha, Omkar Anustoop, Sriram Puttagunta, Vipul Mittal, Gopal Sarda arxiv

The advancement of large language models (LLMs) is critically dependent on the availability of high-quality datasets for Supervised Fine-Tuning (SFT), alignment tasks like Direct Preference Optimization (DPO), etc. In this work, we present a comprehensive synthetic data generation framework that facilitates scalable, configurable, and high-fidelity generation of synthetic data tailored for these training paradigms. Our approach employs a modular and configuration-based pipeline capable of modeling complex dialogue flows with minimal manual intervention. This framework uses a dual-stage quality tagging mechanism, combining heuristic rules and LLM-based evaluations, to automatically filter and score data extracted from OASST-formatted conversations, ensuring the curation of high-quality dialogue samples. The resulting datasets are structured under a flexible schema supporting both SFT and DPO use cases, enabling seamless integration into diverse training workflows. Together, these innovations offer a robust solution for generating and managing synthetic conversational data at scale, significantly reducing the overhead of data preparation in LLM training pipelines.

📄 PDF Abstract BibTeX arXiv:2508.15432

Code (0)

등록된 구현이 없습니다.

Tasks

Synthetic Data Generation

Similar Papers 제목 키워드 기반

OmniSVG: A Unified Scalable Vector Graphics Generation Model

2025-04-08 · Yiying Yang, Wei Cheng, Sijin Chen, Xianfang Zeng 외

Scalable Vector Graphics (SVG) is an important image format widely adopted in graphic design because of their resolution independence and editability. The study of generating high-quality SVG has continuously drawn atten…

modelVector Graphics

MA3DSG: Multi-Agent 3D Scene Graph Generation for Large-Scale Indoor Environments

2026-02-04 · Yirum Kim, Jaewoo Kim, Ue-Hwan Kim arxiv

Current 3D scene graph generation (3DSGG) approaches heavily rely on a single-agent assumption and small-scale environments, exhibiting limited scalability to real-world scenarios. In this work, we introduce Multi-Agent …

Scene Graph Generation

RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance

2025-10-26 · Jiuniu Wang, Gongjie Zhang, Quanhao Qian, Junlong Gao 외 arxiv

Scalable Vector Graphics (SVGs) are fundamental to digital design and robot control, encoding not only visual structure but also motion paths in interactive drawings. In this work, we introduce RoboSVG, a unified multimo…

Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing

2026-08-03 · Junliang Ye, Kenkun Liu, Guocun Wang, Yang Li 외 hf

Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimoda…

Image Generation3D Generation

Towards Automating the Retrospective Generation of BIM Models: A Unified Framework for 3D Semantic Reconstruction of the Built Environment

2024-06-03 · Ka Lung Cheung, Chi Chung Lee

The adoption of Building Information Modeling (BIM) is beneficial in construction projects. However, it faces challenges due to the lack of a unified and scalable framework for converting 3D model details into BIM. This …

3D Architecture3D Semantic Segmentation3D Surface GenerationExtracting Buildings In Remote Sensing Images+2