paper-with-me

홈 › Papers

ForgeHLS: A Large-Scale, Open-Source Dataset for High-Level Synthesis

2025-07-04 · Zedong Peng, Zeju Li, Mingzhe Gao, Qiang Xu, Chen Zhang, Jieru Zhao arxiv

High-Level Synthesis (HLS) plays a crucial role in modern hardware design by transforming high-level code into optimized hardware implementations. However, progress in applying machine learning (ML) to HLS optimization has been hindered by a shortage of sufficiently large and diverse datasets. To bridge this gap, we introduce ForgeHLS, a large-scale, open-source dataset explicitly designed for ML-driven HLS research. ForgeHLS comprises over 400k diverse designs generated from 846 kernels covering a broad range of application domains, consuming over 200k CPU hours during dataset construction. Each kernel includes systematically automated pragma insertions (loop unrolling, pipelining, array partitioning), combined with extensive design space exploration using Bayesian optimization. Compared to existing datasets, ForgeHLS significantly enhances scale, diversity, and design coverage. We further define and evaluate representative downstream tasks in Quality of Result (QoR) prediction and automated pragma exploration, clearly demonstrating ForgeHLS utility for developing and improving ML-based HLS optimization methodologies. The dataset and code are public at https://github.com/zedong-peng/ForgeHLS.

📄 PDF Abstract BibTeX arXiv:2507.03255

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DiffHLS: Differential Learning for High-Level Synthesis QoR Prediction with GNNs and LLM Code Embeddings

2026-04-10 · Zedong Peng, Zeju Li, Qiang Xu, Jieru Zhao arxiv

High-Level Synthesis (HLS) compiles C/C++ into RTL, but exploring pragma-driven optimization choices remains expensive because each design point requires time-consuming synthesis. We propose \textbf{\DiffHLS}, a differen…

Graph Neural Network

ScaleEdit-12M: Scaling Open-Source Image Editing Data Generation via Multi-Agent Framework

2026-03-21 · Guanzhou Chen, Erfei Cui, Changyao Tian, Danni Yang 외 arxiv

Instruction-based image editing has emerged as a key capability for unified multimodal models (UMMs), yet constructing large-scale, diverse, and high-quality editing datasets without costly proprietary APIs remains chall…

Image Editing

OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing

2025-12-08 · Haoyang He, Jie Wang, Jiangning Zhang, Zhucun Xue 외 arxiv

The quality and diversity of instruction-based image editing datasets are continuously increasing, yet large-scale, high-quality datasets for instruction-based video editing remain scarce. To address this gap, we introdu…

Image Editing

MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens

2024-06-17 · Anas Awadalla, Le Xue, Oscar Lo, Manli Shu 외

Multimodal interleaved datasets featuring free-form interleaved sequences of images and text are crucial for training frontier large multimodal models (LMMs). Despite the rapid progression of open-source LMMs, there rema…

DAPO: An Open-Source LLM Reinforcement Learning System at Scale

2025-03-18 · Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 외

Inference scaling empowers LLMs with unprecedented reasoning ability, with reinforcement learning as the core technique to elicit complex reasoning. However, key technical details of state-of-the-art reasoning LLMs are c…

reinforcement-learningReinforcement Learning