paper-with-me

Papers

VerilogDB: The Largest, Highest-Quality Dataset with a Preprocessing Framework for LLM-based RTL Generation

2025-07-09 · Paul E. Calzada, Zahin Ibnat, Tanvir Rahman, Kamal Kandula, Danyu Lu, Sujan Kumar Saha, Farimah Farahmandi, Mark Tehranipoor arxiv

Large Language Models (LLMs) are gaining popularity for hardware design automation, particularly through Register Transfer Level (RTL) code generation. In this work, we examine the current literature on RTL generation using LLMs and identify key requirements for training and fine-tuning datasets. We construct a robust Verilog dataset through an automated three-pronged process involving database (DB) creation and management with PostgreSQL, data collection from code hosting sites like OpenCores and GitHub, and data preprocessing to verify the codes' syntax, run logic synthesis, and extract relevant module metadata. We implement a scalable and efficient DB infrastructure to support analysis and detail our preprocessing pipeline to enforce high-quality data before DB insertion. The resulting dataset comprises 20,392 Verilog samples, 751 MB of Verilog code data, which is the largest high-quality Verilog dataset for LLM fine-tuning to our knowledge. We further evaluate the dataset, address associated challenges, and explore potential applications for future research and development in LLM-based hardware generation.

📄 PDF Abstract BibTeX arXiv:2507.13369

Code (0)

등록된 구현이 없습니다.

Tasks

Code Generation

Similar Papers 제목 키워드 기반

NEWSFARM: the Largest Chinese Corpus for Long News Summarization

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Recently, driven by a large number of datasets, the field of natural language processing(NLP) has developed rapidly. However, the lack of large-scale and high-quality Chinese datasets is still a critical bottleneck for f…

News SummarizationSemantic SimilaritySemantic Textual SimilarityText Summarization

JESC: Japanese-English Subtitle Corpus

2017-10-29 · LREC 2018 5 · Reid Pryzant, Yongjoo Chung, Dan Jurafsky, Denny Britz

In this paper we describe the Japanese-English Subtitle Corpus (JESC). JESC is a large Japanese-English parallel corpus covering the underrepresented domain of conversational dialogue. It consists of more than 3.2 millio…

Machine TranslationTranslation

Studying Without a Syllabus: Task-Agnostic Environment Preprocessing

2026-09-09 · Vinay Samuel, Varun Ursekar, Vijay S. Kalmath, Apaar Shanker 외 hf

Before an LLM agent tackles tasks in a new environment, it can inspect available corpora and tools and construct reusable resources such as indices, scripts, or procedural guidance. Most automated adaptation methods, how…

From PDF to RAG-Ready: Evaluating Document Conversion Frameworks for Domain-Specific Question Answering

2026-03-30 · José Guilherme Marques dos Santos, Ricardo Yang, Rui Humberto Pereira, Alexandre Sousa 외 arxiv

Retrieval-Augmented Generation (RAG) systems depend critically on the quality of document preprocessing, yet no prior study has evaluated PDF processing frameworks by their impact on downstream question-answering accurac…

Question Answering

Hitting the High Notes: Subset Selection for Maximizing Expected Order Statistics

2020-12-01 · NeurIPS 2020 12 · Aranyak Mehta, Uri Nadav, Alexandros Psomas, Aviad Rubinstein

We consider the fundamental problem of selecting $k$ out of $n$ random variables in a way that the expected highest or second-highest value is maximized. This question captures several applications where we have uncertai…

RetrievalVocal Bursts Intensity Prediction