SeDa: A Unified System for Dataset Discovery and Multi-Entity Augmented Semantic Exploration
The continuous expansion of open data platforms and research repositories has led to a fragmented dataset ecosystem, posing significant challenges for cross-source data discovery and interpretation. To address these challenges, we introduce SeDa--a unified framework for dataset discovery, semantic annotation, and multi-entity augmented navigation. SeDa integrates more than 7.6 million datasets from over 200 platforms, spanning governmental, academic, and industrial domains. The framework first performs semantic extraction and standardization to harmonize heterogeneous metadata representations. On this basis, a topic-tagging mechanism constructs an extensible tag graph that supports thematic retrieval and cross-domain association, while a provenance assurance module embedded within the annotation process continuously validates dataset sources and monitors link availability to ensure reliability and traceability. Furthermore, SeDa employs a multi-entity augmented navigation strategy that organizes datasets within a knowledge space of sites, institutions, and enterprises, enabling contextual and provenance-aware exploration beyond traditional search paradigms. Comparative experiments with popular dataset search platforms, such as ChatPD and Google Dataset Search, demonstrate that SeDa achieves superior coverage, timeliness, and traceability. Taken together, SeDa establishes a foundation for trustworthy, semantically enriched, and globally scalable dataset exploration.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
SEDAR: a Large Scale French-English Financial Domain Parallel Corpus
This paper describes the acquisition, preprocessing and characteristics of SEDAR, a large scale English-French parallel corpus for the financial domain. Our extensive experiments on machine translation show that SEDAR is…
Domain AdaptationMachine TranslationSentenceTranslationSedano: A News Stream Processor for Business
We present Sedano, a system for processing and indexing a continuous stream of business-related news. Sedano defines pipelines whose stages analyze and enrich news items (e.g., newspaper articles and press releases). New…
ArticlesEntity LinkingDOSE-I: A Multimodal Biosignal Dataset of Procedural Sedation for Endoscopy -- Technical Report
In this document, we describe characteristics and technical details of the multimodal biosignal dataset DOSE-I of procedural sedation for endoscopy published on zenodo. The DOSE-I dataset includes 78.5 hours of recording…
Artifact DetectionPhysEDA: Physics-Aware Learning Framework for Efficient EDA With Manhattan Distance Decay
Electronic design automation (EDA) addresses placement, routing, timing analysis, and power-integrity verification for integrated circuits. Learning methods -- attention (Transformer) and reinforcement learning (RL) -- h…
Reinforcement LearningPain Detection in Masked Faces during Procedural Sedation
Pain monitoring is essential to the quality of care for patients undergoing a medical procedure with sedation. An automated mechanism for detecting pain could improve sedation dose titration. Previous studies on facial p…
Medical Procedure