paper-with-me

홈 › Papers

IGB: Addressing The Gaps In Labeling, Features, Heterogeneity, and Size of Public Graph Datasets for Deep Learning Research

2023-02-27 · Arpandeep Khatua, Vikram Sharma Mailthody, Bhagyashree Taleka, Tengfei Ma, Xiang Song, Wen-mei Hwu

Graph neural networks (GNNs) have shown high potential for a variety of real-world, challenging applications, but one of the major obstacles in GNN research is the lack of large-scale flexible datasets. Most existing public datasets for GNNs are relatively small, which limits the ability of GNNs to generalize to unseen data. The few existing large-scale graph datasets provide very limited labeled data. This makes it difficult to determine if the GNN model's low accuracy for unseen data is inherently due to insufficient training data or if the model failed to generalize. Additionally, datasets used to train GNNs need to offer flexibility to enable a thorough study of the impact of various factors while training GNN models. In this work, we introduce the Illinois Graph Benchmark (IGB), a research dataset tool that the developers can use to train, scrutinize and systematically evaluate GNN models with high fidelity. IGB includes both homogeneous and heterogeneous academic graphs of enormous sizes, with more than 40% of their nodes labeled. Compared to the largest graph datasets publicly available, the IGB provides over 162X more labeled data for deep learning practitioners and developers to create and evaluate models with higher accuracy. The IGB dataset is a collection of academic graphs designed to be flexible, enabling the study of various GNN architectures, embedding generation techniques, and analyzing system performance issues for node classification tasks. IGB is open-sourced, supports DGL and PyG frameworks, and comes with releases of the raw text that we believe foster emerging language models and GNN research projects. An early public version of IGB is available at https://github.com/IllinoisGraphBenchmark/IGB-Datasets.

📄 PDF Abstract BibTeX arXiv:2302.13522

Code (1)

IllinoisGraphBenchmark/IGB-Datasets 공식 구현 pytorch

Tasks

Node Classification

Similar Papers 제목 키워드 기반

One-shot Federated Learning Methods: A Practical Guide

2025-02-13 · Xiang Liu, Zhenheng Tang, Xia Li, Yijun Song 외

One-shot Federated Learning (OFL) is a distributed machine learning paradigm that constrains client-server communication to a single round, addressing privacy and communication overhead issues associated with multiple ro…

Federated Learning

Unsupervised Domain Adaptation for 3D LiDAR Semantic Segmentation Using Contrastive Learning and Multi-Model Pseudo Labeling

2025-07-24 · Abhishek Kaushik, Norbert Haala, Uwe Soergel arxiv

Addressing performance degradation in 3D LiDAR semantic segmentation due to domain shifts (e.g., sensor type, geographical location) is crucial for autonomous systems, yet manual annotation of target data is prohibitive.…

Unsupervised Domain AdaptationLIDAR Semantic SegmentationContrastive Learning

Ensemble of Large Language Models for Curated Labeling and Rating of Free-text Data

2025-01-14 · Jiaxing Qiu, Dongliang Guo, Papini Natalie, Peace Noelle 외

Free-text responses are commonly collected in psychological studies, providing rich qualitative insights that quantitative measures may not capture. Labeling curated topics of research interest in free-text data by multi…

Sensitivity

LLM Cache Bandit Revisited: Addressing Query Heterogeneity for Cost-Effective LLM Inference

2025-09-19 · Hantao Yang, Hong Xie, Defu Lian, Enhong Chen arxiv

This paper revisits the LLM cache bandit problem, with a special focus on addressing the query heterogeneity for cost-effective LLM inference. Previous works often assume uniform query sizes. Heterogeneous query sizes in…

Prompt-Based Rule Discovery and Boosting for Interactive Weakly-Supervised Learning

2022-05-01 · ACL 2022 5 · Rongzhi Zhang, Yue Yu, Pranav Shetty, Le Song 외

Weakly-supervised learning (WSL) has shown promising results in addressing label scarcity on many NLP tasks, but manually designing a comprehensive, high-quality labeling rule set is tedious and difficult. We study inter…

Weakly-supervised Learning