Papers Data Summarization
“Data Summarization” 태그가 달린 논문 97편 · 필터 해제
PCA-Guided Quantile Sampling: Preserving Data Structure in Large-Scale Subsampling
We introduce Principal Component Analysis guided Quantile Sampling (PCA QS), a novel sampling framework designed to preserve both the statistical and geometric structure of large scale datasets. Unlike conventional PCA, …
Data SummarizationAI Agents vs. Agentic AI: A Conceptual Taxonomy, Applications and Challenges
This study critically distinguishes between AI Agents and Agentic AI, offering a structured conceptual taxonomy, application mapping, and challenge analysis to clarify their divergent design philosophies and capabilities…
AI AgentData SummarizationHallucinationPrompt Engineering+2Dynamic data summarization for hierarchical spatial clustering
Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN) finds meaningful patterns in spatial data by considering density and spatial proximity. As the clustering algorithm is inherently designe…
ClusteringData SummarizationFair Clustering for Data Summarization: Improved Approximation Algorithms and Complexity Insights
Data summarization tasks are often modeled as $k$-clustering problems, where the goal is to choose $k$ data points, called cluster centers, that best represent the dataset by minimizing a clustering objective. A popular …
ClusteringData SummarizationFairnessGet Rid of Isolation: A Continuous Multi-task Spatio-Temporal Learning Framework
Spatiotemporal learning has become a pivotal technique to enable urban intelligence. Traditional spatiotemporal models mostly focus on a specific task by assuming a same distribution between training and testing sets. Ho…
Data SummarizationMulti-Task LearningLinear Submodular Maximization with Bandit Feedback
Submodular optimization with bandit feedback has recently been studied in a variety of contexts. In a number of real-world applications such as diversified recommender systems and data summarization, the submodular funct…
Data SummarizationRecommendation SystemsGIST: Greedy Independent Set Thresholding for Diverse Data Summarization
We introduce a novel subset selection problem called min-distance diversification with monotone submodular utility ($\textsf{MDMS}$), which has a wide variety of applications in machine learning, e.g., data sampling and …
Data SummarizationDiversityfeature selectionimage-classification+1Federated Combinatorial Multi-Agent Multi-Armed Bandits
This paper introduces a federated learning framework tailored for online combinatorial optimization with bandit feedback. In this setting, agents select subsets of arms, observe noisy rewards for these subsets without ac…
Combinatorial OptimizationData SummarizationFederated LearningMulti-Armed BanditsLLMSense: Harnessing LLMs for High-level Reasoning Over Spatiotemporal Sensor Traces
Most studies on machine learning in sensing systems focus on low-level perception tasks that process raw sensory data within a short time window. However, many practical applications, such as human routine modeling and o…
Data SummarizationWorld KnowledgeGreedyML: A Parallel Algorithm for Maximizing Constrained Submodular Functions
We describe a parallel approximation algorithm for maximizing monotone submodular functions subject to hereditary constraints on distributed memory multiprocessors. Our work is motivated by the need to solve submodular o…
Data SummarizationDiffRed: Dimensionality Reduction guided by stable rank
In this work, we propose a novel dimensionality reduction technique, DiffRed, which first projects the data matrix, A, along first $k_1$ principal components and the residual matrix $A^{*}$ (left after subtracting its $k…
Data SummarizationData VisualizationDimensionality ReductionAnalysis of Persian News Agencies on Instagram, A Words Co-occurrence Graph-based Approach
The rise of the Internet and the exponential increase in data have made manual data summarization and analysis a challenging task. Instagram social network is a prominent social network widely utilized in Iran for inform…
Community DetectionData SummarizationKeyword ExtractionDynamic Non-monotone Submodular Maximization
Maximizing submodular functions has been increasingly used in many applications of machine learning, such as data summarization, recommendation systems, and feature selection. Moreover, there has been a growing interest …
Data Summarizationfeature selectionRecommendation SystemsVideo SummarizationDynamic Spatio-Temporal Summarization using Information Based Fusion
In the era of burgeoning data generation, managing and storing large-scale time-varying datasets poses significant challenges. With the rise of supercomputing capabilities, the volume of data produced has soared, intensi…
Data SummarizationDecision MakingRobust Approximation Algorithms for Non-monotone $k$-Submodular Maximization under a Knapsack Constraint
The problem of non-monotone $k$-submodular maximization under a knapsack constraint ($\kSMK$) over the ground set size $n$ has been raised in many applications in machine learning, such as data summarization, information…
Data SummarizationData Summarization beyond Monotonicity: Non-monotone Two-Stage Submodular Maximization
The objective of a two-stage submodular maximization problem is to reduce the ground set using provided training functions that are submodular, with the aim of ensuring that optimizing new objective functions over the re…
Data SummarizationTime-to-Pattern: Information-Theoretic Unsupervised Learning for Scalable Time Series Summarization
Data summarization is the process of generating interpretable and representative subsets from a dataset. Existing time series summarization approaches often search for recurring subsequences using a set of manually devis…
Data SummarizationDiversityTime SeriesOn the Usefulness of Synthetic Tabular Data Generation
Despite recent advances in synthetic data generation, the scientific community still lacks a unified consensus on its usefulness. It is commonly believed that synthetic data can be used for both data exchange and boostin…
Data AugmentationData SummarizationPrivacy PreservingSynthetic Data Generation+1ChartSumm: A Comprehensive Benchmark for Automatic Chart Summarization of Long and Short Summaries
Automatic chart to text summarization is an effective tool for the visually impaired people along with providing precise insights of tabular data in natural language to the user. A large and well-structured dataset is al…
Data SummarizationHallucinationAchieving Long-term Fairness in Submodular Maximization through Randomization
Submodular function optimization has numerous applications in machine learning and data analysis, including data summarization which aims to identify a concise and diverse set of data points from a large dataset. It is i…
Data SummarizationFairness