paper-with-me

Papers

What is the $\textit{intrinsic}$ dimension of your binary data? -- and how to compute it quickly

2024-04-09 · Tom Hanika, Tobias Hille

Dimensionality is an important aspect for analyzing and understanding (high-dimensional) data. In their 2006 ICDM paper Tatti et al. answered the question for a (interpretable) dimension of binary data tables by introducing a normalized correlation dimension. In the present work we revisit their results and contrast them with a concept based notion of intrinsic dimension (ID) recently introduced for geometric data sets. To do this, we present a novel approximation for this ID that is based on computing concepts only up to a certain support value. We demonstrate and evaluate our approximation using all available datasets from Tatti et al., which have between 469 and 41271 extrinsic dimensions.

📄 PDF Abstract BibTeX arXiv:2404.06326

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Consistence beats causality in recommender systems

2015-01-15 · Zhu Xuzhen, Tian Hui, Hu Zheng, Zhang Ping 외

The explosive growth of information challenges people's capability in finding out items fitting to their own interests. Recommender systems provide an efficient solution by automatically push possibly relevant items to u…

Recommendation Systems

CDH-Bench: A Commonsense-Driven Hallucination Benchmark for Evaluating Visual Fidelity in Vision-Language Models

2026-03-30 · Kesheng Chen, Yamin Hu, Qi Zhou, Zhenqian Zhu 외 arxiv

Vision-language models (VLMs) achieve strong performance on many benchmarks, yet a basic reliability question remains underexplored: when visual evidence conflicts with commonsense, do models follow what is shown or what…

Question Answering

What is the dimension of your binary data?

2019-02-04 · Nikolaj Tatti, Taneli Mielikainen, Aristides Gionis, Heikki Mannila

Many 0/1 datasets have a very large number of variables; on the other hand, they are sparse and the dependency structure of the variables is simpler than the number of variables would suggest. Defining the effective dime…

Clustering

PushupBench: Your VLM is not good at counting pushups

2026-04-25 · Shengzhi Li, Jiarun Chen, Karun Sharma, Jiaqi Su 외 arxiv

Large vision-language models (VLMs) can recognize \textit{what} happens in video but fail to count \textit{how many} times. We introduce \textbf{PushupBench}, 446 long-form clips (avg. 36.7s) for evaluating repetition co…

A Scale-Invariant Diagnostic Approach Towards Understanding Dynamics of Deep Neural Networks

2024-07-12 · Ambarish Moharil, Damian Tamburri, Indika Kumara, Willem-Jan van den Heuvel 외

This paper introduces a scale-invariant methodology employing \textit{Fractal Geometry} to analyze and explain the nonlinear dynamics of complex connectionist systems. By leveraging architectural self-similarity in Deep …

Diagnostic