paper-with-me

홈 › Papers

Contextualizing the Limits of Model & Evaluation Dataset Curation on Semantic Similarity Classification Tasks

2023-11-03 · Daniel Theron

This paper demonstrates how the limitations of pre-trained models and open evaluation datasets factor into assessing the performance of binary semantic similarity classification tasks. As (1) end-user-facing documentation around the curation of these datasets and pre-trained model training regimes is often not easily accessible and (2) given the lower friction and higher demand to quickly deploy such systems in real-world contexts, our study reinforces prior work showing performance disparities across datasets, embedding techniques and distance metrics, while highlighting the importance of understanding how data is collected, curated and analyzed in semantic similarity classification.

📄 PDF Abstract BibTeX arXiv:2311.04927

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationFrictionSemantic SimilaritySemantic Textual Similarity

Similar Papers 제목 키워드 기반

Reproducibility Report: Contextualizing Hate Speech Classifiers with Post-hoc Explanation

2021-05-24 · Kiran Purohit, Owais Iqbal, Ankan Mullick

The presented report evaluates Contextualizing Hate Speech Classifiers with Post-hoc Explanation paper within the scope of ML Reproducibility Challenge 2020. Our work focuses on both aspects constituting the paper: the m…

Can Stories Help LLMs Reason? Curating Information Space Through Narrative

2024-10-25 · Vahid Sadiri Javadi, Johanne R. Trippas, Yash Kumar Lal, Lucie Flek

Narratives are widely recognized as a powerful tool for structuring information and facilitating comprehension of complex ideas in various domains such as science communication. This paper investigates whether incorporat…

Math

ACAV100M: Automatic Curation of Large-Scale Datasets for Audio-Visual Video Representation Learning

2021-01-26 · ICCV 2021 10 · Sangho Lee, Jiwan Chung, Youngjae Yu, Gunhee Kim 외

The natural association between visual observations and their corresponding sound provides powerful self-supervisory signals for learning video representations, which makes the ever-growing amount of online videos an att…

Representation Learning

ICC: Quantifying Image Caption Concreteness for Multimodal Dataset Curation

2024-03-02 · Moran Yanuka, Morris Alper, Hadar Averbuch-Elor, Raja Giryes

Web-scale training on paired text-image data is becoming increasingly central to multimodal learning, but is challenged by the highly noisy nature of datasets in the wild. Standard data filtering approaches succeed in re…

Sentence

Ontology-aligned structuring and reuse of multimodal materials data and workflows towards automatic reproduction

2026-01-18 · Sepideh Baghaee Ravari, Abril Azocar Guzman, Sarath Menon, Stefan Sandfeld 외 arxiv

Reproducibility of computational results remains a challenge in materials science, as simulation workflows and parameters are often reported only in unstructured text and tables. While literature data are valuable for va…