paper-with-me

홈 › Papers

TVD: A Reproducible and Multiply Aligned TV Series Dataset

2014-05-01 · LREC 2014 5 · Anindya Roy, Camille Guinaudeau, Herv{\'e} Bredin, Claude Barras

We introduce a new dataset built around two TV series from different genres, The Big Bang Theory, a situation comedy and Game of Thrones, a fantasy drama. The dataset has multiple tracks extracted from diverse sources, including dialogue (manual and automatic transcripts, multilingual subtitles), crowd-sourced textual descriptions (brief episode summaries, longer episode outlines) and various metadata (speakers, shots, scenes). The paper describes the dataset and provide tools to reproduce it for research purposes provided one has legally acquired the DVD set of the series. Tools are also provided to temporally align a major subset of dialogue and description tracks, in order to combine complementary information present in these tracks for enhanced accessibility. For alignment, we consider tracks as comparable corpora and first apply an existing algorithm for aligning such corpora based on dynamic time warping and TFIDF-based similarity scores. We improve this baseline algorithm using contextual information, WordNet-based word similarity and scene location information. We report the performance of these algorithms on a manually aligned subset of the data. To highlight the interest of the database, we report a use case involving rich speech retrieval and propose other uses.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Dynamic Time WarpingInformation RetrievalRetrievalSentiment AnalysisWord Similarity

Similar Papers 제목 키워드 기반

An Open Source and Reproducible Implementation of LSTM and GRU Networks for Time Series Forecasting

2022-06-22 · Eng. Proc. 2022 6 · Gissel Velarde, Pedro Brañez, Alejandro Bueno, Rodrigo Heredia 외

This paper introduces an open source and reproducible implementation of Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) networks for time series forecasting. We evaluated LSTM and GRU networks because of the…

Time SeriesTime Series AnalysisTime Series Forecasting

NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention

2024-03-02 · Tianyi Zhang, Jonah Wonkyu Yi, Bowen Yao, Zhaozhuo Xu 외

Large language model inference on Central Processing Units (CPU) is challenging due to the vast quantities of expensive Multiply-Add (MAD) matrix operations in the attention computations. In this paper, we argue that the…

16kCPULanguage ModelingLanguage Modelling+1

A Multi-Modal Dataset for Ground Reaction Force Estimation Using Consumer Wearable Sensors

2026-03-19 · Parvin Ghaffarzadeh, Debarati Chakraborty, Koorosh Aslansefat, Ali Dostan 외 arxiv

This Data Descriptor presents a fully open, multi-modal dataset for estimating vertical ground reaction force (vGRF) from consumer-grade Apple Watch sensors with laboratory force plate ground truth. Ten healthy adults ag…

Matrix Profile for Time-Series Anomaly Detection: A Reproducible Open-Source Benchmark on TSB-AD

2026-04-02 · Chin-Chia Michael Yeh arxiv

Matrix Profile (MP) methods are an interpretable and scalable family of distance-based methods for time-series anomaly detection, but strong benchmark performance still depends on design choices beyond a vanilla nearest-…

Anomaly Detection

GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

2026-09-21 · Yiran Wang, Xingyilang Yin, Junfu Pu, Guangzhi Wang 외 hf

Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, and precise action control over multiple temporal horizons. Existing d…