paper-with-me

홈 › Papers

DGME-T: Directional Grid Motion Encoding for Transformer-Based Historical Camera Movement Classification

2025-10-17 · Tingyu Lin, Armin Dadras, Florian Kleber, Robert Sablatnig arxiv

Camera movement classification (CMC) models trained on contemporary, high-quality footage often degrade when applied to archival film, where noise, missing frames, and low contrast obscure motion cues. We bridge this gap by assembling a unified benchmark that consolidates two modern corpora into four canonical classes and restructures the HISTORIAN collection into five balanced categories. Building on this benchmark, we introduce DGME-T, a lightweight extension to the Video Swin Transformer that injects directional grid motion encoding, derived from optical flow, via a learnable and normalised late-fusion layer. DGME-T raises the backbone's top-1 accuracy from 81.78% to 86.14% and its macro F1 from 82.08% to 87.81% on modern clips, while still improving the demanding World-War-II footage from 83.43% to 84.62% accuracy and from 81.72% to 82.63% macro F1. A cross-domain study further shows that an intermediate fine-tuning stage on modern data increases historical performance by more than five percentage points. These results demonstrate that structured motion priors and transformer representations are complementary and that even a small, carefully calibrated motion head can substantially enhance robustness in degraded film analysis. Related resources are available at https://github.com/linty5/DGME-T.

📄 PDF Abstract BibTeX arXiv:2510.15725

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BERT-like Pre-training for Symbolic Piano Music Classification Tasks

2021-07-12 · Yi-Hui Chou, I-Chun Chen, Chin-Jui Chang, Joann Ching 외

This article presents a benchmark study of symbolic piano music classification using the masked language modelling approach of the Bidirectional Encoder Representations from Transformers (BERT). Specifically, we consider…

ClassificationEmotion ClassificationLanguage ModellingMelody Extraction+1

Grid4D: 4D Decomposed Hash Encoding for High-Fidelity Dynamic Gaussian Splatting

2024-10-28 · Jiawei Xu, Zexin Fan, Jian Yang, Jin Xie

Recently, Gaussian splatting has received more and more attention in the field of static scene rendering. Due to the low computational overhead and inherent flexibility of explicit representations, plane-based explicit m…

GridPE: Unifying Positional Encoding in Transformers with a Grid Cell-Inspired Framework

2024-06-11 · Boyang Li, Yulin Wu, Nuoxian Huang, Wenjia Zhang

Understanding spatial location and relationships is a fundamental capability for modern artificial intelligence systems. Insights from human spatial cognition provide valuable guidance in this domain. Neuroscientific dis…

Do Emotions Influence Moral Judgment in Large Language Models?

2026-04-21 · Mohammad Saim, Tianyu Jiang arxiv

Large language models have been extensively studied for emotion recognition and moral reasoning as distinct capabilities, yet the extent to which emotions influence moral judgment remains underexplored. In this work, we …

Emotion Recognition

Bidirectional Feature-aligned Motion Transformation for Efficient Dynamic Point Cloud Compression

2025-09-18 · Xuan Deng, Xingtao Wang, Xiandong Meng, Longguang Wang 외 arxiv

Efficient dynamic point cloud compression (DPCC) critically depends on accurate motion estimation and compensation. However, the inherently irregular structure and substantial local variations of point clouds make this t…

Point Clouds