paper-with-me

홈 › Papers

A Simple Transformer-Based Model for Ego4D Natural Language Queries Challenge

2022-11-16 · Sicheng Mo, Fangzhou Mu, Yin Li

This report describes Badgers@UW-Madison, our submission to the Ego4D Natural Language Queries (NLQ) Challenge. Our solution inherits the point-based event representation from our prior work on temporal action localization, and develops a Transformer-based model for video grounding. Further, our solution integrates several strong video features including SlowFast, Omnivore and EgoVLP. Without bells and whistles, our submission based on a single model achieves 12.64% Mean R@1 and is ranked 2nd on the public leaderboard. Meanwhile, our method garners 28.45% (18.03%) R@5 at tIoU=0.3 (0.5), surpassing the top-ranked solution by up to 5.5 absolute percentage points.

📄 PDF Abstract BibTeX arXiv:2211.08704

Code (1)

SichengMo/Ego4D_NLQ_Actionformer 공식 구현 pytorch

Tasks

Action LocalizationNatural Language QueriesTemporal Action LocalizationVideo Grounding

Similar Papers 제목 키워드 기반

ReLER@ZJU-Alibaba Submission to the Ego4D Natural Language Queries Challenge 2022

2022-07-01 · Naiyuan Liu, Xiaohan Wang, Xiaobo Li, Yi Yang 외

In this report, we present the ReLER@ZJU-Alibaba submission to the Ego4D Natural Language Queries (NLQ) Challenge in CVPR 2022. Given a video clip and a text query, the goal of this challenge is to locate a temporal mome…

Data AugmentationDiversityNatural Language Queries

Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent

2025-04-07 · Bo Chen, Zhenmei Shi, Zhao Song, Jiahao Zhang

Recent advancements in Transformer-based architectures have led to impressive breakthroughs in natural language processing tasks, with models such as GPT-4, Claude, and Gemini demonstrating human-level reasoning abilitie…

Logical Reasoning

Calibration & Reconstruction: Deep Integrated Language for Referring Image Segmentation

2024-04-12 · Yichen Yan, Xingjian He, Sihan Chen, Jing Liu

Referring image segmentation aims to segment an object referred to by natural language expression from an image. The primary challenge lies in the efficient propagation of fine-grained semantic information from textual f…

DecoderImage SegmentationSemantic Segmentation

Open-Vocabulary DETR with Conditional Matching

2022-03-22 · Yuhang Zang, Wei Li, Kaiyang Zhou, Chen Huang 외

Open-vocabulary object detection, which is concerned with the problem of detecting novel objects guided by natural language, has gained increasing attention from the community. Ideally, we would like to extend an open-vo…

Language Modellingobject-detectionObject DetectionOpen-vocabulary object detection+1

Natural Language Models for Data Visualization Utilizing nvBench Dataset

2023-10-02 · Shuo Wang, Carlos Crespo-Quinones

Translation of natural language into syntactically correct commands for data visualization is an important application of natural language models and could be leveraged to many different tasks. A closely related effort i…

Data VisualizationNatural Language QueriesTranslation