paper-with-me

Papers

New Capability to Look Up an ASL Sign from a Video Example

2024-07-18 · Carol Neidle, Augustine Opoku, Carey Ballard, Yang Zhou, Xiaoxiao He, Gregory Dimitriadis, Dimitris Metaxas

Looking up an unknown sign in an ASL dictionary can be difficult. Most ASL dictionaries are organized based on English glosses, despite the fact that (1) there is no convention for assigning English-based glosses to ASL signs; and (2) there is no 1-1 correspondence between ASL signs and English words. Furthermore, what if the user does not know either the meaning of the target sign or its possible English translation(s)? Some ASL dictionaries enable searching through specification of articulatory properties, such as handshapes, locations, movement properties, etc. However, this is a cumbersome process and does not always result in successful lookup. Here we describe a new system, publicly shared on the Web, to enable lookup of a video of an ASL sign (e.g., a webcam recording or a clip from a continuous signing video). The user submits a video for analysis and is presented with the five most likely sign matches, in decreasing order of likelihood, so that the user can confirm the selection and then be taken to our ASLLRP Sign Bank entry for that sign. Furthermore, this video lookup is also integrated into our newest version of SignStream(R) software to facilitate linguistic annotation of ASL video data, enabling the user to directly look up a sign in the video being annotated, and, upon confirmation of the match, to directly enter into the annotation the gloss and features of that sign, greatly increasing the efficiency and consistency of linguistic annotations of ASL video data.

📄 PDF Abstract BibTeX arXiv:2407.13571

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

First Frame Is the Place to Go for Video Content Customization

2025-11-19 · Jingxi Chen, Zongxia Li, Zhichao Liu, Guangyao Shi 외 arxiv

What role does the first frame play in video generation models? Traditionally, it's viewed as the spatial-temporal starting point of a video, merely a seed for subsequent animation. In this work, we reveal a fundamentall…

Video Generation

Looking Fast and Slow: Memory-Guided Mobile Video Object Detection

2019-03-25 · Mason Liu, Menglong Zhu, Marie White, Yinxiao Li 외

Models and examples built with TensorFlow

object-detectionObject DetectionObject RecognitionReal-Time Object Detection+1

PandaGPT: One Model To Instruction-Follow Them All

2023-05-25 · Yixuan Su, Tian Lan, Huayang Li, Jialu Xu 외

We present PandaGPT, an approach to emPower large lANguage moDels with visual and Auditory instruction-following capabilities. Our pilot experiments show that PandaGPT can perform complex tasks such as detailed image des…

AllImage DescriptionInstruction Following

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models

2026-05-30 · Shengjun Zhang, Zhang Zhang, Simin Huang, Zhenyu Tang 외 arxiv

Recent advancements in video-based world models have demonstrated an unprecedented ability to synthesize high-fidelity visual sequences. However, a fundamental gap persists between visually plausible video generation and…

Video GenerationVideo Alignment

Look Once to Hear: Target Speech Hearing with Noisy Examples

2024-05-10 · Bandhav Veluri, Malek Itani, Tuochao Chen, Takuya Yoshioka 외

In crowded settings, the human brain can focus on speech from a target speaker, given prior knowledge of how they sound. We introduce a novel intelligent hearable system that achieves this capability, enabling target spe…

CPUSpeech Extraction