paper-with-me

Papers text annotation

“text annotation” 태그가 달린 논문 127편 · 필터 해제

EduCoder: An Open-Source Annotation System for Education Transcript Data

2025-07-07 · Guanzhong Pan, Mei Tan, HyunJi Nam, Lucía Langlois 외

We introduce EduCoder, a domain-specialized tool designed to support utterance-level annotation of educational dialogue. While general-purpose text annotation tools for NLP and qualitative research abound, few address th…

text annotation

Probably Approximately Correct Labels

2025-06-12 · Emmanuel J. Candès, Andrew Ilyas, Tijana Zrnic

Obtaining high-quality labeled datasets is often costly, requiring either extensive human annotation or expensive experiments. We propose a method that supplements such "expert" labels with AI predictions from pre-traine…

Protein Foldingtext annotation

AvatarShield: Visual Reinforcement Learning for Human-Centric Video Forgery Detection

2025-05-21 · Zhipei Xu, Xuanyu Zhang, Xing Zhou, Jian Zhang

The rapid advancement of Artificial Intelligence Generated Content (AIGC) technologies, particularly in video generation, has led to unprecedented creative capabilities but also increased threats to information integrity…

reinforcement-learningReinforcement Learningtext annotationVideo Forensics+1

Comparing LLM Text Annotation Skills: A Study on Human Rights Violations in Social Media Data

2025-05-15 · Poli Apollinaire Nemkova, Solomon Ubani, Mark V. Albert

In the era of increasingly sophisticated natural language processing (NLP) systems, large language models (LLMs) have demonstrated remarkable potential for diverse applications, including tasks requiring nuanced textual …

Binary Classificationtext annotation

It's a (Blind) Match! Towards Vision-Language Correspondence without Parallel Data

2025-03-31 · CVPR 2025 1 · Dominik Schnaus, Nikita Araslanov, Daniel Cremers

The platonic representation hypothesis suggests that vision and language embeddings become more homogeneous as model and dataset sizes increase. In particular, pairwise distances within each modality become more similar.…

text annotation

The Risks of Using Large Language Models for Text Annotation in Social Science Research

2025-03-27 · Hao Lin, Yongjun Zhang

Generative artificial intelligence (GenAI) or large language models (LLMs) have the potential to revolutionize computational social science, particularly in automated textual analysis. In this paper, we conduct a systema…

text annotation

Mask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation

2025-03-25 · Tianhao Qi, Jianlong Yuan, Wanquan Feng, Shancheng Fang 외

Sora has unveiled the immense potential of the Diffusion Transformer (DiT) architecture in single-scene video generation. However, the more challenging task of multi-scene video generation, which offers broader applicati…

text annotationVideo Generation

Evaluation of Safety Cognition Capability in Vision-Language Models for Autonomous Driving

2025-03-09 · Enming Zhang, Peizhe Gong, Xingyuan Dai, Yisheng Lv 외

Assessing the safety of vision-language models (VLMs) in autonomous driving is particularly important; however, existing work mainly focuses on traditional benchmark evaluations. As interactive components within autonomo…

Autonomous Drivingtext annotation

Memory Is All You Need: Testing How Model Memory Affects LLM Performance in Annotation Tasks

2025-03-06 · Joan C. Timoneda, Sebastián Vallejo Vera

Generative Large Language Models (LLMs) have shown promising results in text annotation using zero-shot and few-shot learning. Yet these approaches do not allow the model to retain information from previous annotations, …

AllFew-Shot Learningtext annotation

NOTA: Multimodal Music Notation Understanding for Visual Large Language Model

2025-02-17 · Mingni Tang, Jiajia Li, Lu Yang, Zhiqiang Zhang 외

Symbolic music is represented in two distinct forms: two-dimensional, visually intuitive score images, and one-dimensional, standardized text annotation sequences. While large language models have shown extraordinary pot…

cross-modal alignmentLanguage ModelingLanguage ModellingLarge Language Model+1

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators

2025-02-10 · Shang Liu, Hanzhao Wang, Zhongyao Ma, Xiaocheng Li

Human-annotated preference data play an important role in aligning large language models (LLMs). In this paper, we investigate the questions of assessing the performance of human annotators and incentivizing them to prov…

text annotation

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

2025-01-02 · Zhangyang Qi, Zhixiong Zhang, Ye Fang, Jiaqi Wang 외

In recent years, 2D Vision-Language Models (VLMs) have made significant strides in image-text understanding tasks. However, their performance in 3D spatial comprehension, which is critical for embodied intelligence, rema…

Scene Understandingtext annotationVisual Prompting

Mask^2DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation

2025-01-01 · CVPR 2025 1 · Tianhao Qi, Jianlong Yuan, Wanquan Feng, Shancheng Fang 외

Sora has unveiled the immense potential of the Diffusion Transformer (DiT) architecture in single-scene video generation. However, the more challenging task of multi-scene video generation, which offers broader appli…

text annotationVideo Generation

From Human Annotation to LLMs: SILICON Annotation Workflow for Management Research

2024-12-19 · Xiang Cheng, Raveesh Mayya, João Sedoc

Unstructured text data annotation and analysis are fundamental to management research, often relying on human annotators through crowdsourcing platforms. While Large Language Models (LLMs) promise to provide a cost-effec…

AttributeManagementModel SelectionNavigate+1

Rethinking Emotion Annotations in the Era of Large Language Models

2024-12-10 · Minxue Niu, Yara El-Tawil, Amrit Romana, Emily Mower Provost

Modern affective computing systems rely heavily on datasets with human-annotated emotion labels, for training and evaluation. However, human annotations are expensive to obtain, sensitive to study design, and difficult t…

Natural Language Understandingtext annotation

Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity

2024-12-09 · CVPR 2025 1 · Huaxin Zhang, Xiaohao Xu, Xiang Wang, Jialong Zuo 외

How can we enable models to comprehend video anomalies occurring over varying temporal scales and contexts? Traditional Video Anomaly Understanding (VAU) methods focus on frame-level anomaly prediction, often missing the…

Anomaly Detectiontext annotationVideo SegmentationVideo Semantic Segmentation

GigaHands: A Massive Annotated Dataset of Bimanual Hand Activities

2024-12-05 · CVPR 2025 1 · Rao Fu, Dingxi Zhang, Alex Jiang, Wanjia Fu 외

Understanding bimanual human hand activities is a critical problem in AI and robotics. We cannot build large models of bimanual activities because existing datasets lack the scale, coverage of diverse hand activities, an…

DiversityMotion Captioningtext annotation

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks

2024-12-03 · Jinjin Cai, Kexin Meng, Baijian Yang, Gang Shao

Remote sensing scene classification (RSSC) is a critical task with diverse applications in land use and resource management. While unimodal image-based approaches show promise, they often struggle with limitations such a…

ClassificationScene Classificationtext annotationzero-shot-classification+1

Harmful YouTube Video Detection: A Taxonomy of Online Harm and MLLMs as Alternative Annotators

2024-11-06 · Claire Wonjeong Jo, Miki Wesołowska, Magdalena Wojcieszak

Short video platforms, such as YouTube, Instagram, or TikTok, are used by billions of users globally. These platforms expose users to harmful content, ranging from clickbait or physical harms to misinformation or online …

Binary ClassificationMisinformationtext annotation

VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding

2024-10-11 · Houlun Chen, Xin Wang, Hong Chen, Zeyang Zhang 외

Existing Video Corpus Moment Retrieval (VCMR) is limited to coarse-grained understanding, which hinders precise video moment localization when given fine-grained queries. In this paper, we propose a more challenging fine…

HallucinationMoment RetrievalRetrievaltext annotation+2
1–20 / 127 다음 →