Papers text annotation
“text annotation” 태그가 달린 논문 127편 · 필터 해제
EduCoder: An Open-Source Annotation System for Education Transcript Data
We introduce EduCoder, a domain-specialized tool designed to support utterance-level annotation of educational dialogue. While general-purpose text annotation tools for NLP and qualitative research abound, few address th…
text annotationProbably Approximately Correct Labels
Obtaining high-quality labeled datasets is often costly, requiring either extensive human annotation or expensive experiments. We propose a method that supplements such "expert" labels with AI predictions from pre-traine…
Protein Foldingtext annotationAvatarShield: Visual Reinforcement Learning for Human-Centric Video Forgery Detection
The rapid advancement of Artificial Intelligence Generated Content (AIGC) technologies, particularly in video generation, has led to unprecedented creative capabilities but also increased threats to information integrity…
reinforcement-learningReinforcement Learningtext annotationVideo Forensics+1Comparing LLM Text Annotation Skills: A Study on Human Rights Violations in Social Media Data
In the era of increasingly sophisticated natural language processing (NLP) systems, large language models (LLMs) have demonstrated remarkable potential for diverse applications, including tasks requiring nuanced textual …
Binary Classificationtext annotationIt's a (Blind) Match! Towards Vision-Language Correspondence without Parallel Data
The platonic representation hypothesis suggests that vision and language embeddings become more homogeneous as model and dataset sizes increase. In particular, pairwise distances within each modality become more similar.…
text annotationThe Risks of Using Large Language Models for Text Annotation in Social Science Research
Generative artificial intelligence (GenAI) or large language models (LLMs) have the potential to revolutionize computational social science, particularly in automated textual analysis. In this paper, we conduct a systema…
text annotationMask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation
Sora has unveiled the immense potential of the Diffusion Transformer (DiT) architecture in single-scene video generation. However, the more challenging task of multi-scene video generation, which offers broader applicati…
text annotationVideo GenerationEvaluation of Safety Cognition Capability in Vision-Language Models for Autonomous Driving
Assessing the safety of vision-language models (VLMs) in autonomous driving is particularly important; however, existing work mainly focuses on traditional benchmark evaluations. As interactive components within autonomo…
Autonomous Drivingtext annotationMemory Is All You Need: Testing How Model Memory Affects LLM Performance in Annotation Tasks
Generative Large Language Models (LLMs) have shown promising results in text annotation using zero-shot and few-shot learning. Yet these approaches do not allow the model to retain information from previous annotations, …
AllFew-Shot Learningtext annotationNOTA: Multimodal Music Notation Understanding for Visual Large Language Model
Symbolic music is represented in two distinct forms: two-dimensional, visually intuitive score images, and one-dimensional, standardized text annotation sequences. While large language models have shown extraordinary pot…
cross-modal alignmentLanguage ModelingLanguage ModellingLarge Language Model+1How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators
Human-annotated preference data play an important role in aligning large language models (LLMs). In this paper, we investigate the questions of assessing the performance of human annotators and incentivizing them to prov…
text annotationGPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
In recent years, 2D Vision-Language Models (VLMs) have made significant strides in image-text understanding tasks. However, their performance in 3D spatial comprehension, which is critical for embodied intelligence, rema…
Scene Understandingtext annotationVisual PromptingMask^2DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation
Sora has unveiled the immense potential of the Diffusion Transformer (DiT) architecture in single-scene video generation. However, the more challenging task of multi-scene video generation, which offers broader appli…
text annotationVideo GenerationFrom Human Annotation to LLMs: SILICON Annotation Workflow for Management Research
Unstructured text data annotation and analysis are fundamental to management research, often relying on human annotators through crowdsourcing platforms. While Large Language Models (LLMs) promise to provide a cost-effec…
AttributeManagementModel SelectionNavigate+1Rethinking Emotion Annotations in the Era of Large Language Models
Modern affective computing systems rely heavily on datasets with human-annotated emotion labels, for training and evaluation. However, human annotations are expensive to obtain, sensitive to study design, and difficult t…
Natural Language Understandingtext annotationHolmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity
How can we enable models to comprehend video anomalies occurring over varying temporal scales and contexts? Traditional Video Anomaly Understanding (VAU) methods focus on frame-level anomaly prediction, often missing the…
Anomaly Detectiontext annotationVideo SegmentationVideo Semantic SegmentationGigaHands: A Massive Annotated Dataset of Bimanual Hand Activities
Understanding bimanual human hand activities is a critical problem in AI and robotics. We cannot build large models of bimanual activities because existing datasets lack the scale, coverage of diverse hand activities, an…
DiversityMotion Captioningtext annotationMultimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks
Remote sensing scene classification (RSSC) is a critical task with diverse applications in land use and resource management. While unimodal image-based approaches show promise, they often struggle with limitations such a…
ClassificationScene Classificationtext annotationzero-shot-classification+1Harmful YouTube Video Detection: A Taxonomy of Online Harm and MLLMs as Alternative Annotators
Short video platforms, such as YouTube, Instagram, or TikTok, are used by billions of users globally. These platforms expose users to harmful content, ranging from clickbait or physical harms to misinformation or online …
Binary ClassificationMisinformationtext annotationVERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding
Existing Video Corpus Moment Retrieval (VCMR) is limited to coarse-grained understanding, which hinders precise video moment localization when given fine-grained queries. In this paper, we propose a more challenging fine…
HallucinationMoment RetrievalRetrievaltext annotation+2