Annotation Methodologies for Vision and Language Dataset Creation
Annotated datasets are commonly used in the training and evaluation of tasks involving natural language and vision (image description generation, action recognition and visual question answering). However, many of the existing datasets reflect problems that emerge in the process of data selection and annotation. Here we point out some of the difficulties and problems one confronts when creating and validating annotated vision and language datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionImage DescriptionQuestion AnsweringTemporal Action LocalizationVisual Question AnsweringVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
One-shot to Weakly-Supervised Relation Classification using Language Models
Relation classification aims at detecting a particular relation type between two entities in text, whose methods mostly requires annotated data. Data annotation is either a manual process for supervised learning, or auto…
RelationRelation ClassificationMapping global dynamics of benchmark creation and saturation in artificial intelligence
Benchmarks are crucial to measuring and steering progress in artificial intelligence (AI). However, recent studies raised concerns over the state of AI benchmarking, reporting issues such as benchmark overfitting, benchm…
BenchmarkingGenerative Models for Synthetic Data: Transforming Data Mining in the GenAI Era
Generative models such as Large Language Models, Diffusion Models, and generative adversarial networks have recently revolutionized the creation of synthetic data, offering scalable solutions to data scarcity, privacy, a…
Synthetic Data GenerationTowards Temporal Knowledge-Base Creation for Fine-Grained Opinion Analysis with Language Models
We propose a scalable method for constructing a temporal opinion knowledge base with large language models (LLMs) as automated annotators. Despite the demonstrated utility of time-series opinion analysis of text for down…
Fine-Grained Opinion AnalysisPrompt EngineeringQuestion AnsweringOpinion MiningDetecting annotation noise in automatically labelled data
We introduce a method for error detection in automatically annotated text, aimed at supporting the creation of high-quality language resources at affordable cost. Our method combines an unsupervised generative model with…
Active LearningDomain AdaptationLanguage ModelingLanguage Modelling+1