paper-with-me

Papers

Borrowing or Codeswitching? Annotating for Finer-Grained Distinctions in Language Mixing

2022-06-10 · LREC 2022 6 · Elena Alvarez Mellado, Constantine Lignos

We present a new corpus of Twitter data annotated for codeswitching and borrowing between Spanish and English. The corpus contains 9,500 tweets annotated at the token level with codeswitches, borrowings, and named entities. This corpus differs from prior corpora of codeswitching in that we attempt to clearly define and annotate the boundary between codeswitching and borrowing and do not treat common "internet-speak" ('lol', etc.) as codeswitching when used in an otherwise monolingual context. The result is a corpus that enables the study and modeling of Spanish-English borrowing and codeswitching on Twitter in one dataset. We present baseline scores for modeling the labels of this corpus using Transformer-based language models. The annotation itself is released with a CC BY 4.0 license, while the text it applies to is distributed in compliance with the Twitter terms of service.

📄 PDF Abstract BibTeX arXiv:2206.04973

Code (1)

lirondos/borrowing-or-codeswitching 공식 구현

Methods 이 논문이 사용한 방법론

Golden Queue Managers 설명 없음

Similar Papers 제목 키워드 기반

Let's Split Up: Zero-Shot Classifier Edits for Fine-Grained Video Understanding

2026-02-18 · Kaiting Liu, Hazel Doughty arxiv

Video recognition models are typically trained on fixed taxonomies which are often too coarse, collapsing distinctions in object, manner or outcome under a single label. As tasks and definitions evolve, such models canno…

Finer Grained Entity Typing with TypeNet

2017-11-15 · Shikhar Murty, Patrick Verga, Luke Vilnis, Andrew McCallum

We consider the challenging problem of entity typing over an extremely fine grained set of types, wherein a single mention or entity can have many simultaneous and often hierarchically-structured types. Despite the impor…

Entity Typing

CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions

2025-07-08 · Yuchen Huang, Zhiyuan Fan, Zhitao He, Sandeep Polisetty 외

Pretrained vision-language models (VLMs) such as CLIP excel in multimodal understanding but struggle with contextually relevant fine-grained visual features, making it difficult to distinguish visually similar yet cultur…

Contrastive Learning

Codeswitching Detection via Lexical Features in Conditional Random Fields

2016-11-01 · WS 2016 11 · Prajwol Shrestha
Automatic Speech Recognition (ASR)Sentiment AnalysisSpeech Recognition

Exploring Social Bias in Chatbots using Stereotype Knowledge

2019-08-01 · WS 2019 8 · Nayeon Lee, Andrea Madotto, Pascale Fung

Exploring social bias in chatbot is an important, yet relatively unexplored problem. In this paper, we propose an approach to understand social bias in chatbots by leveraging stereotype knowledge. It allows interesting c…

Chatbot