paper-with-me

홈 › Papers

Language Diversity: Visible to Humans, Exploitable by Machines

2022-03-09 · ACL 2022 5 · Gábor Bella, Erdenebileg Byambadorj, Yamini Chandrashekar, Khuyagbaatar Batsuren, Danish Ashgar Cheema, Fausto Giunchiglia

The Universal Knowledge Core (UKC) is a large multilingual lexical database with a focus on language diversity and covering over a thousand languages. The aim of the database, as well as its tools and data catalogue, is to make the somewhat abstract notion of diversity visually understandable for humans and formally exploitable by machines. The UKC website lets users explore millions of individual words and their meanings, but also phenomena of cross-lingual convergence and divergence, such as shared interlingual meanings, lexicon similarities, cognate clusters, or lexical gaps. The UKC LiveLanguage Catalogue, in turn, provides access to the underlying lexical data in a computer-processable form, ready to be reused in cross-lingual applications.

📄 PDF Abstract BibTeX arXiv:2203.04723

Code (0)

등록된 구현이 없습니다.

Tasks

Diversity

Similar Papers 제목 키워드 기반

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization

2026-06-08 · Mohammad Beigi, Ming Jin, Lifu Huang arxiv

Reward hacking is usually studied after it becomes visible, once a model earns high proxy reward while failing the intended task. We instead study what proxy RL teaches before that failure appears. We introduce Proxy Rew…

Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions

2021-08-20 · Hammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt 외

There is burgeoning interest in designing AI-based systems to assist humans in designing computing systems, including tools that automatically generate computer code. The most notable of these comes in the form of the fi…

Code GenerationDiversityLanguage ModelingLanguage Modelling

Reverse CAPTCHA: Evaluating LLM Susceptibility to Invisible Unicode Instruction Injection

2026-02-26 · Marcus Graves arxiv

We introduce Reverse CAPTCHA, an evaluation framework that tests whether large language models follow invisible Unicode-encoded instructions embedded in otherwise normal-looking text. Unlike traditional CAPTCHAs that dis…

Diffusion Models as Artists: Are we Closing the Gap between Humans and Machines?

2023-01-27 · Victor Boutin, Thomas Fel, Lakshya Singhal, Rishav Mukherji 외

An important milestone for AI is the development of algorithms that can produce drawings that are indistinguishable from those of humans. Here, we adapt the 'diversity vs. recognizability' scoring framework from Boutin e…

DiagnosticDiversity

We Can't Understand AI Using our Existing Vocabulary

2025-02-11 · John Hewitt, Robert Geirhos, Been Kim

This position paper argues that, in order to understand AI, we cannot rely on our existing vocabulary of human words. Instead, we should strive to develop neologisms: new words that represent precise human concepts that …

Diversity