paper-with-me

홈 › Papers

``ye word kis lang ka hai bhai?'' Testing the Limits of Word level Language Identification

2014-12-01 · WS 2014 12 · Sp Gella, ana, Kalika Bali, Monojit Choudhury
📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language IdentificationTransliteration

Similar Papers 제목 키워드 기반

Technical Evaluation of a Disruptive Approach in Homomorphic AI

2025-06-13 · Eric Filiol

We present a technical evaluation of a new, disruptive cryptographic approach to data security, known as HbHAI (Hash-based Homomorphic Artificial Intelligence). HbHAI is based on a novel class of key-dependent hash funct…

A Vocabulary-Free Multilingual Neural Tokenizer for End-to-End Task Learning

2022-04-22 · RepL4NLP (ACL) 2022 5 · Md Mofijul Islam, Gustavo Aguilar, Pragaash Ponnusamy, Clint Solomon Mathialagan 외

Subword tokenization is a commonly used input pre-processing step in most recent NLP models. However, it limits the models' ability to leverage end-to-end task learning. Its frequency-based vocabulary creation compromise…

DiversitySentiment Analysis

Word Embeddings from Large-Scale Greek Web Content

2018-10-08 · Stamatis Outsios, Konstantinos Skianis, Polykarpos Meladianos, Christos Xypolopoulos 외

Word embeddings are undoubtedly very useful components in many NLP tasks. In this paper, we present word embeddings and other linguistic resources trained on the largest to date digital Greek language corpus. We also pre…

Word Embeddings

Robust and Consistent Estimation of Word Embedding for Bangla Language by fine-tuning Word2Vec Model

2020-10-26 · Rifat Rahman

Word embedding or vector representation of word holds syntactical and semantic characteristics of a word which can be an informative feature for any machine learning-based models of natural language processing. There are…

Keyword ExtractionWord Embeddings

From Characters to Words: Hierarchical Pre-trained Language Model for Open-vocabulary Language Understanding

2023-05-23 · Li Sun, Florian Luisier, Kayhan Batmanghelich, Dinei Florencio 외

Current state-of-the-art models for natural language understanding require a preprocessing step to convert raw text into discrete tokens. This process known as tokenization relies on a pre-built vocabulary of words or su…

Language ModelingLanguage ModellingNatural Language Understanding