paper-with-me

Papers Word Translation

“Word Translation” 태그가 달린 논문 125편 · 필터 해제

Cross-Domain Bilingual Lexicon Induction via Pretrained Language Models

2025-05-29 · Qiuyu Ding, Zhiqiang Cao, Hailong Cao, Tiejun Zhao

Bilingual Lexicon Induction (BLI) is generally based on common domain data to obtain monolingual word embedding, and by aligning the monolingual word embeddings to obtain the cross-lingual embeddings which are used to ge…

Bilingual Lexicon InductionWord EmbeddingsWord Translation

Semantic Pivots Enable Cross-Lingual Transfer in Large Language Models

2025-05-22 · Kaiyu He, Tong Zhou, Yubo Chen, Delai Qiu 외

Large language models (LLMs) demonstrate remarkable ability in cross-lingual tasks. Understanding how LLMs acquire this ability is crucial for their interpretability. To quantify the cross-lingual ability of LLMs accurat…

AttributeCross-Lingual TransferTranslationWord Translation

Deep Reasoning Translation via Reinforcement Learning

2025-04-14 · Jiaan Wang, Fandong Meng, Jie zhou

Recently, deep reasoning LLMs (e.g., OpenAI o1/o3 and DeepSeek-R1) have shown promising performance in various complex tasks. Free translation is an important and interesting task in the multilingual world, which require…

reinforcement-learningReinforcement LearningTranslationWord Translation

Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in Transformers

2024-11-13 · Clément Dumas, Chris Wendler, Veniamin Veselovsky, Giovanni Monea 외

A central question in multilingual language modeling is whether large language models (LLMs) develop a universal concept representation, disentangled from specific languages. In this paper, we address this question by an…

Language ModelingLanguage ModellingTranslationWord Translation

Towards Cross-Cultural Machine Translation with Retrieval-Augmented Generation from Multilingual Knowledge Graphs

2024-10-17 · Simone Conia, Daniel Lee, Min Li, Umar Farooq Minhas 외

Translating text that contains entity names is a challenging task, as cultural-related references can vary significantly across languages. These variations may also be caused by transcreation, an adaptation process that …

Knowledge GraphsMachine TranslationRetrievalRetrieval-augmented Generation+3

Optimizing Rare Word Accuracy in Direct Speech Translation with a Retrieval-and-Demonstration Approach

2024-09-13 · Siqi Li, Danni Liu, Jan Niehues

Direct speech translation (ST) models often struggle with rare words. Incorrect translation of these words can have severe consequences, impacting translation quality and user trust. While rare word translation is inhere…

In-Context LearningRetrievalSparse LearningSpeech-to-Text+2

Improving Rare Word Translation With Dictionaries and Attention Masking

2024-08-17 · Kenneth J. Sible, David Chiang

In machine translation, rare words continue to be a problem for the dominant encoder-decoder architecture, especially in low-resource and out-of-domain translation settings. Human translators solve this problem with mono…

DecoderMachine TranslationTranslationWord Translation

A Japanese-Chinese Parallel Corpus Using Crowdsourcing for Web Mining

2024-05-15 · Masaaki Nagata, Makoto Morishita, Katsuki Chousa, Norihito Yasuda

Using crowdsourcing, we collected more than 10,000 URL pairs (parallel top page pairs) of bilingual websites that contain parallel documents and created a Japanese-Chinese parallel corpus of 4.6M sentence pairs from thes…

SentenceTranslationWord Translation

Enhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word Alignment

2024-04-03 · Zhongtao Miao, Qiyu Wu, Kaiyan Zhao, Zilong Wu 외

The field of cross-lingual sentence embeddings has recently experienced significant advancements, but research concerning low-resource languages has lagged due to the scarcity of parallel corpora. This paper shows that c…

RetrievalSentenceSentence EmbeddingSentence-Embedding+4

LexC-Gen: Generating Data for Extremely Low-Resource Languages with Large Language Models and Bilingual Lexicons

2024-02-21 · Zheng-Xin Yong, Cristina Menghini, Stephen H. Bach

Data scarcity in low-resource languages can be addressed with word-to-word translations from labeled task data in high-resource languages using bilingual lexicons. However, bilingual lexicons often have limited lexical o…

Sentiment AnalysisTopic ClassificationTranslationWord Translation

Self-Augmented In-Context Learning for Unsupervised Word Translation

2024-02-15 · Yaoyiran Li, Anna Korhonen, Ivan Vulić

Recent work has shown that, while large language models (LLMs) demonstrate strong word translation or bilingual lexicon induction (BLI) capabilities in few-shot setups, they still cannot match the performance of 'traditi…

Bilingual Lexicon InductionCross-Lingual Word EmbeddingsFew-Shot LearningIn-Context Learning+8

Fingerspelling PoseNet: Enhancing Fingerspelling Translation with Pose-Based Transformer Models

2023-11-20 · Pooya Fayyazsanavi, Negar Nejatishahidin, Jana Kosecka

We address the task of American Sign Language fingerspelling translation using videos in the wild. We exploit advances in more accurate hand pose estimation and propose a novel architecture that leverages the transformer…

DecoderHand Pose EstimationLanguage ModelingLanguage Modelling+4

ProMap: Effective Bilingual Lexicon Induction via Language Model Prompting

2023-10-28 · Abdellah El Mekki, Muhammad Abdul-Mageed, ElMoatez Billah Nagoudi, Ismail Berrada 외

Bilingual Lexicon Induction (BLI), where words are translated between two languages, is an important NLP task. While noticeable progress on BLI in rich resource languages using static word embeddings has been achieved. T…

Bilingual Lexicon InductionLanguage ModelingLanguage ModellingRe-Ranking+3

On Bilingual Lexicon Induction with Large Language Models

2023-10-21 · Yaoyiran Li, Anna Korhonen, Ivan Vulić

Bilingual Lexicon Induction (BLI) is a core task in multilingual NLP that still, to a large extent, relies on calculating cross-lingual word representations. Inspired by the global paradigm shift in NLP towards Large Lan…

Bilingual Lexicon InductionCross-Lingual Word EmbeddingsFew-Shot LearningIn-Context Learning+7

Tik-to-Tok: Translating Language Models One Token at a Time: An Embedding Initialization Strategy for Efficient Language Adaptation

2023-10-05 · François Remy, Pieter Delobelle, Bettina Berendt, Kris Demuynck 외

Training monolingual language models for low and mid-resource languages is made challenging by limited and often inadequate pretraining data. In this study, we propose a novel model conversion strategy to address this is…

Word Translation

Media of Langue: The Interface for Exploring Word Translation Network/Space

2023-08-25 · Goki Muramoto, Atsuki Sato, Takayoshi Koyama

In the human activity of word translation, two languages face each other, mutually searching their own language system for the semantic place of words in the other language. We discover the huge network formed by the cha…

TranslationWord Translation

Universal Language Modelling agent

2023-06-10 · Anees Aslam

Large Language Models are designed to understand complex Human Language. Yet, Understanding of animal language has long intrigued researchers striving to bridge the communication gap between humans and other species. Thi…

Language ModellingWord Translation

A European and Brazilian cross-national investigation into the Portuguese translation of soundscape perceptual attributes within the SATP project

2023-06-10 · Applied Acoustics 2023 6 · Sónia Monteiro Antunes, Ranny Loureiro Xavier Nascimento Michalski, Maria Luiza de Ulhôa Carvalho, Sónia Alves 외

This paper presents a cross-national investigation into the soundscape perceptual attributes translation from English into European and Brazilian Portuguese. It is a study within the scope of a larger project – the Sound…

Audio Emotion RecognitionSoundscape evaluationTranslationWord Translation

TransDocs: Optical Character Recognition with word to word translation

2023-04-15 · Abhishek Bamotra, Phani Krishna Uppala

While OCR has been used in various applications, its output is not always accurate, leading to misfit words. This research work focuses on improving the optical character recognition (OCR) with ML techniques with integra…

Deep LearningDocument TranslationMachine TranslationOptical Character Recognition+3

PEACH: Pre-Training Sequence-to-Sequence Multilingual Models for Translation with Semi-Supervised Pseudo-Parallel Document Generation

2023-04-03 · Alireza Salemi, Amirhossein Abaskohi, Sara Tavakoli, Yadollah Yaghoobzadeh 외

Multilingual pre-training significantly improves many multilingual NLP tasks, including machine translation. Most existing methods are based on some variants of masked language modeling and text-denoising objectives on m…

DenoisingLanguage ModelingLanguage ModellingMachine Translation+5
1–20 / 125 다음 →