Papers Masked Language Modeling
“Masked Language Modeling” 태그가 달린 논문 475편 · 필터 해제
Generating Synthetic Free-text Medical Records with Low Re-identification Risk using Masked Language Modeling
The vast amount of available medical records has the potential to improve healthcare and biomedical research. However, privacy restrictions make these data accessible for internal use only. Recent works have addressed th…
Causal Language ModelingDe-identificationDiversityLanguage Modeling+4DomURLs_BERT: Pre-trained BERT-based Model for Malicious Domains and URLs Detection and Classification
Detecting and classifying suspicious or malicious domain names and URLs is fundamental task in cybersecurity. To leverage such indicators of compromise, cybersecurity vendors and practitioners often maintain and update b…
Language ModelingLanguage ModellingMasked Language ModelingMulti-class ClassificationVidLPRO: A $\underline{Vid}$eo-$\underline{L}$anguage $\underline{P}$re-training Framework for $\underline{Ro}$botic and Laparoscopic Surgery
We introduce VidLPRO, a novel video-language (VL) pre-training framework designed specifically for robotic and laparoscopic surgery. While existing surgical VL models primarily rely on contrastive learning, we propose a …
Computational EfficiencyContrastive LearningLanguage ModelingLanguage Modelling+4N-gram Prediction and Word Difference Representations for Language Modeling
Causal language modeling (CLM) serves as the foundational framework underpinning remarkable successes of recent large language models (LLMs). Despite its success, the training approach for next word prediction poses a po…
Causal Language ModelingLanguage ModelingLanguage ModellingMachine Translation+4Dynamic Motion Synthesis: Masked Audio-Text Conditioned Spatio-Temporal Transformers
Our research presents a novel motion generation framework designed to produce whole-body motion sequences conditioned on multiple modalities simultaneously, specifically text and audio inputs. Leveraging Vector Quantized…
Language ModelingLanguage ModellingMasked Language ModelingMotion Generation+1How transformers learn structured data: insights from hierarchical filtering
Understanding the learning process and the embedded computation in transformers is becoming a central goal for the development of interpretable AI. In the present study, we introduce a hierarchical filtering procedure fo…
Language ModelingLanguage ModellingMasked Language ModelingMistral-SPLADE: LLMs for better Learned Sparse Retrieval
Learned Sparse Retrievers (LSR) have evolved into an effective retrieval strategy that can bridge the gap between traditional keyword-based sparse retrievers and embedding-based dense retrievers. At its core, learned spa…
DecoderLanguage ModelingLanguage ModellingLarge Language Model+3Unlocking Efficiency: Adaptive Masking for Gene Transformer Models
Gene transformer models such as Nucleotide Transformer, DNABert, and LOGO are trained to learn optimal gene sequence representations by using the Masked Language Modeling (MLM) training objective over the complete Human …
Language ModelingLanguage ModellingMasked Language ModelingRepresentation LearningMIDI-to-Tab: Guitar Tablature Inference via Masked Language Modeling
Guitar tablatures enrich the structure of traditional music notation by assigning each note to a string and fret of a guitar in a particular tuning, indicating precisely where to play the note on the instrument. The prob…
DecoderLanguage ModelingLanguage ModellingMasked Language ModelingAutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
Domain reweighting is an emerging research area aimed at adjusting the relative weights of different data sources to improve the effectiveness and efficiency of LLM pre-training. We show that data mixtures that perform w…
Bilevel OptimizationLanguage ModellingMasked Language ModelingMMCLIP: Cross-modal Attention Masked Modelling for Medical Language-Image Pre-Training
Vision-and-language pretraining (VLP) in the medical field utilizes contrastive learning on image-text pairs to achieve effective transfer across tasks. Yet, current VLP approaches with the masked modeling strategy face …
Contrastive LearningLanguage ModelingLanguage ModellingMasked Language ModelingA Novel Two-Step Fine-Tuning Pipeline for Cold-Start Active Learning in Text Classification Tasks
This is the first work to investigate the effectiveness of BERT-based contextual embeddings in active learning (AL) tasks on cold-start scenarios, where traditional fine-tuning is infeasible due to the absence of labeled…
Active LearningDomain AdaptationLanguage ModelingLanguage Modelling+3Pre-Training and Prompting for Few-Shot Node Classification on Text-Attributed Graphs
The text-attributed graph (TAG) is one kind of important real-world graph-structured data with each node associated with raw texts. For TAGs, traditional few-shot node classification methods directly conduct training on …
Few-Shot LearningGraph Neural NetworkLanguage ModelingLanguage Modelling+3Promises and Pitfalls of Generative Masked Language Modeling: Theoretical Framework and Practical Guidelines
Autoregressive language models are the currently dominant paradigm for text generation, but they have some fundamental limitations that cannot be remedied by scale-for example inherently sequential and unidirectional gen…
Language ModelingLanguage ModellingMachine TranslationMasked Language Modeling+1Pseudo-perplexity in One Fell Swoop for Protein Fitness Estimation
Protein language models trained on the masked language modeling objective learn to predict the identity of hidden amino acid residues within a sequence using the remaining observable sequence as context. They do so by em…
Computational EfficiencyLanguage ModelingLanguage ModellingMasked Language ModelingHistorical Ink: Semantic Shift Detection for 19th Century Spanish
This paper explores the evolution of word meanings in 19th-century Spanish texts, with an emphasis on Latin American Spanish, using computational linguistics techniques. It addresses the Semantic Shift Detection (SSD) ta…
Masked Language ModelingSemantic Shift DetectionSemantic SimilarityLLMcap: Large Language Model for Unsupervised PCAP Failure Detection
The integration of advanced technologies into telecommunication networks complicates troubleshooting, posing challenges for manual error identification in Packet Capture (PCAP) data. This manual approach, requiring subst…
Language ModelingLanguage ModellingLarge Language ModelMasked Language Modeling+1ESALE: Enhancing Code-Summary Alignment Learning for Source Code Summarization
(Source) code summarization aims to automatically generate succinct natural language summaries for given code snippets. Such summaries play a significant role in promoting developers to understand and maintain code. Insp…
Code SummarizationDecoderLanguage ModelingLanguage Modelling+4Adapting Multilingual LLMs to Low-Resource Languages with Knowledge Graphs via Adapters
This paper explores the integration of graph knowledge from linguistic ontologies into multilingual Large Language Models (LLMs) using adapters to improve performance for low-resource languages (LRLs) in sentiment analys…
Knowledge GraphsLanguage ModelingLanguage ModellingMasked Language Modeling+7Retrieval-style In-Context Learning for Few-shot Hierarchical Text Classification
Hierarchical text classification (HTC) is an important task with broad applications, while few-shot HTC has gained increasing interest recently. While in-context learning (ICL) with large language models (LLMs) has achie…
Contrastive Learningfew-shot-htcFew-shot HTCFew-Shot Learning+7