Papers Task 2
“Task 2” 태그가 달린 논문 572편 · 필터 해제
Pun Intended: Multi-Agent Translation of Wordplay with Contrastive Learning and Phonetic-Semantic Embeddings
Translating wordplay across languages presents unique challenges that have long confounded both professional human translators and machine translation systems. This research proposes a novel approach for translating puns…
Contrastive LearningMachine TranslationTask 2TranslationBUT System for the MLC-SLM Challenge
We present a two-speaker automatic speech recognition (ASR) system that combines DiCoW -- a diarization-conditioned variant of Whisper -- with DiariZen, a diarization pipeline built on top of Pyannote. We first evaluate …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Domain Adaptationspeech-recognition+2Developing a High-performance Framework for Speech Emotion Recognition in Naturalistic Conditions Challenge for Emotional Attribute Prediction
Speech emotion recognition (SER) in naturalistic conditions presents a significant challenge for the speech processing community. Challenges include disagreement in labeling among annotators and imbalanced data distribut…
AttributeEmotion RecognitionMulti-Task LearningSpeech Emotion Recognition+1MCP-Zero: Active Tool Discovery for Autonomous LLM Agents
True intelligence requires active capability acquisition, yet current LLM agents inject pre-defined tool schemas into prompts, reducing models to passive selectors and falling short of robust general-purpose agency. We i…
RetrievalSemantic SimilaritySemantic Textual SimilarityTask 2Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing
Visual Speech Recognition (VSR) transcribes speech by analyzing lip movements. Recently, Large Language Models (LLMs) have been integrated into VSR systems, leading to notable performance improvements. However, the poten…
speech-recognitionSpeech RecognitionTask 2Visual Speech RecognitionPrompt Tuning Vision Language Models with Margin Regularizer for Few-Shot Learning under Distribution Shifts
Recently, Vision-Language foundation models like CLIP and ALIGN, which are pre-trained on large-scale data have shown remarkable zero-shot generalization to diverse datasets with different classes and even domains. In th…
Few-Shot LearningTask 2Zero-shot GeneralizationLLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
Assessing the capacity of Large Language Models (LLMs) to plan and reason within the constraints of interactive environments is crucial for developing capable AI agents. We introduce $\textbf{LLM-BabyBench}$, a new bench…
Task 2TUMS: Enhancing Tool-use Abilities of LLMs with Multi-structure Handlers
Recently, large language models(LLMs) have played an increasingly important role in solving a wide range of NLP tasks, leveraging their capabilities of natural language understanding and generating. Integration with exte…
Natural Language UnderstandingTask 21$^{st}$ Place Solution of WWW 2025 EReL@MIR Workshop Multimodal CTR Prediction Challenge
The WWW 2025 EReL@MIR Workshop Multimodal CTR Prediction Challenge focuses on effectively applying multimodal embedding features to improve click-through rate (CTR) prediction in recommender systems. This technical repor…
Click-Through Rate PredictionRecommendation SystemsTask 2Team ACK at SemEval-2025 Task 2: Beyond Word-for-Word Machine Translation for English-Korean Pairs
Translating knowledge-intensive and entity-rich text between English and Korean requires transcreation to preserve language-specific and cultural nuances beyond literal, phonetic or word-for-word conversion. We evaluate …
Machine TranslationTask 2TranslationFeature Fusion Revisited: Multimodal CTR Prediction for MMCTR Challenge
With the rapid advancement of Multimodal Large Language Models (MLLMs), an increasing number of researchers are exploring their application in recommendation systems. However, the high latency associated with large model…
Click-Through Rate PredictionInformation RetrievalPredictionRecommendation Systems+2FinBERT-QA: Financial Question Answering with pre-trained BERT Language Models
Motivated by the emerging demand in the financial industry for the automatic analysis of unstructured and structured data at scale, Question Answering (QA) systems can provide lucrative and competitive advantages to comp…
Answer SelectionInformation RetrievalQuestion AnsweringRe-Ranking+2BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts
Mixture-of-Experts (MoE) have emerged as a powerful architecture for large language models (LLMs), enabling efficient scaling of model capacity while maintaining manageable computational costs. The key advantage lies in …
Backdoor AttackMixture-of-ExpertsTask 2Quadratic Interest Network for Multimodal Click-Through Rate Prediction
Multimodal click-through rate (CTR) prediction is a key technique in industrial recommender systems. It leverages heterogeneous modalities such as text, images, and behavioral logs to capture high-order feature interacti…
Click-Through Rate PredictionMultimodal RecommendationPredictionRecommendation Systems+2Data Augmentation Using Neural Acoustic Fields With Retrieval-Augmented Pre-training
This report details MERL's system for room impulse response (RIR) estimation submitted to the Generative Data Augmentation Workshop at ICASSP 2025 for Augmenting RIR Data (Task 1) and Improving Speaker Distance Estimatio…
Data AugmentationRetrievalRoom Impulse Response (RIR)Task 2HausaNLP at SemEval-2025 Task 2: Entity-Aware Fine-tuning vs. Prompt Engineering in Entity-Aware Machine Translation
This paper presents our findings for SemEval 2025 Task 2, a shared task on entity-aware machine translation (EA-MT). The goal of this task is to develop translation models that can accurately translate English sentences …
Machine TranslationPrompt EngineeringTask 2TranslationTowards Universal Learning-based Model for Cardiac Image Reconstruction: Summary of the CMRxRecon2024 Challenge
Cardiovascular magnetic resonance (CMR) imaging offers diverse contrasts for non-invasive assessment of cardiac function and myocardial characterization. However, CMR often requires the acquisition of many contrasts, and…
BenchmarkingImage ReconstructionOut-of-Distribution GeneralizationPrompt Learning+1Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation
Reliable evaluation of AI models is critical for scientific progress and practical application. While existing VLM benchmarks provide general insights into model capabilities, their heterogeneous designs and limited focu…
BenchmarkingLanguage ModelingLanguage ModellingTask 2MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction
World models that forecast environmental changes from actions are vital for autonomous driving models with strong generalization. The prevailing driving world model mainly build on video prediction model. Although these …
2kAutonomous DrivingTask 2Video PredictionFine-Tuning Open-Source Large Language Models to Improve Their Performance on Radiation Oncology Tasks: A Feasibility Study to Investigate Their Potential Clinical Applications in Radiation Oncology
Background: The radiation oncology clinical practice involves many steps relying on the dynamic interplay of abundant text data. Large language models have displayed remarkable capabilities in processing complex text inf…
DiagnosticTask 2