Papers Instrument Recognition
“Instrument Recognition” 태그가 달린 논문 45편 · 필터 해제
Lost in Motion: Vision Language Models Fail the Dynamic Gauges Test
The digital transformation of industrial manufacturing increasingly relies on the ability of autonomous robots to interact with legacy infrastructure, particularly analog gauges. Vision-Language Models (VLMs) have the po…
Instrument RecognitionEvaluating Large Vision-language Models for Surgical Tool Detection
Surgery is a highly complex process, and artificial intelligence has emerged as a transformative force in supporting surgical guidance and decision-making. However, the unimodal nature of most current AI systems limits t…
Zero-shot GeneralizationSurgical tool detectionInstrument RecognitionVideo Dataset for Surgical Phase, Keypoint, and Instrument Recognition in Laparoscopic Surgery (PhaKIR)
Robotic- and computer-assisted minimally invasive surgery (RAMIS) is increasingly relying on computer vision methods for reliable instrument recognition and surgical workflow understanding. Developing such systems often …
Surgical phase recognitionInstrument RecognitionInstance SegmentationScene UnderstandingPersian Musical Instruments Classification Using Polyphonic Data Augmentation
Musical instrument classification is essential for music information retrieval (MIR) and generative music systems. However, research on non-Western traditions, particularly Persian music, remains limited. We address this…
Instrument RecognitionInformation RetrievalData AugmentationMusic GenerationIdentifying Surgical Instruments in Laparoscopy Using Deep Learning Instance Segmentation
Recorded videos from surgeries have become an increasingly important information source for the field of medical endoscopy, since the recorded footage shows every single detail of the surgery. However, while video record…
Instrument RecognitionInstance SegmentationComparative validation of surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation in endoscopy: Results of the PhaKIR 2024 challenge
Reliable recognition and localization of surgical instruments in endoscopic video recordings are foundational for a wide range of applications in computer- and robot-assisted minimally invasive surgery (RAMIS), including…
Surgical phase recognitionInstrument RecognitionInstance SegmentationScene UnderstandingA Hierarchical Deep Learning Approach for Minority Instrument Detection
Identifying instrument activities within audio excerpts is vital in music information retrieval, with significant implications for music cataloging and discovery. Prior deep learning endeavors in musical instrument recog…
Deep LearningDiversityInformation RetrievalInstrument Recognition+1Can DeepSeek Reason Like a Surgeon? An Empirical Evaluation for Vision-Language Understanding in Robotic-Assisted Surgery
The DeepSeek models have shown exceptional performance in general scene understanding, question-answering (QA), and text generation tasks, owing to their efficient training paradigm and strong reasoning capabilities. In …
Action UnderstandingInstrument RecognitionLarge Language ModelPosition+4M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP
Contrastive language-audio pre-training (CLAP) has addressed audio-language tasks such as audio-text retrieval by aligning audio and text in a common feature space. While CLAP addresses general audio-language tasks, its …
Audio captioningAudio ClassificationAudio TaggingAudio to Text Retrieval+13SurgRAW: Multi-Agent Workflow with Chain-of-Thought Reasoning for Surgical Intelligence
Integration of Vision-Language Models (VLMs) in surgical intelligence is hindered by hallucinations, domain knowledge gaps, and limited understanding of task interdependencies within surgical scenes, undermining clinical…
Action RecognitionInstrument RecognitionRAGRetrieval-augmented GenerationMasked Latent Prediction and Classification for Self-Supervised Audio Representation Learning
Recently, self-supervised learning methods based on masked latent prediction have proven to encode input data into powerful representations. However, during training, the learned latent space can be further transformed t…
Audio ClassificationAudio TaggingClassificationEnvironmental Sound Classification+9MIRFLEX: Music Information Retrieval Feature Library for Extraction
This paper introduces an extendable modular system that compiles a range of music feature extraction models to aid music information retrieval research. The features include musical elements like key, downbeats, and genr…
BenchmarkingInformation RetrievalInstrument RecognitionMusic Information Retrieval+2Deep Learning for Surgical Instrument Recognition and Segmentation in Robotic-Assisted Surgeries: A Systematic Review
Applying deep learning (DL) for annotating surgical instruments in robot-assisted minimally invasive surgeries (MIS) represents a significant advancement in surgical technology. This systematic review examines 48 studies…
Instrument RecognitionSurgical tool detectionPitVis-2023 Challenge: Workflow Recognition in videos of Endoscopic Pituitary Surgery
The field of computer vision applied to videos of minimally invasive surgery is ever-growing. Workflow recognition pertains to the automated recognition of various aspects of a surgery: including which surgical steps are…
Instrument RecognitionI can listen but cannot read: An evaluation of two-tower multimodal systems for instrument recognition
Music two-tower multimodal systems integrate audio and text modalities into a joint audio-text space, enabling direct comparison between songs and their corresponding labels. These systems enable new approaches for class…
Instrument RecognitionRetrievalzero-shot-classificationZero-Shot LearningA Stem-Agnostic Single-Decoder System for Music Source Separation Beyond Four Stems
Despite significant recent progress across multiple subtasks of audio source separation, few music source separation systems support separation beyond the four-stem vocals, drums, bass, and other (VDBO) setup. Of the ver…
Audio Source SeparationDecoderInstrument RecognitionMusic Source SeparationWeakSurg: Weakly supervised surgical instrument segmentation using temporal equivariance and semantic continuity
For robotic surgical videos, instrument presence annotations are typically recorded with video streams, which offering the potential to reduce the manually annotated costs for segmentation. However, weakly supervised sur…
Instance SegmentationInstrument RecognitionRepresentation LearningSegmentation+2Dynamic Convolutional Neural Networks as Efficient Pre-trained Audio Models
The introduction of large-scale audio datasets, such as AudioSet, paved the way for Transformers to conquer the audio domain and replace CNNs as the state-of-the-art neural network architecture for many tasks. Audio Spec…
Audio ClassificationAudio TaggingInstrument RecognitionKnowledge DistillationSelf-refining of Pseudo Labels for Music Source Separation with Noisy Labeled Data
Music source separation (MSS) faces challenges due to the limited availability of correctly-labeled individual instrument tracks. With the push to acquire larger datasets to improve MSS performance, the inevitability of …
Instrument RecognitionMusic Source SeparationTransfer Learning and Bias Correction with Pre-trained Audio Embeddings
Deep neural network models have become the dominant approach to a large variety of tasks within music information retrieval (MIR). These models generally require large amounts of (annotated) training data to achieve high…
Information RetrievalInstrument RecognitionMusic Information RetrievalRetrieval+1