paper-with-me

Papers Instrument Recognition

“Instrument Recognition” 태그가 달린 논문 45편 · 필터 해제

Lost in Motion: Vision Language Models Fail the Dynamic Gauges Test

2026-04-19 · Tairan Fu, Francisco Javier Santos-Martín, Javier Conde, Pedro Reviriego 외 arxiv

The digital transformation of industrial manufacturing increasingly relies on the ability of autonomous robots to interact with legacy infrastructure, particularly analog gauges. Vision-Language Models (VLMs) have the po…

Instrument Recognition

Evaluating Large Vision-language Models for Surgical Tool Detection

2026-01-23 · Nakul Poudel, Richard Simon, Cristian A. Linte arxiv

Surgery is a highly complex process, and artificial intelligence has emerged as a transformative force in supporting surgical guidance and decision-making. However, the unimodal nature of most current AI systems limits t…

Zero-shot GeneralizationSurgical tool detectionInstrument Recognition

Video Dataset for Surgical Phase, Keypoint, and Instrument Recognition in Laparoscopic Surgery (PhaKIR)

2025-11-09 · Tobias Rueckert, Raphaela Maerkl, David Rauber, Leonard Klausmann 외 arxiv

Robotic- and computer-assisted minimally invasive surgery (RAMIS) is increasingly relying on computer vision methods for reliable instrument recognition and surgical workflow understanding. Developing such systems often …

Surgical phase recognitionInstrument RecognitionInstance SegmentationScene Understanding

Persian Musical Instruments Classification Using Polyphonic Data Augmentation

2025-11-07 · Diba Hadi Esfangereh, Mohammad Hossein Sameti, Sepehr Harfi Moridani, Leili Javidpour 외 arxiv

Musical instrument classification is essential for music information retrieval (MIR) and generative music systems. However, research on non-Western traditions, particularly Persian music, remains limited. We address this…

Instrument RecognitionInformation RetrievalData AugmentationMusic Generation

Identifying Surgical Instruments in Laparoscopy Using Deep Learning Instance Segmentation

2025-08-29 · Sabrina Kletz, Klaus Schoeffmann, Jenny Benois-Pineau, Heinrich Husslein arxiv

Recorded videos from surgeries have become an increasingly important information source for the field of medical endoscopy, since the recorded footage shows every single detail of the surgery. However, while video record…

Instrument RecognitionInstance Segmentation

Comparative validation of surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation in endoscopy: Results of the PhaKIR 2024 challenge

2025-07-22 · Tobias Rueckert, David Rauber, Raphaela Maerkl, Leonard Klausmann 외 arxiv

Reliable recognition and localization of surgical instruments in endoscopic video recordings are foundational for a wide range of applications in computer- and robot-assisted minimally invasive surgery (RAMIS), including…

Surgical phase recognitionInstrument RecognitionInstance SegmentationScene Understanding

A Hierarchical Deep Learning Approach for Minority Instrument Detection

2025-06-26 · Dylan Sechet, Francesca Bugiotti, Matthieu Kowalski, Edouard d'Hérouville 외

Identifying instrument activities within audio excerpts is vital in music information retrieval, with significant implications for music cataloging and discovery. Prior deep learning endeavors in musical instrument recog…

Deep LearningDiversityInformation RetrievalInstrument Recognition+1

Can DeepSeek Reason Like a Surgeon? An Empirical Evaluation for Vision-Language Understanding in Robotic-Assisted Surgery

2025-03-29 · Boyi Ma, Yanguang Zhao, Jie Wang, Guankun Wang 외

The DeepSeek models have shown exceptional performance in general scene understanding, question-answering (QA), and text generation tasks, owing to their efficient training paradigm and strong reasoning capabilities. In …

Action UnderstandingInstrument RecognitionLarge Language ModelPosition+4

M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP

2025-03-28 · Daisuke Niizumi, Daiki Takeuchi, Masahiro Yasuda, Binh Thien Nguyen 외

Contrastive language-audio pre-training (CLAP) has addressed audio-language tasks such as audio-text retrieval by aligning audio and text in a common feature space. While CLAP addresses general audio-language tasks, its …

Audio captioningAudio ClassificationAudio TaggingAudio to Text Retrieval+13

SurgRAW: Multi-Agent Workflow with Chain-of-Thought Reasoning for Surgical Intelligence

2025-03-13 · Chang Han Low, Ziyue Wang, Tianyi Zhang, Zhitao Zeng 외

Integration of Vision-Language Models (VLMs) in surgical intelligence is hindered by hallucinations, domain knowledge gaps, and limited understanding of task interdependencies within surgical scenes, undermining clinical…

Action RecognitionInstrument RecognitionRAGRetrieval-augmented Generation

Masked Latent Prediction and Classification for Self-Supervised Audio Representation Learning

2025-02-17 · ICASSP 2025 3 · Aurian Quelennec, Pierre Chouteau, Geoffroy Peeters, Slim Essid

Recently, self-supervised learning methods based on masked latent prediction have proven to encode input data into powerful representations. However, during training, the learned latent space can be further transformed t…

Audio ClassificationAudio TaggingClassificationEnvironmental Sound Classification+9

MIRFLEX: Music Information Retrieval Feature Library for Extraction

2024-11-01 · Anuradha Chopra, Abhinaba Roy, Dorien Herremans

This paper introduces an extendable modular system that compiles a range of music feature extraction models to aid music information retrieval research. The features include musical elements like key, downbeats, and genr…

BenchmarkingInformation RetrievalInstrument RecognitionMusic Information Retrieval+2

Deep Learning for Surgical Instrument Recognition and Segmentation in Robotic-Assisted Surgeries: A Systematic Review

2024-10-09 · Fatimaelzahraa Ali Ahmed, Mahmoud Yousef, Mariam Ali Ahmed, Hasan Omar Ali 외

Applying deep learning (DL) for annotating surgical instruments in robot-assisted minimally invasive surgeries (MIS) represents a significant advancement in surgical technology. This systematic review examines 48 studies…

Instrument RecognitionSurgical tool detection

PitVis-2023 Challenge: Workflow Recognition in videos of Endoscopic Pituitary Surgery

2024-09-02 · Adrito Das, Danyal Z. Khan, Dimitrios Psychogyios, Yitong Zhang 외

The field of computer vision applied to videos of minimally invasive surgery is ever-growing. Workflow recognition pertains to the automated recognition of various aspects of a surgery: including which surgical steps are…

Instrument Recognition

I can listen but cannot read: An evaluation of two-tower multimodal systems for instrument recognition

2024-07-25 · Yannis Vasilakis, Rachel Bittner, Johan Pauwels

Music two-tower multimodal systems integrate audio and text modalities into a joint audio-text space, enabling direct comparison between songs and their corresponding labels. These systems enable new approaches for class…

Instrument RecognitionRetrievalzero-shot-classificationZero-Shot Learning

A Stem-Agnostic Single-Decoder System for Music Source Separation Beyond Four Stems

2024-06-26 · Karn N. Watcharasupat, Alexander Lerch

Despite significant recent progress across multiple subtasks of audio source separation, few music source separation systems support separation beyond the four-stem vocals, drums, bass, and other (VDBO) setup. Of the ver…

Audio Source SeparationDecoderInstrument RecognitionMusic Source Separation

WeakSurg: Weakly supervised surgical instrument segmentation using temporal equivariance and semantic continuity

2024-03-14 · Qiyuan Wang, Yanzhe Liu, Shang Zhao, Rong Liu 외

For robotic surgical videos, instrument presence annotations are typically recorded with video streams, which offering the potential to reduce the manually annotated costs for segmentation. However, weakly supervised sur…

Instance SegmentationInstrument RecognitionRepresentation LearningSegmentation+2

Dynamic Convolutional Neural Networks as Efficient Pre-trained Audio Models

2023-10-24 · Florian Schmid, Khaled Koutini, Gerhard Widmer

The introduction of large-scale audio datasets, such as AudioSet, paved the way for Transformers to conquer the audio domain and replace CNNs as the state-of-the-art neural network architecture for many tasks. Audio Spec…

Audio ClassificationAudio TaggingInstrument RecognitionKnowledge Distillation

Self-refining of Pseudo Labels for Music Source Separation with Noisy Labeled Data

2023-07-24 · Junghyun Koo, Yunkee Chae, Chang-Bin Jeon, Kyogu Lee

Music source separation (MSS) faces challenges due to the limited availability of correctly-labeled individual instrument tracks. With the push to acquire larger datasets to improve MSS performance, the inevitability of …

Instrument RecognitionMusic Source Separation

Transfer Learning and Bias Correction with Pre-trained Audio Embeddings

2023-07-20 · Changhong Wang, Gaël Richard, Brian McFee

Deep neural network models have become the dominant approach to a large variety of tasks within music information retrieval (MIR). These models generally require large amounts of (annotated) training data to achieve high…

Information RetrievalInstrument RecognitionMusic Information RetrievalRetrieval+1
1–20 / 45 다음 →