paper-with-me

홈 › Papers

Deep Learning-Driven Multimodal Detection and Movement Analysis of Objects in Culinary

2025-08-21 · Tahoshin Alam Ishat, Mohammad Abdul Qayum arxiv

This is a research exploring existing models and fine tuning them to combine a YOLOv8 segmentation model, a LSTM model trained on hand point motion sequence and a ASR (whisper-base) to extract enough data for a LLM (TinyLLaMa) to predict the recipe and generate text creating a step by step guide for the cooking procedure. All the data were gathered by the author for a robust task specific system to perform best in complex and challenging environments proving the extension and endless application of computer vision in daily activities such as kitchen work. This work extends the field for many more crucial task of our day to day life.

📄 PDF Abstract BibTeX arXiv:2509.00033

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Truth and Nothing but the Truth: Multimodal Analysis for Deception Detection

2019-03-11 · Mimansa Jaiswal, Sairam Tabibu, Rajiv Bajpai

We propose a data-driven method for automatic deception detection in real-life trial data using visual and verbal cues. Using OpenFace with facial action unit recognition, we analyze the movement of facial features of th…

Deception DetectionFacial Action Unit DetectionLexical Analysis

Eye4Ref: A Multimodal Eye Movement Dataset of Referentially Complex Situations

2020-05-01 · LREC 2020 5 · {\"O}zge Alacam, Eugen Ruppert, Amr Rekaby Salama, Tobias Staron 외

Eye4Ref is a rich multimodal dataset of eye-movement recordings collected from referentially complex situated settings where the linguistic utterances and their visual referential world were available to the listener. It…

Sentence

A Multimodal Emotion Recognition System: Integrating Facial Expressions, Body Movement, Speech, and Spoken Language

2024-12-23 · Kris Kraack

Traditional psychological evaluations rely heavily on human observation and interpretation, which are prone to subjectivity, bias, fatigue, and inconsistency. To address these limitations, this work presents a multimodal…

DiagnosticEmotion RecognitionMultimodal Emotion Recognition

Detection and Tracking of General Movable Objects in Large 3D Maps

2017-12-22 · Nils Bore, Johan Ekekrantz, Patric Jensfelt, John Folkesson

This paper studies the problem of detection and tracking of general objects with long-term dynamics, observed by a mobile robot moving in a large environment. A key problem is that due to the environment scale, it can on…

AnimalFormer: Multimodal Vision Framework for Behavior-based Precision Livestock Farming

2024-06-14 · Ahmed Qazi, Taha Razzaq, Asim Iqbal

We introduce a multimodal vision framework for precision livestock farming, harnessing the power of GroundingDINO, HQSAM, and ViTPose models. This integrated suite enables comprehensive behavioral analytics from video da…

Action DetectionActivity DetectionManagement