Deep Learning-Driven Multimodal Detection and Movement Analysis of Objects in Culinary
This is a research exploring existing models and fine tuning them to combine a YOLOv8 segmentation model, a LSTM model trained on hand point motion sequence and a ASR (whisper-base) to extract enough data for a LLM (TinyLLaMa) to predict the recipe and generate text creating a step by step guide for the cooking procedure. All the data were gathered by the author for a robust task specific system to perform best in complex and challenging environments proving the extension and endless application of computer vision in daily activities such as kitchen work. This work extends the field for many more crucial task of our day to day life.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
The Truth and Nothing but the Truth: Multimodal Analysis for Deception Detection
We propose a data-driven method for automatic deception detection in real-life trial data using visual and verbal cues. Using OpenFace with facial action unit recognition, we analyze the movement of facial features of th…
Deception DetectionFacial Action Unit DetectionLexical AnalysisEye4Ref: A Multimodal Eye Movement Dataset of Referentially Complex Situations
Eye4Ref is a rich multimodal dataset of eye-movement recordings collected from referentially complex situated settings where the linguistic utterances and their visual referential world were available to the listener. It…
SentenceA Multimodal Emotion Recognition System: Integrating Facial Expressions, Body Movement, Speech, and Spoken Language
Traditional psychological evaluations rely heavily on human observation and interpretation, which are prone to subjectivity, bias, fatigue, and inconsistency. To address these limitations, this work presents a multimodal…
DiagnosticEmotion RecognitionMultimodal Emotion RecognitionDetection and Tracking of General Movable Objects in Large 3D Maps
This paper studies the problem of detection and tracking of general objects with long-term dynamics, observed by a mobile robot moving in a large environment. A key problem is that due to the environment scale, it can on…
AnimalFormer: Multimodal Vision Framework for Behavior-based Precision Livestock Farming
We introduce a multimodal vision framework for precision livestock farming, harnessing the power of GroundingDINO, HQSAM, and ViTPose models. This integrated suite enables comprehensive behavioral analytics from video da…
Action DetectionActivity DetectionManagement