paper-with-me

Papers

Temporal and Semantic Evaluation Metrics for Foundation Models in Post-Hoc Analysis of Robotic Sub-tasks

2024-03-25 · Jonathan Salfity, Selma Wanna, Minkyu Choi, Mitch Pryor

Recent works in Task and Motion Planning (TAMP) show that training control policies on language-supervised robot trajectories with quality labeled data markedly improves agent task success rates. However, the scarcity of such data presents a significant hurdle to extending these methods to general use cases. To address this concern, we present an automated framework to decompose trajectory data into temporally bounded and natural language-based descriptive sub-tasks by leveraging recent prompting strategies for Foundation Models (FMs) including both Large Language Models (LLMs) and Vision Language Models (VLMs). Our framework provides both time-based and language-based descriptions for lower-level sub-tasks that comprise full trajectories. To rigorously evaluate the quality of our automatic labeling framework, we contribute an algorithm SIMILARITY to produce two novel metrics, temporal similarity and semantic similarity. The metrics measure the temporal alignment and semantic fidelity of language descriptions between two sub-task decompositions, namely an FM sub-task decomposition prediction and a ground-truth sub-task decomposition. We present scores for temporal similarity and semantic similarity above 90%, compared to 30% of a randomized baseline, for multiple robotic environments, demonstrating the effectiveness of our proposed framework. Our results enable building diverse, large-scale, language-supervised datasets for improved robotic TAMP.

📄 PDF Abstract BibTeX arXiv:2403.17238

Code (1)

jsalfity/task_decomposition 공식 구현

Tasks

DescriptiveMotion PlanningSemantic SimilaritySemantic Textual SimilarityTask and Motion Planning

Similar Papers 제목 키워드 기반

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing

2025-01-13 · Varun Biyyala, Bharat Chanderprakash Kathuria, Jialu Li, Youshan Zhang

Video editing models have advanced significantly, but evaluating their performance remains challenging. Traditional metrics, such as CLIP text and image scores, often fall short: text scores are limited by inadequate tra…

Objectobject-detectionObject DetectionObject Tracking+1

Evaluation Metrics for Automated Typographic Poster Generation

2024-02-10 · Sérgio M. Rebelo, J. J. Merelo, João Bicker, Penousal Machado

Computational Design approaches facilitate the generation of typographic design, but evaluating these designs remains a challenging task. In this paper, we propose a set of heuristic metrics for typographic design evalua…

Emotion Recognition

Towards Interpretable Time Series Foundation Models

2025-07-10 · Matthieu Boileau, Philippe Helluy, Jeremy Pawlus, Svitlana Vyetrenko

In this paper, we investigate the distillation of time series reasoning capabilities into small, instruction-tuned language models as a step toward building interpretable time series foundation models. Leveraging a synth…

Time Series

ASurvey: Spatiotemporal Consistency in Video Generation

2025-02-25 · Zhiyu Yin, Kehai Chen, Xuefeng Bai, Ruili Jiang 외

Video generation, by leveraging a dynamic visual generation method, pushes the boundaries of Artificial Intelligence Generated Content (AIGC). Video generation presents unique challenges beyond static image generation, r…

Image GenerationVideo Generation

Towards Generalizable Scene Change Detection

2024-09-10 · Jaewoo Kim, UeHwan Kim

Scene Change Detection (SCD) is vital for applications such as visual surveillance and mobile robotics. However, current SCD methods exhibit a bias to the temporal order of training datasets and limited performance on un…

Change DetectionScene Change Detection