paper-with-me

Papers

Evaluating the evaluators: Towards human-aligned metrics for missing markers reconstruction

2024-10-18 · Taras Kucherenko, Derek Peristy, Judith Bütepage

Animation data is often obtained through optical motion capture systems, which utilize a multitude of cameras to establish the position of optical markers. However, system errors or occlusions can result in missing markers, the manual cleaning of which can be time-consuming. This has sparked interest in machine learning-based solutions for missing marker reconstruction in the academic community. Most academic papers utilize a simplistic mean square error as the main metric. In this paper, we show that this metric does not correlate with subjective perception of the fill quality. Additionally, we introduce and evaluate a set of better-correlated metrics that can drive progress in the field.

📄 PDF Abstract BibTeX arXiv:2410.14334

Code (0)

등록된 구현이 없습니다.

Tasks

Missing Markers ReconstructionPosition

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

A Gamified Evaluation and Recruitment Platform for Low Resource Language Machine Translation Systems

2025-06-13 · Carlos Rafael Catalan

Human evaluators provide necessary contributions in evaluating large language models. In the context of Machine Translation (MT) systems for low-resource languages (LRLs), this is made even more apparent since popular au…

Machine Translation

MILE-RefHumEval: A Reference-Free, Multi-Independent LLM Framework for Human-Aligned Evaluation

2026-02-10 · Nalin Srun, Parisa Rastin, Guénaël Cabanes, Lydia Boudjeloud Assala arxiv

We introduce MILE-RefHumEval, a reference-free framework for evaluating Large Language Models (LLMs) without ground-truth annotations or evaluator coordination. It leverages an ensemble of independently prompted evaluato…

Image Captioning

Omni-Judge: Can Omni-LLMs Serve as Human-Aligned Judges for Text-Conditioned Audio-Video Generation?

2026-02-02 · Susan Liang, Chao Huang, Filippos Bellos, Yolo Yunlong Tang 외 arxiv

State-of-the-art text-to-video generation models such as Sora 2 and Veo 3 can now produce high-fidelity videos with synchronized audio directly from a textual prompt, marking a new milestone in multi-modal generation. Ho…

Text-to-Video Generation

SLMEval: Entropy-Based Calibration for Human-Aligned Evaluation of Large Language Models

2025-05-21 · Roland Daynauth, Christopher Clarke, Krisztian Flautner, Lingjia Tang 외

The LLM-as-a-Judge paradigm offers a scalable, reference-free approach for evaluating language models. Although several calibration techniques have been proposed to better align these evaluators with human judgment, prio…

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation

2026-05-07 · Shichao Kan, Xuyang Zhang, Haojie Zhang, Zhe Zhu 외 arxiv

Evaluating image captions without references remains challenging because global embedding similarity often misses fine-grained mismatches such as hallucinated objects, missing attributes, or incorrect relations. We propo…

Image-text matching