paper-with-me

홈 › Papers

A Learning Algorithm That Attains the Human Optimum in a Repeated Human-Machine Interaction Game

2025-01-15 · Jason T. Isa, Lillian J. Ratliff, Samuel A. Burden

When humans interact with learning-based control systems, a common goal is to minimize a cost function known only to the human. For instance, an exoskeleton may adapt its assistance in an effort to minimize the human's metabolic cost-of-transport. Conventional approaches to synthesizing the learning algorithm solve an inverse problem to infer the human's cost. However, these problems can be ill-posed, hard to solve, or sensitive to problem data. Here we show a game-theoretic learning algorithm that works solely by observing human actions to find the cost minimum, avoiding the need to solve an inverse problem. We evaluate the performance of our algorithm in an extensive set of human subjects experiments, demonstrating consistent convergence to the minimum of a prescribed human cost function in scalar and multidimensional instantiations of the game. We conclude by outlining future directions for theoretical and empirical extensions of our results.

📄 PDF Abstract BibTeX arXiv:2501.08626

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

A Contextual-Bandit Oversight Game with Two-Sided Informational Asymmetry

2026-06-30 · Yunjin Tong arxiv

We study runtime human oversight of an AI agent when private information runs in both directions: the human privately knows her reward function, while the AI privately knows the quality of the action it proposes. This is…

Reinforcement Learning

Breaking Algorithmic Collusion in Human-AI Ecosystems

2025-11-26 · Natalie Collina, Eshwar Ram Arunachaleswaran, Meena Jagadeesan arxiv

AI agents are increasingly deployed in ecosystems where they repeatedly interact not only with each other but also with humans. In this work, we study these human-AI ecosystems from a theoretical perspective, focusing on…

A Multi-Plane Block-Coordinate Frank-Wolfe Algorithm for Training Structural SVMs with a Costly max-Oracle

2014-08-28 · CVPR 2015 6 · Neel Shah, Vladimir Kolmogorov, Christoph H. Lampert

Structural support vector machines (SSVMs) are amongst the best performing models for structured computer vision tasks, such as semantic image segmentation or human pose estimation. Training SSVMs, however, is computatio…

Image SegmentationPose EstimationSemantic SegmentationStructured Prediction

TDFlow: Agentic Workflows for Test Driven Development

2025-10-27 · Kevin Han, Siddharth Maddikayala, Tim Knappe, Om Patel 외 arxiv

We introduce TDFlow, a novel test-driven agentic workflow that frames repository-scale software engineering as a test-resolution task, specifically designed to solve human-written tests. Given a set of tests, TDFlow repe…

Program Repair

A stepped sampling method for video detection using LSTM

2021-07-18 · Dengshan Li, Rujing Wang, Chengjun Xie

Artificial neural networks that simulate human achieves great successes. From the perspective of simulating human memory method, we propose a stepped sampler based on the "repeated input". We repeatedly inputted data to …