paper-with-me

홈 › Papers

Self-Supervised Behavior Cloned Transformers are Path Crawlers for Text Games

2023-12-07 · Ruoyao Wang, Peter Jansen

In this work, we introduce a self-supervised behavior cloning transformer for text games, which are challenging benchmarks for multi-step reasoning in virtual environments. Traditionally, Behavior Cloning Transformers excel in such tasks but rely on supervised training data. Our approach auto-generates training data by exploring trajectories (defined by common macro-action sequences) that lead to reward within the games, while determining the generality and utility of these trajectories by rapidly training small models then evaluating their performance on unseen development games. Through empirical analysis, we show our method consistently uncovers generalizable training data, achieving about 90\% performance of supervised systems across three benchmark text games.

📄 PDF Abstract BibTeX arXiv:2312.04657

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Behavior Cloned Transformers are Neurosymbolic Reasoners

2022-10-13 · Ruoyao Wang, Peter Jansen, Marc-Alexandre Côté, Prithviraj Ammanabrolu

In this work, we explore techniques for augmenting interactive agents with information from symbolic modules, much like humans use tools like calculators and GPS systems to assist with arithmetic and navigation. We test …

Common Sense Reasoning

Self-Supervised Vision Transformers Learn Visual Concepts in Histopathology

2022-03-01 · Richard J. Chen, Rahul G. Krishnan

Tissue phenotyping is a fundamental task in learning objective characterizations of histopathologic biomarkers within the tumor-immune microenvironment in cancer pathology. However, whole-slide imaging (WSI) is a complex…

DiversityKnowledge DistillationTransfer Learning

Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent Space

2025-09-30 · Houjun Liu, Shikhar Murty, Christopher D. Manning, Róbert Csordás arxiv

Current approaches for scaling inference-time compute in transformers train them to emit explicit chain-of-thought tokens before producing an answer. While these methods are powerful, they are limited because they cannot…

Recommender Transformers with Behavior Pathways

2022-06-13 · Zhiyu Yao, Xinyang Chen, Sinan Wang, Qinyan Dai 외

Sequential recommendation requires the recommender to capture the evolving behavior characteristics from logged user behavior data for accurate recommendations. However, user behavior sequences are viewed as a script wit…

Sequential Recommendation

Self-supervised pretraining of vision transformers for animal behavioral analysis and neural encoding

2025-07-13 · Yanchen Wang, Han Yu, Ari Blau, Yizi Zhang 외

The brain can only be fully understood through the lens of the behavior it generates -- a guiding principle in modern neuroscience research that nevertheless presents significant technical challenges. Many studies captur…

Action SegmentationContrastive LearningPose Estimation