Learning State Representations in Complex Systems with Multimodal Data
Representation learning becomes especially important for complex systems with multimodal data sources such as cameras or sensors. Recent advances in reinforcement learning and optimal control make it possible to design control algorithms on these latent representations, but the field still lacks a large-scale standard dataset for unified comparison. In this work, we present a large-scale dataset and evaluation framework for representation learning for the complex task of landing an airplane. We implement and compare several approaches to representation learning on this dataset in terms of the quality of simple supervised learning tasks and disentanglement scores. The resulting representations can be used for further tasks such as anomaly detection, optimal control, model-based reinforcement learning, and other applications.
Code (0)
등록된 구현이 없습니다.
Tasks
Anomaly DetectionDisentanglementModel-based Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Representation LearningSimilar Papers 제목 키워드 기반
Improving Multimodal fusion via Mutual Dependency Maximisation
Multimodal sentiment analysis is a trending area of research, and the multimodal fusion is one of its most active topic. Acknowledging humans communicate through a variety of channels (i.e visual, acoustic, linguistic), …
Multimodal Sentiment AnalysisSentiment AnalysisDo Recommender Systems Really Leverage Multimodal Content? A Comprehensive Analysis on Multimodal Representations for Recommendation
Multimodal Recommender Systems aim to improve recommendation accuracy by integrating heterogeneous content, such as images and textual metadata. While effective, it remains unclear whether their gains stem from true mult…
Adversarial Multimodal Domain Transfer for Video-Level Sentiment Analysis
Video-level sentiment analysis is a challenging task and requires systems to obtain discriminative multimodal representations that can capture difference in sentiments across various modalities. However, due to diverse …
Multimodal Sentiment AnalysisSentiment AnalysisEffMulti: Efficiently Modeling Complex Multimodal Interactions for Emotion Analysis
Humans are skilled in reading the interlocutor's emotion from multimodal signals, including spoken words, simultaneous speech, and facial expressions. It is still a challenge to effectively decode emotions from the compl…
Emotion RecognitionEvaluating Multimodal Representations on Visual Semantic Textual Similarity
The combination of visual and textual representations has produced excellent results in tasks such as image captioning and visual question answering, but the inference capabilities of multimodal representations are large…
BenchmarkingImage CaptioningNatural Language InferenceQuestion Answering+3