paper-with-me

홈 › Papers

Graph-attention-based Casual Discovery with Trust Region-navigated Clipping Policy Optimization

2024-12-27 · Shixuan Liu, Yanghe Feng, Keyu Wu, Guangquan Cheng, Jincai Huang, Zhong Liu

In many domains of empirical sciences, discovering the causal structure within variables remains an indispensable task. Recently, to tackle with unoriented edges or latent assumptions violation suffered by conventional methods, researchers formulated a reinforcement learning (RL) procedure for causal discovery, and equipped REINFORCE algorithm to search for the best-rewarded directed acyclic graph. The two keys to the overall performance of the procedure are the robustness of RL methods and the efficient encoding of variables. However, on the one hand, REINFORCE is prone to local convergence and unstable performance during training. Neither trust region policy optimization, being computationally-expensive, nor proximal policy optimization (PPO), suffering from aggregate constraint deviation, is decent alternative for combinatory optimization problems with considerable individual subactions. We propose a trust region-navigated clipping policy optimization method for causal discovery that guarantees both better search efficiency and steadiness in policy optimization, in comparison with REINFORCE, PPO and our prioritized sampling-guided REINFORCE implementation. On the other hand, to boost the efficient encoding of variables, we propose a refined graph attention encoder called SDGAT that can grasp more feature information without priori neighbourhood information. With these improvements, the proposed method outperforms former RL method in both synthetic and benchmark datasets in terms of output results and optimization robustness.

📄 PDF Abstract BibTeX arXiv:2412.19578

Code (0)

등록된 구현이 없습니다.

Tasks

Causal DiscoveryGraph AttentionReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
REINFORCE REINFORCE is a Monte Carlo variant of a policy gradient algorithm in reinforcement learning. The agent collects samples of an episode using its current policy, and uses it to…
Entropy Regularization 설명 없음
PPO Proximal Policy Optimization, or PPO, is a policy gradient method for reinforcement learning. The motivation was to have an algorithm with the data efficiency and reliable…

Similar Papers 제목 키워드 기반

Intervention-based Recurrent Casual Model for Non-stationary Video Causal Discovery

2021-09-29 · Yuke Li, Kenneth Li, Pin Wang, Donglai Wei 외

Non-stationary casual structures are prevalent in real-world physical systems. For example, the stacked blocks interacted with one another until they fall apart, while the billiard balls are moving independently until th…

Causal DiscoverycounterfactualCounterfactual Reasoning

Prevention of Terrorist Crimes in the North Caucasus Region

2021-08-25 · Ivan Kucherkov, Mattia Masolletti

The relevance of the topic is dictated by the fact that in recent decades, the threat to international security emanating from terrorism has increased many times. Terrorist organizations have become full-fledged subjects…

The Nijmegen Corpus of Casual Czech

2014-05-01 · LREC 2014 5 · Mirjam Ernestus, Lucie Ko{\v{c}}kov{\'a}-Amortov{\'a}, Petr Pollak

This article introduces a new speech corpus, the Nijmegen Corpus of Casual Czech (NCCCz), which contains more than 30 hours of high-quality recordings of casual conversations in Common Czech, among ten groups of three ma…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

ATRACT: A Trustworthy Robotic Autonomous system to support Casualty Triage

2026-05-16 · Tasweer Ahmad, Rafael Pina, Sandip Pradhan, Arindam Sikdar 외 arxiv

At a time when drones are increasingly associated with hostile operations, we re-purpose them for humanitarian and life-saving applications. However, adapting search and rescue drones for battlefield triage remains extre…

Action ClassificationData Augmentation

Position Debiasing Fine-Tuning for Causal Perception in Long-Term Dialogue

2024-06-04 · Shixuan Fan, Wei Wei, Wendi Li, Xian-Ling Mao 외

The core of the dialogue system is to generate relevant, informative, and human-like responses based on extensive dialogue history. Recently, dialogue generation domain has seen mainstream adoption of large language mode…

Dialogue GenerationPositionResponse GenerationSentence