Graph-attention-based Casual Discovery with Trust Region-navigated Clipping Policy Optimization
In many domains of empirical sciences, discovering the causal structure within variables remains an indispensable task. Recently, to tackle with unoriented edges or latent assumptions violation suffered by conventional methods, researchers formulated a reinforcement learning (RL) procedure for causal discovery, and equipped REINFORCE algorithm to search for the best-rewarded directed acyclic graph. The two keys to the overall performance of the procedure are the robustness of RL methods and the efficient encoding of variables. However, on the one hand, REINFORCE is prone to local convergence and unstable performance during training. Neither trust region policy optimization, being computationally-expensive, nor proximal policy optimization (PPO), suffering from aggregate constraint deviation, is decent alternative for combinatory optimization problems with considerable individual subactions. We propose a trust region-navigated clipping policy optimization method for causal discovery that guarantees both better search efficiency and steadiness in policy optimization, in comparison with REINFORCE, PPO and our prioritized sampling-guided REINFORCE implementation. On the other hand, to boost the efficient encoding of variables, we propose a refined graph attention encoder called SDGAT that can grasp more feature information without priori neighbourhood information. With these improvements, the proposed method outperforms former RL method in both synthetic and benchmark datasets in terms of output results and optimization robustness.
Code (0)
등록된 구현이 없습니다.
Tasks
Causal DiscoveryGraph AttentionReinforcement Learning (RL)Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Intervention-based Recurrent Casual Model for Non-stationary Video Causal Discovery
Non-stationary casual structures are prevalent in real-world physical systems. For example, the stacked blocks interacted with one another until they fall apart, while the billiard balls are moving independently until th…
Causal DiscoverycounterfactualCounterfactual ReasoningPrevention of Terrorist Crimes in the North Caucasus Region
The relevance of the topic is dictated by the fact that in recent decades, the threat to international security emanating from terrorism has increased many times. Terrorist organizations have become full-fledged subjects…
The Nijmegen Corpus of Casual Czech
This article introduces a new speech corpus, the Nijmegen Corpus of Casual Czech (NCCCz), which contains more than 30 hours of high-quality recordings of casual conversations in Common Czech, among ten groups of three ma…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionATRACT: A Trustworthy Robotic Autonomous system to support Casualty Triage
At a time when drones are increasingly associated with hostile operations, we re-purpose them for humanitarian and life-saving applications. However, adapting search and rescue drones for battlefield triage remains extre…
Action ClassificationData AugmentationPosition Debiasing Fine-Tuning for Causal Perception in Long-Term Dialogue
The core of the dialogue system is to generate relevant, informative, and human-like responses based on extensive dialogue history. Recently, dialogue generation domain has seen mainstream adoption of large language mode…
Dialogue GenerationPositionResponse GenerationSentence