paper-with-me

홈 › Papers

Jointly Optimizing State Operation Prediction and Value Generation for Dialogue State Tracking

2020-10-24 · Yan Zeng, Jian-Yun Nie

We investigate the problem of multi-domain Dialogue State Tracking (DST) with open vocabulary. Existing approaches exploit BERT encoder and copy-based RNN decoder, where the encoder predicts the state operation, and the decoder generates new slot values. However, in such a stacked encoder-decoder structure, the operation prediction objective only affects the BERT encoder and the value generation objective mainly affects the RNN decoder. In this paper, we propose a purely Transformer-based framework, where a single BERT works as both the encoder and the decoder. In so doing, the operation prediction objective and the value generation objective can jointly optimize this BERT for DST. At the decoding step, we re-use the hidden states of the encoder in the self-attention mechanism of the corresponding decoder layers to construct a flat encoder-decoder architecture for effective parameter updating. Experimental results show that our approach substantially outperforms the existing state-of-the-art framework, and it also achieves very competitive performance to the best ontology-based approaches.

📄 PDF Abstract BibTeX arXiv:2010.14061

Code (2)

zengyan-97/Transformer-DST 공식 구현 pytorch
bcaitech1/p3-dst-chatting-day pytorch

Tasks

DecoderDialogue State TrackingMulti-domain Dialogue State Tracking

Methods 이 논문이 사용한 방법론

DST Dynamic sparse training methods train neural networks in a sparse manner, starting with an initial sparse mask, and periodically updating the mask based on some criteria.
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Adam 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.

Similar Papers 제목 키워드 기반

Multi-Agent Reinforcement Learning for Joint Police Patrol and Dispatch

2024-09-03 · Matthew Repasky, He Wang, Yao Xie

Police patrol units need to split their time between performing preventive patrol and being dispatched to serve emergency incidents. In the existing literature, patrol and dispatch decisions are often studied separately.…

Multi-agent Reinforcement Learningreinforcement-learningReinforcement Learning

Q-Delta: Beyond Key-Value Associative State Evolution

2026-06-07 · Sumin Park, Seojin Kim, Noseong Park arxiv

Linear attention reformulates sequence modeling as recurrent state evolution, enabling efficient linear-time inference. Under the key-value associative paradigm, existing approaches restrict the role of the query to the …

Value prediction

Environment Parameter Gradient Theorem for Policy-Environment Co-Design in Reinforcement Learning

2026-07-14 · Amber Srivastava arxiv

Reinforcement learning (RL) is traditionally concerned with learning a control policy for a fixed environment. In many engineering systems, however, the environment itself is alterable: physical or operational parameters…

Reinforcement Learning

Decision-Value Attribution in Predict-then-Optimize Systems

2026-06-29 · Konstantinos Ziliaskopoulos, Alexander Vinel, Alice E. Smith arxiv

Predictive models are increasingly embedded in operational decision-making, yet standard explanation methods typically explain forecasts rather than the decisions those forecasts induce. This distinction is important in …

Supervised PCA: A Multiobjective Approach

2020-11-10 · Alexander Ritchie, Laura Balzano, Daniel Kessler, Chandra S. Sripada 외

Methods for supervised principal component analysis (SPCA) aim to incorporate label information into principal component analysis (PCA), so that the extracted features are more useful for a prediction task of interest. P…

Prediction