paper-with-me

홈 › Papers

Unveiling the Secrets: How Masking Strategies Shape Time Series Imputation

2024-05-26 · Linglong Qian, Yiyuan Yang, Wenjie Du, Jun Wang, Zina Ibrahim

Time series imputation is a critical challenge in data mining, particularly in domains like healthcare and environmental monitoring, where missing data can compromise analytical outcomes. This study investigates the influence of diverse masking strategies, normalization timing, and missingness patterns on the performance of eleven state-of-the-art imputation models across three diverse datasets. Specifically, we evaluate the effects of pre-masking versus in-mini-batch masking, augmentation versus overlaying of artificial missingness, and pre-normalization versus post-normalization. Our findings reveal that masking strategies profoundly affect imputation accuracy, with dynamic masking providing robust augmentation benefits and overlay masking better simulating real-world missingness patterns. Sophisticated models, such as CSDI, exhibited sensitivity to preprocessing configurations, while simpler models like BRITS delivered consistent and efficient performance. We highlight the importance of aligning preprocessing pipelines and masking strategies with dataset characteristics to improve robustness under diverse conditions, including high missing rates. This study provides actionable insights for designing imputation pipelines and underscores the need for transparent and comprehensive experimental designs.

📄 PDF Abstract BibTeX arXiv:2405.17508

Code (1)

LinglongQian/ExperimentalDesignAnalysis 공식 구현 pytorch

Tasks

ImputationTime Series

Similar Papers 제목 키워드 기반

Towards Improved Input Masking for Convolutional Neural Networks

2022-11-26 · ICCV 2023 1 · Sriram Balasubramanian, Soheil Feizi

The ability to remove features from the input of machine learning models is very important to understand and interpret model predictions. However, this is non-trivial for vision models since masking out parts of the inpu…

Data Augmentation

Understanding Secret Leakage Risks in Code LLMs: A Tokenization Perspective

2026-04-20 · Meifang Chen, Zhe Yang, Huang Nianchen, Yizhan Huang 외 arxiv

Code secrets are sensitive assets for software developers, and their leakage poses significant cybersecurity risks. While the rapid development of AI code assistants powered by Code Large Language Models (CLLMs), CLLMs a…

Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets

2026-07-09 · Jack Hopkins, Dipika Khullar, Fabien Roger arxiv

Black box auditing of language models is an essential pre-deployment tool, but it may miss subtle forms of misalignment and hidden information. To better elicit hidden information during an auditing process, we introduce…

RedacBench: Can AI Erase Your Secrets?

2026-03-02 · Hyunjun Jeon, Kyuyoung Kim, Jinwoo Shin arxiv

Modern language models can readily extract sensitive information from unstructured text, making redaction -- the selective removal of such information -- critical for data security. However, existing benchmarks for redac…

Unveiling the Secrets of Engaging Conversations: Factors that Keep Users Hooked on Role-Playing Dialog Agents

2024-02-18 · Shuai Zhang, Yu Lu, Junwen Liu, JIA YU 외

With the growing humanlike nature of dialog agents, people are now engaging in extended conversations that can stretch from brief moments to substantial periods of time. Understanding the factors that contribute to susta…