paper-with-me

Papers

ROME: Robustifying Memory-Efficient NAS via Topology Disentanglement and Gradient Accumulation

2020-11-23 · ICCV 2023 1 · Xiaoxing Wang, Xiangxiang Chu, Yuda Fan, Zhexi Zhang, Bo Zhang, Xiaokang Yang, Junchi Yan

Albeit being a prevalent architecture searching approach, differentiable architecture search (DARTS) is largely hindered by its substantial memory cost since the entire supernet resides in the memory. This is where the single-path DARTS comes in, which only chooses a single-path submodel at each step. While being memory-friendly, it also comes with low computational costs. Nonetheless, we discover a critical issue of single-path DARTS that has not been primarily noticed. Namely, it also suffers from severe performance collapse since too many parameter-free operations like skip connections are derived, just like DARTS does. In this paper, we propose a new algorithm called RObustifying Memory-Efficient NAS (ROME) to give a cure. First, we disentangle the topology search from the operation search to make searching and evaluation consistent. We then adopt Gumbel-Top2 reparameterization and gradient accumulation to robustify the unwieldy bi-level optimization. We verify ROME extensively across 15 benchmarks to demonstrate its effectiveness and robustness.

📄 PDF Abstract BibTeX arXiv:2011.11233

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementNeural Architecture Search

Similar Papers 제목 키워드 기반

Robust Training in High Dimensions via Block Coordinate Geometric Median Descent

2021-06-16 · Anish Acharya, Abolfazl Hashemi, Prateek Jain, Sujay Sanghavi 외

Geometric median (\textsc{Gm}) is a classical method in statistics for achieving a robust estimation of the uncorrupted data; under gross corruption, it achieves the optimal breakdown point of 0.5. However, its computati…

Image ClassificationVocal Bursts Intensity Prediction

Robustifying Markowitz

2022-12-28 · Wolfgang Karl Härdle, Yegor Klochkov, Alla Petukhina, Nikita Zhivotovskiy

Markowitz mean-variance portfolios with sample mean and covariance as input parameters feature numerous issues in practice. They perform poorly out of sample due to estimation error, they experience extreme weights toget…

Time SeriesTime Series Analysis

Latent Disentanglement in Mesh Variational Autoencoders Improves the Diagnosis of Craniofacial Syndromes and Aids Surgical Planning

2023-09-05 · Simone Foti, Alexander J. Rickart, Bongjin Koo, Eimear O' Sullivan 외

The use of deep learning to undertake shape analysis of the complexities of the human head holds great promise. However, there have traditionally been a number of barriers to accurate modelling, especially when operating…

Disentanglement

TimeROME-DLM: Temporal Causal Tracing and Low-Rank Inference-Time Knowledge Editing for Masked Diffusion Language Models

2026-06-11 · Zhengtao Yao, Liuyang Song, Hongbo Zhang, Chenhao Wei 외 arxiv

Masked diffusion language models (MDLMs) such as LLaDA now rival autoregressive (AR) LLMs, but every existing knowledge-editing and unlearning method (ROME, MEMIT, etc.) targets AR transformers and either makes assumptio…

knowledge editing

Evaluating the Disentanglement of Deep Generative Models through Manifold Topology

2020-06-05 · ICLR 2021 1 · Sharon Zhou, Eric Zelikman, Fred Lu, Andrew Y. Ng 외

Learning disentangled representations is regarded as a fundamental task for improving the generalization, robustness, and interpretability of generative models. However, measuring disentanglement has been challenging and…

Disentanglement