paper-with-me

Papers 2k

“2k” 태그가 달린 논문 288편 · 필터 해제

MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization

2025-07-14 · Mingkai Jia, Wei Yin, Xiaotao Hu, Jiaxin Guo 외

Vector Quantized Variational Autoencoders (VQ-VAEs) are fundamental models that compress continuous visual data into discrete tokens. Existing methods have tried to improve the quantization strategy for better reconstruc…

2kImage GenerationImage ReconstructionQuantization

MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization

2025-07-10 · Mingkai Jia, Wei Yin, Xiaotao Hu, Jiaxin Guo 외

Vector Quantized Variational Autoencoders (VQ-VAEs) are fundamental models that compress continuous visual data into discrete tokens. Existing methods have tried to improve the quantization strategy for better reconstruc…

2kQuantization

Understanding and Improving Length Generalization in Recurrent Models

2025-07-03 · Ricardo Buitrago Ruiz, Albert Gu

Recently, recurrent models such as state space models and linear attention have become popular due to their linear complexity in the sequence length. Thanks to their recurrent nature, in principle they can process arbitr…

2kState Space Models

A strengthened bound on the number of states required to characterize maximum parsimony distance

2025-06-11 · Mareike Fischer, Steven Kelk, Sofia Vazquez Alferez

In this article we prove that the distance $d_{\mathrm{MP}}(T_1,T_2) = k$ between two unrooted binary phylogenetic trees $T_1, T_2$ on the same set of taxa can be defined by a character that is convex on one of $T_1, T_2…

2k

Structured Variational $D$-Decomposition for Accurate and Stable Low-Rank Approximation

2025-06-10 · Ronald Katende

We introduce the $D$-decomposition, a non-orthogonal matrix factorization of the form $A \approx P D Q$, where $P \in \mathbb{R}^{n \times k}$, $D \in \mathbb{R}^{k \times k}$, and $Q \in \mathbb{R}^{k \times n}$. The de…

2k

Latent Wavelet Diffusion: Enabling 4K Image Synthesis for Free

2025-05-31 · Luigi Sigillo, Shengfeng He, Danilo Comminiello

High-resolution image synthesis remains a core challenge in generative modeling, particularly in balancing computational efficiency with the preservation of fine-grained visual detail. We present Latent Wavelet Diffusion…

2k4kComputational EfficiencyDenoising+1

Tradeoffs between Mistakes and ERM Oracle Calls in Online and Transductive Online Learning

2025-05-30 · Idan Attias, Steve Hanneke, Arvind Ramaswami

We study online and transductive online learning when the learner interacts with the concept class only via Empirical Risk Minimization (ERM) or weak consistency oracles on arbitrary instance subsets. This contrasts with…

2k

Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models

2025-05-29 · Yiran Guo, Lijie Xu, Jie Liu, Dan Ye 외

Enhancing the reasoning capabilities of large language models effectively using reinforcement learning (RL) remains a crucial challenge. Existing approaches primarily adopt two contrasting advantage estimation granularit…

2k4kGSM8KReinforcement Learning (RL)

Test-Time Training Done Right

2025-05-29 · Tianyuan Zhang, Sai Bi, Yicong Hong, Kai Zhang 외

Test-Time Training (TTT) models context dependencies by adapting part of the model's weights (referred to as fast weights) during inference. This fast weight, akin to recurrent states in RNNs, stores temporary memories o…

2kNovel View Synthesis

MMP-2K: A Benchmark Multi-Labeled Macro Photography Image Quality Assessment Database

2025-05-25 · Jiashuo Chang, Zhengyi Li, Jianxun Lou, Zhen Qiu 외

Macro photography (MP) is a specialized field of photography that captures objects at an extremely close range, revealing tiny details. Although an accurate macro photography image quality assessment (MPIQA) metric can b…

2kDiversityImage Quality Assessment

Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions

2025-05-23 · Olivier Toubia, George Z. Gui, Tianyi Peng, Daniel J. Merlau 외

LLM-based digital twin simulation, where large language models are used to emulate individual human behavior, holds great promise for research in AI, social science, and digital experimentation. However, progress in this…

2kBenchmarking

PIIvot: A Lightweight NLP Anonymization Framework for Question-Anchored Tutoring Dialogues

2025-05-22 · Matthew Zent, Digory Smith, Simon Woodhead

Personally identifiable information (PII) anonymization is a high-stakes task that poses a barrier to many open-science data sharing initiatives. While PII identification has made large strides in recent years, in practi…

2k

Unlocking the Potential of Difficulty Prior in RL-based Multimodal Reasoning

2025-05-19 · Mingrui Chen, Haogeng Liu, Hao Liang, Huaibo Huang 외

In this work, we investigate how explicitly modeling problem's difficulty prior information shapes the effectiveness of reinforcement learning based fine-tuning for multimodal reasoning. Our exploration mainly comprises …

2kMathematical ReasoningMultimodal Reasoning

UIShift: Enhancing VLM-based GUI Agents through Self-supervised Reinforcement Learning

2025-05-18 · Longxi Gao, Li Zhang, Mengwei Xu

Training effective Vision Language Models (VLMs) for GUI agents typically relies on supervised fine-tuning (SFT) over large-scale annotated datasets, where the collection process is labor-intensive and error-prone. In th…

2kReinforcement Learning (RL)

ViMRHP: A Vietnamese Benchmark Dataset for Multimodal Review Helpfulness Prediction via Human-AI Collaborative Annotation

2025-05-12 · Truc Mai-Thanh Nguyen, Dat Minh Nguyen, Son T. Luu, Kiet Van Nguyen

Multimodal Review Helpfulness Prediction (MRHP) is an essential task in recommender systems, particularly in E-commerce platforms. Determining the helpfulness of user-generated reviews enhances user experience and improv…

2kRecommendation Systems

Calibrating Translation Decoding with Quality Estimation on LLMs

2025-04-26 · Di wu, Yibin Lei, Christof Monz

Neural machine translation (NMT) systems typically employ maximum a posteriori (MAP) decoding to select the highest-scoring translation from the distribution mass. However, recent evidence highlights the inadequacy of MA…

2kMachine TranslationNMTTranslation

aiXamine: Simplified LLM Safety and Security

2025-04-21 · Fatih Deniz, Dorde Popovic, Yazan Boshmaf, Euisuh Jeong 외

Evaluating Large Language Models (LLMs) for safety and security remains a complex task, often requiring users to navigate a fragmented landscape of ad hoc benchmarks, datasets, metrics, and reporting formats. To address …

2kAdversarial RobustnessFairnessHallucination+2

Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis

2025-04-20 · Jingjing Ren, Wenbo Li, Zhongdao Wang, Haoze Sun 외

Demand for 2K video synthesis is rising with increasing consumer expectations for ultra-clear visuals. While diffusion transformers (DiTs) have demonstrated remarkable capabilities in high-quality video generation, scali…

2kKnowledge DistillationTransfer LearningVideo Generation

On Linear Representations and Pretraining Data Frequency in Language Models

2025-04-16 · Jack Merullo, Noah A. Smith, Sarah Wiegreffe, Yanai Elazar

Pretraining data has a direct impact on the behaviors and quality of language models (LMs), but we only understand the most basic principles of this relationship. While most work focuses on pretraining data's effect on d…

2kIn-Context LearningRelation

Rethinking the Generation of High-Quality CoT Data from the Perspective of LLM-Adaptive Question Difficulty Grading

2025-04-16 · Qianjin Yu, Keyu Wu, Zihan Chen, Chushu Zhang 외

Recently, DeepSeek-R1 (671B) (DeepSeek-AIet al., 2025) has demonstrated its excellent reasoning ability in complex tasks and has publiclyshared its methodology. This provides potentially high-quality chain-of-thought (Co…

2kCode GenerationMath
1–20 / 288 다음 →