paper-with-me

홈 › Papers

Inference Time Alignment with Reward-Guided Tree Search

2024-06-21 · Chia-Yu Hung, Navonil Majumder, Ambuj Mehrish, Soujanya Poria

Inference-time computation methods enhance the performance of Large Language Models (LLMs) by leveraging additional computational resources to achieve superior results. Common techniques, such as Best-of-N sampling, Majority Voting, and variants of tree-search algorithms have proven to be effective in boosting the performance of LLMs. These approaches strategically trade increased computational resources for improved model responses. In this work, we proposed DARWIN, an inference-time alignment method that leverages the guidance of a reward model to achieve alignment through a reward-guided tree search. Empirical evidences indicates that our method outperforms other inference-time alignment methods such as Best-of-N and ARGS on two widely accepted alignment benchmarks AlpacaEval 2 and MT-Bench. Furthermore, we show that our inference-time approach achieves performance comparable to preference-tuned models on both benchmarks, highlighting the effectiveness of trading inference-time compute for enhanced performance during inference. We have released our codes at https://github.com/declare-lab/darwin.

📄 PDF Abstract BibTeX arXiv:2406.15193

Code (1)

declare-lab/darwin 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Gradient-Guided Reward Optimization for Inference-time Alignment

2026-06-08 · Hankun Lin, Ruqi Zhang arxiv

Ensuring the reliability of Large Language Models (LLMs) under distribution drift requires inference-time adaptation. While inference-time alignment methods such as Best-of-$N$ and rejection sampling are widely used, the…

Diffusion Tree Sampling: Scalable inference-time alignment of diffusion models

2025-06-25 · Vineet Jain, Kusha Sareen, Mohammad Pedramfar, Siamak Ravanbakhsh

Adapting a pretrained diffusion model to new objectives at inference time remains an open problem in generative modeling. Existing steering methods suffer from inaccurate value estimation, especially at high noise levels…

Image GenerationText to Image GenerationText-to-Image Generation

Safe Inference-Time Alignment via Lagrangian Reward Augmentation

2026-07-02 · Yaswanth Chittepu, Ativ Joshi, Sohini Chintala, Scott Niekum arxiv

Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weight updates. However, existing inference-time alignment methods typically optimize …

Inference-Time Alignment in Diffusion Models with Reward-Guided Generation: Tutorial and Review

2025-01-16 · Masatoshi Uehara, Yulai Zhao, Chenyu Wang, Xiner Li 외

This tutorial provides an in-depth guide on inference-time guidance and alignment methods for optimizing downstream reward functions in diffusion models. While diffusion models are renowned for their generative modeling …

DenoisingProtein Design

InfAlign: Inference-aware language model alignment

2024-12-27 · Ananth Balashankar, Ziteng Sun, Jonathan Berant, Jacob Eisenstein 외

Language model alignment is a critical step in training modern generative language models. Alignment targets to improve win rate of a sample from the aligned model against the base model. Today, we are increasingly using…

Language ModelingLanguage ModellingmodelModels Alignment