paper-with-me

홈 › Papers

RankPO: Preference Optimization for Job-Talent Matching

2025-03-13 · Yafei Zhang, Murray Wang, Yu Wang, Xiaohui Wang

Matching job descriptions (JDs) with suitable talent requires models capable of understanding not only textual similarities between JDs and candidate resumes but also contextual factors such as geographical location and academic seniority. To address this challenge, we propose a two-stage training framework for large language models (LLMs). In the first stage, a contrastive learning approach is used to train the model on a dataset constructed from real-world matching rules, such as geographical alignment and research area overlap. While effective, this model primarily learns patterns that defined by the matching rules. In the second stage, we introduce a novel preference-based fine-tuning method inspired by Direct Preference Optimization (DPO), termed Rank Preference Optimization (RankPO), to align the model with AI-curated pairwise preferences emphasizing textual understanding. Our experiments show that while the first-stage model achieves strong performance on rule-based data (nDCG@20 = 0.706), it lacks robust textual understanding (alignment with AI annotations = 0.46). By fine-tuning with RankPO, we achieve a balanced model that retains relatively good performance in the original tasks while significantly improving the alignment with AI preferences. The code and data are available at https://github.com/yflyzhang/RankPO.

📄 PDF Abstract BibTeX arXiv:2503.10723

Code (1)

yflyzhang/rankpo 공식 구현 pytorch

Tasks

Contrastive Learning

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models

2023-11-16 · Jiongxiao Wang, Junlin Wu, Muhao Chen, Yevgeniy Vorobeychik 외

Reinforcement Learning with Human Feedback (RLHF) is a methodology designed to align Large Language Models (LLMs) with human preferences, playing an important role in LLMs alignment. Despite its advantages, RLHF relies o…

Backdoor AttackData PoisoningRed TeamingSafety Alignment

TalentCLEF at CLEF2026: Skill and Job Title Intelligence for Human Capital Management

2026-07-22 · Luis Gasco, Hermenegildo Fabregat, Laura García-Sardiña, Paula Estrella 외 arxiv

This paper presents the second edition of the TalentCLEF Challenge, which will run as an evaluation lab as part of CLEF 2026. The aim of TalentCLEF is to promote the development of systems and methods that use Natural La…

Enhancing Talent Search Ranking with Role-Aware Expert Mixtures and LLM-based Fine-Grained Job Descriptions

2025-10-05 · Jihang Li, Bing Xu, Zulong Chen, Chuanfei Xu 외 arxiv

Talent search is a cornerstone of modern recruitment systems, yet existing approaches often struggle to capture nuanced job-specific preferences, model recruiter behavior at a fine-grained level, and mitigate noise from …

Multi-Task Learning

Person-Job Fit: Adapting the Right Talent for the Right Job with Joint Representation Learning

2018-10-08 · Chen Zhu, HengShu Zhu, Hui Xiong, Chao Ma 외

Person-Job Fit is the process of matching the right talent for the right job by identifying talent competencies that are required for the job. While many qualitative efforts have been made in related fields, it still lac…

Data VisualizationRepresentation Learning

Entity Personalized Talent Search Models with Tree Interaction Features

2019-02-25 · Cagri Ozcaglar, Sahin Geyik, Brian Schmitz, Prakhar Sharma 외

Talent Search systems aim to recommend potential candidates who are a good match to the hiring needs of a recruiter expressed in terms of the recruiter's search query or job posting. Past work in this domain has focused …

Benchmarking