paper-with-me

홈 › Papers

VietJobs: A Vietnamese Job Advertisement Dataset

2026-03-05 · Hieu Pham Dinh, Hung Nguyen Huy, Mo El-Haj arxiv

VietJobs is the first large-scale, publicly available corpus of Vietnamese job advertisements, comprising 48,092 postings and over 15 million words collected from all 34 provinces and municipalities across Vietnam. The dataset provides extensive linguistic and structured information, including job titles, categories, salaries, skills, and employment conditions, covering 16 occupational domains and multiple employment types (full-time, part-time, and internship). Designed to support research in natural language processing and labour market analytics, VietJobs captures substantial linguistic, regional, and socio-economic diversity. We benchmark several generative large language models (LLMs) on two core tasks: job category classification and salary estimation. Instruction-tuned models such as Qwen2.5-7B-Instruct and Llama-SEA-LION-v3-8B-IT demonstrate notable gains under few-shot and fine-tuned settings, while highlighting challenges in multilingual and Vietnamese-specific modelling for structured labour market prediction. VietJobs establishes a new benchmark for Vietnamese NLP and offers a valuable foundation for future research on recruitment language, socio-economic representation, and AI-driven labour market analysis. All code and resources are available at: https://github.com/VinNLP/VietJobs.

📄 PDF Abstract BibTeX arXiv:2603.05262

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Fake Advertisements Detection Using Automated Multimodal Learning: A Case Study for Vietnamese Real Estate Data

2025-01-18 · Duy Nguyen, Trung T. Nguyen, Cuong V. Nguyen

The popularity of e-commerce has given rise to fake advertisements that can expose users to financial and data risks while damaging the reputation of these e-commerce platforms. For these reasons, detecting and removing …

Fake News Detection

Ensemble Learning for Vietnamese Scene Text Spotting in Urban Environments

2024-04-01 · Hieu Nguyen, Cong-Hoang Ta, Phuong-Thuy Le-Nguyen, Minh-Triet Tran 외

This paper presents a simple yet efficient ensemble learning framework for Vietnamese scene text spotting. Leveraging the power of ensemble learning, which combines multiple models to yield more accurate predictions, our…

Ensemble LearningText DetectionText SpottingVietnamese Scene Text

A Vietnamese Dataset for Evaluating Machine Reading Comprehension

2020-09-30 · Kiet Van Nguyen, Duc-Vu Nguyen, Anh Gia-Tuan Nguyen, Ngan Luu-Thuy Nguyen

Over 97 million people speak Vietnamese as their native language in the world. However, there are few research studies on machine reading comprehension (MRC) for Vietnamese, the task of understanding a text and answering…

ArticlesMachine Reading ComprehensionQuestion AnsweringReading Comprehension+3

A Vietnamese Dataset for Evaluating Machine Reading Comprehension

2020-12-01 · COLING 2020 8 · Kiet Nguyen, Vu Nguyen, Anh Nguyen, Ngan Nguyen

Over 97 million inhabitants speak Vietnamese as the native language in the world. However, there are few research studies on machine reading comprehension (MRC) in Vietnamese, the task of understanding a document or text…

ArticlesMachine Reading ComprehensionQuestion AnsweringReading Comprehension+1

KTVIC: A Vietnamese Image Captioning Dataset on the Life Domain

2024-01-16 · Anh-Cuong Pham, Van-Quang Nguyen, Thi-Hong Vuong, Quang-Thuy Ha

Image captioning is a crucial task with applications in a wide range of domains, including healthcare and education. Despite extensive research on English image captioning datasets, the availability of such datasets for …

Image CaptioningVietnamese Image Captioning