paper-with-me

홈 › Papers

Multilingual hierarchical classification of job advertisements for job vacancy statistics

2024-11-06 · Maciej Beręsewicz, Marek Wydmuch, Herman Cherniaiev, Robert Pater

The goal of this paper is to develop a multilingual classifier and conditional probability estimator of occupation codes for online job advertisements according in accordance with the International Standard Classification of Occupations (ISCO) extended with the Polish Classification of Occupations and Specializations (KZiS), which is analogous to the European Classification of Occupations. In this paper, we utilise a range of data sources, including a novel one, namely the Central Job Offers Database, which is a register of all vacancies submitted to Public Employment Offices. Their staff members code the vacancies according to the ISCO and KZiS. A hierarchical multi-class classifier has been developed based on the transformer architecture. The classifier begins by encoding the jobs found in advertisements to the widest 1-digit occupational group, and then narrows the assignment to a 6-digit occupation code. We show that incorporation of the hierarchical structure of occupations improves prediction accuracy by 1-2 percentage points, particularly for the hand-coded online job advertisements. Finally, a bilingual (Polish and English) and multilingual (24 languages) model is developed based on data translated using closed and open-source software. The open-source software is provided for the benefit of the official statistics community, with a particular focus on international comparability.

📄 PDF Abstract BibTeX arXiv:2411.03779

Code (1)

OJALAB/CBOP-datasets 공식 구현

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Serif or Sans: Visual Font Analytics on Book Covers and Online Advertisements

2019-06-24 · Yuto Shinahara, Takuro Karamatsu, Daisuke Harada, Kota Yamaguchi 외

In this paper, we conduct a large-scale study of font statistics in book covers and online advertisements. Through the statistical study, we try to understand how graphic designers relate fonts and content genres and ide…

Clustering

Learning to Match Job Candidates Using Multilingual Bi-Encoder BERT

2021-09-15 · Dor Lavi

In this talk, we will show how we used Randstad history of candidate placements to generate labeled CV-vacancy pairs dataset. Afterwards we fine-tune a multilingual BERT with bi encoder structure over this dataset, by ad…

Text Zoning and Classification for Job Advertisements in German, French and English

2020-11-01 · EMNLP (NLP+CSS) 2020 11 · Ann-Sophie Gnehm, Simon Clematide

We present experiments to structure job ads into text zones and classify them into pro- fessions, industries and management functions, thereby facilitating social science analyses on labor marked demand. Our main contrib…

Management

Multilingual Hierarchical Attention Networks for Document Classification

2017-07-04 · IJCNLP 2017 11 · Nikolaos Pappas, Andrei Popescu-Belis

Hierarchical attention networks have recently achieved remarkable performance for document classification in a given language. However, when multilingual document collections are considered, training such models separate…

ClassificationComputational EfficiencyDocument ClassificationGeneral Classification+2

Human + AI for Accelerating Ad Localization Evaluation

2025-09-16 · Harshit Rajgarhia, Shivali Dalmia, Mengyang Zhao, Mukherji Abhishek 외 arxiv

Adapting advertisements for multilingual audiences requires more than simple text translation; it demands preservation of visual consistency, spatial alignment, and stylistic integrity across diverse languages and format…

Scene Text DetectionMachine Translation