paper-with-me

Papers

Automatic Product Categorization for Official Statistics

2019-08-01 · WS 2019 8 · Andrea Roberson

The North American Product Classification System (NAPCS) is a comprehensive, hierarchical classification system for products (goods and services) that is consistent across the three North American countries. Beginning in 2017, the Economic Census will use NAPCS to produce economy-wide product tabulations. Respondents are asked to report data from a long, pre-specified list of potential products in a given industry, with some lists containing more than 50 potential products. Businesses have expressed the desire to alternatively supply Universal Product Codes (UPC) to the U. S. Census Bureau. Much work has been done around the categorization of products using product descriptions. No study has applied these efforts for the calculation of official statistics (statistics published by government agencies) using only the text of UPC product descriptions. The question we address in this paper is: Given UPC codes and their associated product descriptions, can we accurately predict NAPCS? We tested the feasibility of businesses submitting a spreadsheet with Universal Product Codes and their associated text descriptions. This novel strategy classified text with very high accuracy rates, all of our algorithms surpassed over 90 percent.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

General ClassificationProduct Categorization

Methods 이 논문이 사용한 방법론

American 설명 없음

Similar Papers 제목 키워드 기반

Leveraging Machine Learning for Official Statistics: A Statistical Manifesto

2024-09-06 · Marco Puts, David Salgado, Piet Daas

It is important for official statistics production to apply ML with statistical rigor, as it presents both opportunities and challenges. Although machine learning has enjoyed rapid technological advances in recent years,…

Surveyvalid

Changing Data Sources in the Age of Machine Learning for Official Statistics

2023-06-07 · Cedric De Boom, Michael Reusens

Data science has become increasingly essential for the production of official statistics, as it enables the automated collection, processing, and analysis of large amounts of data. With such data science practices in pla…

Decision MakingEthics

Connecting Algorithmic Fairness to Quality Dimensions in Machine Learning in Official Statistics and Survey Production

2024-02-14 · Patrick Oliver Schenk, Christoph Kern

National Statistical Organizations (NSOs) increasingly draw on Machine Learning (ML) to improve the timeliness and cost-effectiveness of their products. When introducing ML solutions, NSOs must ensure that high standards…

Fairness

The Applicability of Federated Learning to Official Statistics

2023-07-28 · Joshua Stock, Oliver Hauke, Julius Weißmann, Hannes Federrath

This work investigates the potential of Federated Learning (FL) for official statistics and shows how well the performance of FL models can keep up with centralized learning methods.F L is particularly interesting for of…

Federated Learning

Large-Scale Categorization of Japanese Product Titles Using Neural Attention Models

2017-04-01 · EACL 2017 4 · Y Xia, i, Aaron Levine, Pradipto Das 외

We propose a variant of Convolutional Neural Network (CNN) models, the Attention CNN (ACNN); for large-scale categorization of millions of Japanese items into thirty-five product categories. Compared to a state-of-the-ar…

Feature EngineeringProduct CategorizationText Categorization