paper-with-me

Papers

Large Scale Product Categorization using Structured and Unstructured Attributes

2019-03-01 · Abhinandan Krishnan, Abilash Amarthaluri

Product categorization using text data for eCommerce is a very challenging extreme classification problem with several thousands of classes and several millions of products to classify. Even though multi-class text classification is a well studied problem both in academia and industry, most approaches either deal with treating product content as a single pile of text, or only consider a few product attributes for modelling purposes. Given the variety of products sold on popular eCommerce platforms, it is hard to consider all available product attributes as part of the modeling exercise, considering that products possess their own unique set of attributes based on category. In this paper, we compare hierarchical models to flat models and show that in specific cases, flat models perform better. We explore two Deep Learning based models that extract features from individual pieces of unstructured data from each product and then combine them to create a product signature. We also propose a novel idea of using structured attributes and their values together in an unstructured fashion along with convolutional filters such that the ordering of the attributes and the differing attributes by product categories no longer becomes a modelling challenge. This approach is also more robust to the presence of faulty product attribute names and values and can elegantly generalize to use both closed list and open list attributes.

📄 PDF Abstract BibTeX arXiv:1903.04254

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeGeneral ClassificationMulti Class Text ClassificationProduct Categorizationtext-classificationText Classification

Similar Papers 제목 키워드 기반

Order from Chaos: Comparative Study of Ten Leading LLMs on Unstructured Data Categorization

2025-10-14 · Ariel Kamen arxiv

This study presents a comparative evaluation of ten state-of-the-art large language models (LLMs) applied to unstructured text categorization using the Interactive Advertising Bureau (IAB) 2.2 hierarchical taxonomy. The …

A Scalable Machine Learning Approach for Inferring Probabilistic US-LI-RADS Categorization

2018-06-15 · Imon Banerjee, Hailey H. Choi, Terry Desser, Daniel L. Rubin

We propose a scalable computerized approach for large-scale inference of Liver Imaging Reporting and Data System (LI-RADS) final assessment categories in narrative ultrasound (US) reports. Although our model was trained …

BIG-bench Machine Learning

LLMForecaster: Improving Seasonal Event Forecasts with Unstructured Textual Data

2024-12-03 · Hanyu Zhang, Chuck Arvin, Dmitry Efimov, Michael W. Mahoney 외

Modern time-series forecasting models often fail to make full use of rich unstructured information about the time series themselves. This lack of proper conditioning can lead to obvious model failures; for example, model…

Demand ForecastingTime SeriesTime Series Forecasting

Learning Mutual Fund Categorization using Natural Language Processing

2022-07-11 · Dimitrios Vamvourellis, Mate Attila Toth, Dhruv Desai, Dhagash Mehta 외

Categorization of mutual funds or Exchange-Traded-funds (ETFs) have long served the financial analysts to perform peer analysis for various purposes starting from competitor analysis, to quantifying portfolio diversifica…

Multi-class Classification

LANISTR: Multimodal Learning from Structured and Unstructured Data

2023-05-26 · Sayna Ebrahimi, Sercan O. Arik, Yihe Dong, Tomas Pfister

Multimodal large-scale pretraining has shown impressive performance for unstructured data such as language and image. However, a prevalent real-world scenario involves structured data types, tabular and time-series, alon…

Time Series