paper-with-me

Papers

ATLAS: Benchmarking and Adapting LLMs for Global Trade via Harmonized Tariff Code Classification

2025-09-22 · Pritish Yuvraj, Siva Devarakonda arxiv

Accurate classification of products under the Harmonized Tariff Schedule (HTS) is a critical bottleneck in global trade, yet it has received little attention from the machine learning community. Misclassification can halt shipments entirely, with major postal operators suspending deliveries to the U.S. due to incomplete customs documentation. We introduce the first benchmark for HTS code classification, derived from the U.S. Customs Rulings Online Search System (CROSS). Evaluating leading LLMs, we find that our fine-tuned Atlas model (LLaMA-3.3-70B) achieves 40 percent fully correct 10-digit classifications and 57.5 percent correct 6-digit classifications, improvements of 15 points over GPT-5-Thinking and 27.5 points over Gemini-2.5-Pro-Thinking. Beyond accuracy, Atlas is roughly five times cheaper than GPT-5-Thinking and eight times cheaper than Gemini-2.5-Pro-Thinking, and can be self-hosted to guarantee data privacy in high-stakes trade and compliance workflows. While Atlas sets a strong baseline, the benchmark remains highly challenging, with only 40 percent 10-digit accuracy. By releasing both dataset and model, we aim to position HTS classification as a new community benchmark task and invite future work in retrieval, reasoning, and alignment.

📄 PDF Abstract BibTeX arXiv:2509.18400

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM

2025-10-20 · Haoyu Huang, Hong Ting Tsang, Jiaxin Bai, Xi Peng 외 arxiv

Retrieval-augmented generation (RAG) has shown some success in augmenting large language models (LLMs) with external knowledge. However, as a non-parametric knowledge integration paradigm for LLMs, RAG methods heavily re…

Knowledge Graphs

From Knowledge to Inference: Formalizing Specialized Public Health Reasoning on GlobalHealthAtlas

2026-01-31 · Zhaokun Yan, Shan Xu, Wuzheng Dong, Zhaohan Liu 외 arxiv

Public health reasoning requires population level inference grounded in scientific evidence, expert consensus, and safety constraints. However, it remains underexplored as a structured machine learning problem with limit…

Atlas-Chat: Adapting Large Language Models for Low-Resource Moroccan Arabic Dialect

2024-09-26 · Guokan Shang, Hadi Abdine, Yousef Khoubrane, Amr Mohamed 외

We introduce Atlas-Chat, the first-ever collection of LLMs specifically developed for dialectal Arabic. Focusing on Moroccan Arabic, also known as Darija, we construct our instruction dataset by consolidating existing Da…

Human Behavior Atlas: Benchmarking Unified Psychological and Social Behavior Understanding

2025-10-06 · Keane Ong, Wei Dai, Carol Li, Dewei Feng 외 arxiv

Using intelligent systems to perceive psychological and social behaviors, that is, the underlying affective, cognitive, and pathological states that are manifested through observable behaviors and social interactions, re…

On the Empirical Association between Trade Network Complexity and Global Gross Domestic Product

2022-11-18 · Mayank Kejriwal, Yuesheng Luo

In recent decades, trade between nations has constituted an important component of global Gross Domestic Product (GDP), with official estimates showing that it likely accounted for a quarter of total global production. W…