paper-with-me

Papers Attribute Extraction

“Attribute Extraction” 태그가 달린 논문 82편 · 필터 해제

Scaling E-Commerce Attribute Extraction with Parallel Decoding

2026-09-09 · Nikhita Vedula, Dushyanta Dhyani, Bryan Wang, Shervin Malmasi arxiv

Customers rely on specific product attributes to compare products and make purchasing decisions, but e-commerce catalogs are messy and unstructured, making it difficult to identify which attributes matter most and extrac…

Attribute Value ExtractionAttribute Extraction

How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA

2026-07-17 · Navya Gupta, Bingjie Xu, Avinash Anand, Timothy Liu 외 arxiv

Compositional visual question answering requires Vision-Language Models (VLMs) to execute multiple reasoning operations like object selection, spatial relation resolution, and attribute verification. Despite strong aggre…

Visual Question AnsweringAttribute ExtractionVisual Grounding

SynthAVE: Scalable Synthetic Labeling for E-Commerce with LLM-Arena Validation

2026-07-08 · Andrea Scarinci, Virginia Negri, Brayan Impata, Suleiman Khan 외 arxiv

Fine-tuning large language models (LLMs) for e-commerce attribute extraction requires labeled data representative across thousands of product types, attributes, and multiple languages. This combinatorial scale translates…

Attribute Value ExtractionAttribute Extraction

CheXpercept: A Benchmark for Evaluating Expert-Level Lesion Perception in Chest X-rays

2026-06-19 · Geon Choi, Hangyul Yoon, Nalee Kim, Jeong Yun Jang 외 arxiv

The evaluation of vision-language models (VLMs) for chest X-ray (CXR) analysis has largely been limited to disease-presence classification without visual grounding. Such evaluations fail to verify the expert-level lesion…

Attribute ExtractionDomain AdaptationVisual Grounding

IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products

2026-06-12 · Haonan Qi, Jin Cao, Yongqi Zhang, Xintong Wang 외 arxiv

Industrial products such as valves and circuit breakers are defined by dense technical specifications that govern procurement, compatibility, and safety across supply chains. These specifications are scattered across mul…

Attribute Value ExtractionAttribute ExtractionVisual Reasoning

AREA: Attribute Extraction and Aggregation for CLIP-Based Class-Incremental Learning

2026-05-27 · Zhen-Hao Xie, Yu-Cheng Shi, Da-Wei Zhou arxiv

Class-Incremental Learning (CIL) is important in building real-world learning systems. In CLIP-based CIL, the model performs classification by comparing similarity between visual and textual embeddings obtained from temp…

class-incremental learningAttribute Extraction

Experiments in Agentic AI for Science

2026-05-25 · Judy Fox, Geoffrey Fox arxiv

This paper details two novel frameworks for developing autonomous, agentic AI in scientific workflows. Both systems leverage a hybrid Local Body, Remote Brain architecture via Google Colab, utilizing Python-based local o…

Attribute ExtractionKnowledge Graphs

Fashion Florence: Fine-Tuning Florence-2 for Structured Fashion Attribute Extraction

2026-05-11 · Anushree Berlia arxiv

We present Fashion Florence, a Florence-2 vision-language model fine-tuned with LoRA to extract structured fashion attributes from clothing images. Given a single photograph, the model generates a JSON object containing …

Attribute Extraction

Breaking the Autoregressive Chain: Hyper-Parallel Decoding for Efficient LLM-Based Attribute Value Extraction

2026-04-29 · Theodore Glavas, Nikhita Vedula, Dushyanta Dhyani, Yilun Zhu 외 arxiv

Some text generation tasks, such as Attribute Value Extraction (AVE), require decoding multiple independent sequences from the same document context. While standard autoregressive decoding is slow due to its sequential n…

Attribute Value ExtractionAttribute ExtractionText Generation

AutoPKG: An Automated Framework for Dynamic E-commerce Product-Attribute Knowledge Graph Construction

2026-04-18 · Pollawat Hongwimol, Haoning Shang, Chutong Wang, Zhichao Wan 외 arxiv

Product attribute extraction in e-commerce is bottlenecked by ontologies that are inconsistent, incomplete, and costly to maintain. We present AutoPKG, a multi-agent Large Language Model (LLM) framework that automaticall…

Attribute Extraction

Vision-Language Attribute Disentanglement and Reinforcement for Lifelong Person Re-Identification

2026-03-20 · Kunlun Xu, Haotong Cheng, Jiangmeng Li, Xu Zou 외 arxiv

Lifelong person re-identification (LReID) aims to learn from varying domains to obtain a unified person retrieval model. Existing LReID approaches typically focus on learning from scratch or a visual classification-pretr…

Person Re-IdentificationAttribute ExtractionPerson Retrieval

Logics-Parsing-Omni Technical Report

2026-03-10 · Xin An, Jingyi Cai, Xiangyang Chen, Huayao Liu 외 arxiv

Addressing the challenges of fragmented task definitions and the heterogeneity of unstructured data in multimodal parsing, this paper proposes the Omni Parsing framework. This framework establishes a Unified Taxonomy cov…

Attribute Extraction

Adapting Vision-Language Models for E-commerce Understanding at Scale

2026-02-12 · Matteo Nulli, Vladimir Orshulevich, Tala Bazazo, Christian Herold 외 arxiv

E-commerce product understanding demands by nature, strong multimodal comprehension from text, images, and structured attributes. General-purpose Vision-Language Models (VLMs) enable generalizable multimodal latent model…

Instruction FollowingAttribute Extraction

Dolphin-v2: Universal Document Parsing via Scalable Anchor Prompting

2026-02-05 · Hao Feng, Wei Shi, Ke Zhang, Xiang Fei 외 arxiv

Document parsing has garnered widespread attention as vision-language models (VLMs) advance OCR capabilities. However, the field remains fragmented across dozens of specialized models with varying strengths, forcing user…

Attribute Extraction

Neural Sentinel: Unified Vision Language Model (VLM) for License Plate Recognition with Human-in-the-Loop Continual Learning

2026-02-04 · Karthik Sivakoti arxiv

Traditional Automatic License Plate Recognition (ALPR) systems employ multi-stage pipelines consisting of object detection networks followed by separate Optical Character Recognition (OCR) modules, introducing compoundin…

License Plate RecognitionZero-shot GeneralizationAttribute ExtractionContinual Learning

GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation

2025-12-30 · Yuan Feng, Yue Yang, Xiaohan He, Jiatong Zhao 외 arxiv

Geometric problem solving constitutes a critical branch of mathematical reasoning, requiring precise analysis of shapes and spatial relationships. Current evaluations of geometric reasoning in vision-language models (VLM…

Mathematical ReasoningAttribute Extraction

Beyond Memorization: Selective Learning for Copyright-Safe Diffusion Model Training

2025-12-12 · Divya Kothandaraman, Jaclyn Pytlarz arxiv

Memorization in large-scale text-to-image diffusion models poses significant security and intellectual property risks, enabling adversarial attribute extraction and the unauthorized reproduction of sensitive or proprieta…

Attribute Extraction

Customizing Open Source LLMs for Quantitative Medication Attribute Extraction across Heterogeneous EHR Systems

2025-10-23 · Zhe Fei, Mehmet Yigit Turali, Shreyas Rajesh, Xinyang Dai 외 arxiv

Harmonizing medication data across Electronic Health Record (EHR) systems is a persistent barrier to monitoring medications for opioid use disorder (MOUD). In heterogeneous EHR systems, key prescription attributes are sc…

Attribute Extraction

LLM-RG: Referential Grounding in Outdoor Scenarios using Large Language Models

2025-09-29 · Pranav Saxena, Avigyan Bhattacharya, Ji Zhang, Wenshan Wang arxiv

Referential grounding in outdoor driving scenes is challenging due to large scene variability, many visually similar objects, and dynamic elements that complicate resolving natural-language references (e.g., "the black c…

Attribute ExtractionReferring Expression

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis

2025-09-11 · Jing Hao, Yuxuan Fan, Yanpeng Sun, Kaixin Guo 외 arxiv

Recent advances in large vision-language models (LVLMs) have demonstrated strong performance on general-purpose medical tasks. However, their effectiveness in specialized domains such as dentistry remains underexplored. …

Visual Question AnsweringAttribute Extraction
1–20 / 82 다음 →