Papers Attribute Extraction
“Attribute Extraction” 태그가 달린 논문 82편 · 필터 해제
Scaling E-Commerce Attribute Extraction with Parallel Decoding
Customers rely on specific product attributes to compare products and make purchasing decisions, but e-commerce catalogs are messy and unstructured, making it difficult to identify which attributes matter most and extrac…
Attribute Value ExtractionAttribute ExtractionHow Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA
Compositional visual question answering requires Vision-Language Models (VLMs) to execute multiple reasoning operations like object selection, spatial relation resolution, and attribute verification. Despite strong aggre…
Visual Question AnsweringAttribute ExtractionVisual GroundingSynthAVE: Scalable Synthetic Labeling for E-Commerce with LLM-Arena Validation
Fine-tuning large language models (LLMs) for e-commerce attribute extraction requires labeled data representative across thousands of product types, attributes, and multiple languages. This combinatorial scale translates…
Attribute Value ExtractionAttribute ExtractionCheXpercept: A Benchmark for Evaluating Expert-Level Lesion Perception in Chest X-rays
The evaluation of vision-language models (VLMs) for chest X-ray (CXR) analysis has largely been limited to disease-presence classification without visual grounding. Such evaluations fail to verify the expert-level lesion…
Attribute ExtractionDomain AdaptationVisual GroundingIndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products
Industrial products such as valves and circuit breakers are defined by dense technical specifications that govern procurement, compatibility, and safety across supply chains. These specifications are scattered across mul…
Attribute Value ExtractionAttribute ExtractionVisual ReasoningAREA: Attribute Extraction and Aggregation for CLIP-Based Class-Incremental Learning
Class-Incremental Learning (CIL) is important in building real-world learning systems. In CLIP-based CIL, the model performs classification by comparing similarity between visual and textual embeddings obtained from temp…
class-incremental learningAttribute ExtractionExperiments in Agentic AI for Science
This paper details two novel frameworks for developing autonomous, agentic AI in scientific workflows. Both systems leverage a hybrid Local Body, Remote Brain architecture via Google Colab, utilizing Python-based local o…
Attribute ExtractionKnowledge GraphsFashion Florence: Fine-Tuning Florence-2 for Structured Fashion Attribute Extraction
We present Fashion Florence, a Florence-2 vision-language model fine-tuned with LoRA to extract structured fashion attributes from clothing images. Given a single photograph, the model generates a JSON object containing …
Attribute ExtractionBreaking the Autoregressive Chain: Hyper-Parallel Decoding for Efficient LLM-Based Attribute Value Extraction
Some text generation tasks, such as Attribute Value Extraction (AVE), require decoding multiple independent sequences from the same document context. While standard autoregressive decoding is slow due to its sequential n…
Attribute Value ExtractionAttribute ExtractionText GenerationAutoPKG: An Automated Framework for Dynamic E-commerce Product-Attribute Knowledge Graph Construction
Product attribute extraction in e-commerce is bottlenecked by ontologies that are inconsistent, incomplete, and costly to maintain. We present AutoPKG, a multi-agent Large Language Model (LLM) framework that automaticall…
Attribute ExtractionVision-Language Attribute Disentanglement and Reinforcement for Lifelong Person Re-Identification
Lifelong person re-identification (LReID) aims to learn from varying domains to obtain a unified person retrieval model. Existing LReID approaches typically focus on learning from scratch or a visual classification-pretr…
Person Re-IdentificationAttribute ExtractionPerson RetrievalLogics-Parsing-Omni Technical Report
Addressing the challenges of fragmented task definitions and the heterogeneity of unstructured data in multimodal parsing, this paper proposes the Omni Parsing framework. This framework establishes a Unified Taxonomy cov…
Attribute ExtractionAdapting Vision-Language Models for E-commerce Understanding at Scale
E-commerce product understanding demands by nature, strong multimodal comprehension from text, images, and structured attributes. General-purpose Vision-Language Models (VLMs) enable generalizable multimodal latent model…
Instruction FollowingAttribute ExtractionDolphin-v2: Universal Document Parsing via Scalable Anchor Prompting
Document parsing has garnered widespread attention as vision-language models (VLMs) advance OCR capabilities. However, the field remains fragmented across dozens of specialized models with varying strengths, forcing user…
Attribute ExtractionNeural Sentinel: Unified Vision Language Model (VLM) for License Plate Recognition with Human-in-the-Loop Continual Learning
Traditional Automatic License Plate Recognition (ALPR) systems employ multi-stage pipelines consisting of object detection networks followed by separate Optical Character Recognition (OCR) modules, introducing compoundin…
License Plate RecognitionZero-shot GeneralizationAttribute ExtractionContinual LearningGeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation
Geometric problem solving constitutes a critical branch of mathematical reasoning, requiring precise analysis of shapes and spatial relationships. Current evaluations of geometric reasoning in vision-language models (VLM…
Mathematical ReasoningAttribute ExtractionBeyond Memorization: Selective Learning for Copyright-Safe Diffusion Model Training
Memorization in large-scale text-to-image diffusion models poses significant security and intellectual property risks, enabling adversarial attribute extraction and the unauthorized reproduction of sensitive or proprieta…
Attribute ExtractionCustomizing Open Source LLMs for Quantitative Medication Attribute Extraction across Heterogeneous EHR Systems
Harmonizing medication data across Electronic Health Record (EHR) systems is a persistent barrier to monitoring medications for opioid use disorder (MOUD). In heterogeneous EHR systems, key prescription attributes are sc…
Attribute ExtractionLLM-RG: Referential Grounding in Outdoor Scenarios using Large Language Models
Referential grounding in outdoor driving scenes is challenging due to large scene variability, many visually similar objects, and dynamic elements that complicate resolving natural-language references (e.g., "the black c…
Attribute ExtractionReferring ExpressionTowards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
Recent advances in large vision-language models (LVLMs) have demonstrated strong performance on general-purpose medical tasks. However, their effectiveness in specialized domains such as dentistry remains underexplored. …
Visual Question AnsweringAttribute Extraction