paper-with-me

Papers

GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning

2025-05-28 · Shikhhar Siingh, Abhinav Rawat, Chitta Baral, Vivek Gupta

Publicly significant images from events hold valuable contextual information, crucial for journalism and education. However, existing methods often struggle to extract this relevance accurately. To address this, we introduce GETReason (Geospatial Event Temporal Reasoning), a framework that moves beyond surface-level image descriptions to infer deeper contextual meaning. We propose that extracting global event, temporal, and geospatial information enhances understanding of an image's significance. Additionally, we introduce GREAT (Geospatial Reasoning and Event Accuracy with Temporal Alignment), a new metric for evaluating reasoning-based image understanding. Our layered multi-agent approach, assessed using a reasoning-weighted metric, demonstrates that meaningful insights can be inferred, effectively linking images to their broader event context.

📄 PDF Abstract BibTeX arXiv:2505.21863

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SLIDE: Sliding Localized Information for Document Extraction

2025-03-23 · Divyansh Singh, Manuel Nunez Martinez, Bonnie J. Dorr, Sonja Schmer Galunder

Constructing accurate knowledge graphs from long texts and low-resource languages is challenging, as large language models (LLMs) experience degraded performance with longer input chunks. This problem is amplified in low…

Chunkinggraph constructionKnowledge GraphsQuestion Answering+1

Vision Language Model-based Caption Evaluation Method Leveraging Visual Context Extraction

2024-02-28 · Koki Maeda, Shuhei Kurita, Taiki Miyanishi, Naoaki Okazaki

Given the accelerating progress of vision and language modeling, accurate evaluation of machine-generated image captions remains critical. In order to evaluate captions more closely to human preferences, metrics need to …

Image CaptioningLanguage ModelingLanguage Modelling

LR-FPN: Enhancing Remote Sensing Object Detection with Location Refined Feature Pyramid Network

2024-04-02 · Hanqian Li, Ruinan Zhang, Ye Pan, Junchi Ren 외

Remote sensing target detection aims to identify and locate critical targets within remote sensing images, finding extensive applications in agriculture and urban planning. Feature pyramid networks (FPNs) are commonly us…

Objectobject-detectionObject Detection

How transfer learning is used in generative models for image classification: improved accuracy

2024-12-09 · Signal, Image and Video Processing 2024 12 · Danial Ebrahimzadeh, Sarah S Sharif, Yaser M Banad

Recent breakthroughs in generative neural networks have paved the way for transformative capabilities, particularly in their capacity to generate novel data, notably in the realm of images. The integration of these model…

ClassificationGenerative Adversarial Networkimage-classificationImage Classification+1

Enhancing Clinical Concept Extraction with Contextual Embeddings

2019-02-22 · Yuqi Si, Jingqi Wang, Hua Xu, Kirk Roberts

Neural network-based representations ("embeddings") have dramatically advanced natural language processing (NLP) tasks, including clinical NLP tasks such as concept extraction. Recently, however, more advanced embedding …

Clinical Concept ExtractionLanguage ModellingLarge Language ModelWord Embeddings