paper-with-me

Papers

GRAFT: GRaPH and Table Reasoning for Textual Alignment -- A Benchmark for Structured Instruction Following and Visual Reasoning

2025-08-21 · Abhigya Verma, Sriram Puttagunta, Seganrasan Subramanian, Sravan Ramachandran arxiv

GRAFT is a structured multimodal benchmark designed to probe how well LLMs handle instruction following, visual reasoning, and tasks requiring tight visual textual alignment. The dataset is built around programmatically generated charts and synthetically rendered tables, each paired with a carefully constructed, multi step analytical question that depends solely on what can be inferred from the image itself. Responses are formatted in structured outputs such as JSON or YAML, enabling consistent and fine grained evaluation of both reasoning processes and adherence to output specifications. The benchmark further introduces a taxonomy of reasoning operations ranging from comparison and trend identification to ranking, aggregation, proportional estimation, and anomaly detection to support a comprehensive assessment of model capabilities. Taken together, GRAFT provides a unified and scalable framework for evaluating multimodal LLMs on visually grounded, structured reasoning tasks, offering a more rigorous standard for future benchmarking efforts.

📄 PDF Abstract BibTeX arXiv:2508.15690

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction FollowingAnomaly DetectionVisual Reasoning

Similar Papers 제목 키워드 기반

Temporal Fusion Nexus: A task-agnostic multi-modal embedding model for clinical narratives and irregular time series in post-kidney transplant care

2026-01-13 · Aditya Kumar, Simon Rauch, Mario Cypko, Marcel Naik 외 arxiv

We introduce Temporal Fusion Nexus (TFN), a multi-modal and task-agnostic embedding model to integrate irregular time series and unstructured clinical narratives. We analysed TFN in post-kidney transplant (KTx) care, wit…

Mortality Prediction

GRAFT: A Graph-based Flow-aware Agentic Framework for Document-level Machine Translation

2025-07-04 · Himanshu Dutta, Sunny Manchanda, Prakhar Bapat, Meva Ram Gurjar 외

Document level Machine Translation (DocMT) approaches often struggle with effectively capturing discourse level phenomena. Existing approaches rely on heuristic rules to segment documents into discourse units, which rare…

Document Level Machine TranslationDocument TranslationLarge Language ModelMachine Translation+1

GRAFT: Biological Graph and Hypergraph Benchmarks for Linked Gene Expression and Phenotypic Trait Prediction in Arabidopsis thaliana

2026-06-25 · Manuel Serna-Aguilera, Vanshika Jindal, Fiona L. Goggin, Jiamei Li 외 arxiv

Understanding which genes control which traits in an organism remains one of the central challenges in biology. Despite significant advances in data collection technology, our ability to map genes to traits is still limi…

Graph RegressionGraph Learning

FreeGraftor: Training-Free Cross-Image Feature Grafting for Subject-Driven Text-to-Image Generation

2025-04-22 · Zebin Yao, Lei Ren, Huixing Jiang, Chen Wei 외

Subject-driven image generation aims to synthesize novel scenes that faithfully preserve subject identity from reference images while adhering to textual guidance, yet existing methods struggle with a critical trade-off …

Image GenerationText to Image GenerationText-to-Image Generation

GRAFT: Grid-Aware Load Forecasting with Multi-Source Textual Alignment and Fusion

2025-12-16 · Fangzhou Lin, Guoshun He, Zhenyu Guo, Zhe Huang 외 arxiv

Electric load is simultaneously affected across multiple time scales by exogenous factors such as weather and calendar rhythms, sudden events, and policies. Therefore, this paper proposes GRAFT (GRid-Aware Forecasting wi…