paper-with-me

Papers

STREET: A Multi-Task Structured Reasoning and Explanation Benchmark

2023-02-13 · Danilo Ribeiro, Shen Wang, Xiaofei Ma, Henry Zhu, Rui Dong, Deguang Kong, Juliette Burger, Anjelica Ramos, William Wang, Zhiheng Huang, George Karypis, Bing Xiang, Dan Roth

We introduce STREET, a unified multi-task and multi-domain natural language reasoning and explanation benchmark. Unlike most existing question-answering (QA) datasets, we expect models to not only answer questions, but also produce step-by-step structured explanations describing how premises in the question are used to produce intermediate conclusions that can prove the correctness of a certain answer. We perform extensive evaluation with popular language models such as few-shot prompting GPT-3 and fine-tuned T5. We find that these models still lag behind human performance when producing such structured reasoning steps. We believe this work will provide a way for the community to better train and test systems on multi-step reasoning and explanations in natural language.

📄 PDF Abstract BibTeX arXiv:2302.06729

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음
Attention 설명 없음
Test 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Adafactor Adafactor is a stochastic optimization method based on Adam that reduces memory usage while retaining the empirical benefits of…
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…

Similar Papers 제목 키워드 기반

Eliciting Better Multilingual Structured Reasoning from LLMs through Code

2024-03-05 · Bryan Li, Tamer Alkhouli, Daniele Bonadiman, Nikolaos Pappas 외

The development of large language models (LLM) has shown progress on reasoning, though studies have largely considered either English or simple reasoning tasks. To address this, we introduce a multilingual structured rea…

Machine Translation

SEER: Facilitating Structured Reasoning and Explanation via Reinforcement Learning

2024-01-24 · Guoxin Chen, Kexin Tang, Chao Yang, Fuying Ye 외

Elucidating the reasoning process with structured explanations from question to answer is crucial, as it significantly enhances the interpretability, traceability, and trustworthiness of question-answering (QA) systems. …

Question Answeringreinforcement-learningReinforcement LearningReinforcement Learning (RL)

GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View

2025-06-19 · Fenghua Cheng, Jinxiang Wang, Sen Wang, Zi Huang 외

Multimodal reasoning is a process of understanding, integrating and inferring information across different data modalities. It has recently attracted surging academic attention as a benchmark for Artificial Intelligence …

Multimodal Reasoning

m2sv: A Scalable Benchmark for Map-to-Street-View Spatial Reasoning

2026-01-27 · Yosub Shin, Michael Buriek, Igor Molybog arxiv

Vision--language models (VLMs) achieve strong performance on many multimodal benchmarks but remain brittle on spatial reasoning tasks that require aligning abstract overhead representations with egocentric views. We intr…

Reinforcement LearningSpatial Reasoning

StreetDiff: Multi-view Street Scenes Generation via Cross-view Consistent Multi-view Stable Diffusion with Structure Prompts

2026-09-09 · Qi Zhang, Yanyifan Wang, Weiyuan Zhang, Hui Huang arxiv

Multi-view diffusion models have shown strong performance in scenes with strong geometric priors and sparse semantics, such as indoor rooms or simple outdoor environments (e.g., fields, courtyards). However, they often f…

Scene Generation