paper-with-me

홈 › Papers

Examining the Commitments and Difficulties Inherent in Multimodal Foundation Models for Street View Imagery

2024-08-23 · Zhenyuan Yang, Xuhui Lin, Qinyi He, Ziye Huang, Zhengliang Liu, Hanqi Jiang, Peng Shu, Zihao Wu, Yiwei Li, Stephen Law, Gengchen Mai, Tianming Liu, Tao Yang

The emergence of Large Language Models (LLMs) and multimodal foundation models (FMs) has generated heightened interest in their applications that integrate vision and language. This paper investigates the capabilities of ChatGPT-4V and Gemini Pro for Street View Imagery, Built Environment, and Interior by evaluating their performance across various tasks. The assessments include street furniture identification, pedestrian and car counts, and road width measurement in Street View Imagery; building function classification, building age analysis, building height analysis, and building structure classification in the Built Environment; and interior room classification, interior design style analysis, interior furniture counts, and interior length measurement in Interior. The results reveal proficiency in length measurement, style analysis, question answering, and basic image understanding, but highlight limitations in detailed recognition and counting tasks. While zero-shot learning shows potential, performance varies depending on the problem domains and image complexities. This study provides new insights into the strengths and weaknesses of multimodal foundation models for practical challenges in Street View Imagery, Built Environment, and Interior. Overall, the findings demonstrate foundational multimodal intelligence, emphasizing the potential of FMs to drive forward interdisciplinary applications at the intersection of computer vision and language.

📄 PDF Abstract BibTeX arXiv:2408.12821

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringZero-Shot Learning

Similar Papers 제목 키워드 기반

From Hesitancy Framings to Vaccine Hesitancy Profiles: A Journey of Stance, Ontological Commitments and Moral Foundations

2022-02-18 · Maxwell Weinzierl, Sanda Harabagiu

While billions of COVID-19 vaccines have been administered, too many people remain hesitant. Twitter, with its substantial reach and daily exposure, is an excellent resource for examining how people frame their vaccine h…

Intrinsic Barriers to Explaining Deep Foundation Models

2025-04-21 · Zhen Tan, Huan Liu

Deep Foundation Models (DFMs) offer unprecedented capabilities but their increasing complexity presents profound challenges to understanding their internal workings-a critical need for ensuring trust, safety, and account…

Hard Choices in Artificial Intelligence: Addressing Normative Uncertainty through Sociotechnical Commitments

2019-11-20 · Roel Dobbe, Thomas Krendl Gilbert, Yonatan Mintz

As AI systems become prevalent in high stakes domains such as surveillance and healthcare, researchers now examine how to design and implement them in a safe manner. However, the potential harms caused by systems to stak…

Navigate

Ten Challenging Problems in Federated Foundation Models

2025-02-14 · Tao Fan, Hanlin Gu, Xuemei Cao, Chee Seng Chan 외

Federated Foundation Models (FedFMs) represent a distributed learning paradigm that fuses general competences of foundation models as well as privacy-preserving capabilities of federated learning. This combination allows…

Continual LearningFederated LearningPrivacy PreservingTransfer Learning

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA

2025-05-09 · Karthik Reddy Kanjula, Surya Guthikonda, Nahid Alam, Shayekh Bin Islam

Pretraining datasets are foundational to the development of multimodal models, yet they often have inherent biases and toxic content from the web-scale corpora they are sourced from. In this paper, we investigate the pre…