paper-with-me

홈 › Papers

How Good is Google Bard's Visual Understanding? An Empirical Study on Open Challenges

2023-07-27 · Haotong Qin, Ge-Peng Ji, Salman Khan, Deng-Ping Fan, Fahad Shahbaz Khan, Luc van Gool

Google's Bard has emerged as a formidable competitor to OpenAI's ChatGPT in the field of conversational AI. Notably, Bard has recently been updated to handle visual inputs alongside text prompts during conversations. Given Bard's impressive track record in handling textual inputs, we explore its capabilities in understanding and interpreting visual data (images) conditioned by text questions. This exploration holds the potential to unveil new insights and challenges for Bard and other forthcoming multi-modal Generative models, especially in addressing complex computer vision problems that demand accurate visual and language understanding. Specifically, in this study, we focus on 15 diverse task scenarios encompassing regular, camouflaged, medical, under-water and remote sensing data to comprehensively evaluate Bard's performance. Our primary finding indicates that Bard still struggles in these vision scenarios, highlighting the significant gap in vision-based understanding that needs to be bridged in future developments. We expect that this empirical study will prove valuable in advancing future models, leading to enhanced capabilities in comprehending and interpreting fine-grained visual data. Our project is released on https://github.com/htqin/GoogleBard-VisUnderstand

📄 PDF Abstract BibTeX arXiv:2307.15016

Code (1)

htqin/googlebard-visunderstand 공식 구현

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Multimodal Analysis Of Google Bard And GPT-Vision: Experiments In Visual Reasoning

2023-08-17 · David Noever, Samantha Elizabeth Miller Noever

Addressing the gap in understanding visual comprehension in Large Language Models (LLMs), we designed a challenge-response study, subjecting Google Bard and GPT-Vision to 64 visual tasks, spanning categories like "Visual…

Common Sense ReasoningOptical Character RecognitionVisual Reasoning

Performance Comparison of Large Language Models on VNHSGE English Dataset: OpenAI ChatGPT, Microsoft Bing Chat, and Google Bard

2023-07-05 · Xuan-Quy Dao

This paper presents a performance comparison of three large language models (LLMs), namely OpenAI ChatGPT, Microsoft Bing Chat (BingChat), and Google Bard, on the VNHSGE English dataset. The performance of BingChat, Bard…

Question Answering

How Robust is Google's Bard to Adversarial Image Attacks?

2023-09-21 · Yinpeng Dong, Huanran Chen, Jiawei Chen, Zhengwei Fang 외

Multimodal Large Language Models (MLLMs) that integrate text and other modalities (especially vision) have achieved unprecedented performance in various multimodal tasks. However, due to the unsolved adversarial robustne…

Adversarial RobustnessChatbotFace Detection

Bridging History with AI A Comparative Evaluation of GPT 3.5, GPT4, and GoogleBARD in Predictive Accuracy and Fact Checking

2023-05-13 · Davut Emre Tasar, Ceren Ocal Tasar

The rapid proliferation of information in the digital era underscores the importance of accurate historical representation and interpretation. While artificial intelligence has shown promise in various fields, its potent…

Fact Checking

Can ChatGPT and Bard Generate Aligned Assessment Items? A Reliability Analysis against Human Performance

2023-04-09 · Abdolvahab Khademi

ChatGPT and Bard are AI chatbots based on Large Language Models (LLM) that are slated to promise different applications in diverse areas. In education, these AI technologies have been tested for applications in assessmen…

Automated Essay Scoring