paper-with-me

홈 › Papers

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

2025-01-10 · Ziheng Wu, Zhenghao Chen, Ruipu Luo, Can Zhang, Yuan Gao, Zhentao He, Xian Wang, Haoran Lin, Minghui Qiu

Recently, vision-language models have made remarkable progress, demonstrating outstanding capabilities in various tasks such as image captioning and video understanding. We introduce Valley2, a novel multimodal large language model designed to enhance performance across all domains and extend the boundaries of practical applications in e-commerce and short video scenarios. Notably, Valley2 achieves state-of-the-art (SOTA) performance on e-commerce benchmarks, surpassing open-source models of similar size by a large margin (79.66 vs. 72.76). Additionally, Valley2 ranks second on the OpenCompass leaderboard among models with fewer than 10B parameters, with an impressive average score of 67.4. The code and model weights are open-sourced at https://github.com/bytedance/Valley.

📄 PDF Abstract BibTeX arXiv:2501.05901

Code (1)

bytedance/valley 공식 구현 pytorch

Tasks

Image CaptioningLanguage ModelingLanguage ModellingLarge Language ModelMultimodal Large Language ModelVideo Understanding

Similar Papers 제목 키워드 기반

Valley3: Scaling Omni Foundation Models for E-commerce

2026-05-02 · Zeyu Chen, Guanghao Zhou, Qixiang Yin, Ziwang Zhao 외 arxiv

In this work, we present Valley3, an omni multimodal large language model (MLLM) developed for diverse global e-commerce tasks, with unified understanding and reasoning capabilities across text, images, video, and audio.…

Valley: Video Assistant with Large Language model Enhanced abilitY

2023-06-12 · Ruipu Luo, Ziwang Zhao, Min Yang, Junwei DOng 외

Large language models (LLMs), with their remarkable conversational capabilities, have demonstrated impressive performance across various applications and have emerged as formidable AI assistants. In view of this, it rais…

Action RecognitionInstruction FollowingLanguage ModelingLanguage Modelling+3

Exploring and Exploiting the Asymmetric Valley of Deep Neural Networks

2024-05-21 · Xin-Chun Li, Jin-Lin Tang, Bo Zhang, Lan Li 외

Exploring the loss landscape offers insights into the inherent principles of deep neural networks (DNNs). Recent work suggests an additional asymmetry of the valley beyond the flat and sharp ones, yet without thoroughly …

Federated Learning

Debating for Better Reasoning: An Unsupervised Multimodal Approach

2025-05-20 · Ashutosh Adhikari, Mirella Lapata

As Large Language Models (LLMs) gain expertise across diverse domains and modalities, scalable oversight becomes increasingly challenging, particularly when their capabilities may surpass human evaluators. Debate has eme…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Benchmarking the Hill-Valley Evolutionary Algorithm for the GECCO 2018 Competition on Niching Methods Multimodal Optimization

2018-06-30 · S. C. Maree, T. Alderliesten, D. Thierens, P. A. N. Bosman

This report presents benchmarking results of the latest version of the Hill-Valley Evolutionary Algorithm (HillVallEA) on the CEC2013 niching benchmark suite. The benchmarking follows restrictions required by the GECCO 2…

Benchmarking