paper-with-me

홈 › Papers

Planting a SEED of Vision in Large Language Model

2023-07-16 · Yuying Ge, Yixiao Ge, Ziyun Zeng, Xintao Wang, Ying Shan

We present SEED, an elaborate image tokenizer that empowers Large Language Models (LLMs) with the emergent ability to SEE and Draw at the same time. Research on image tokenizers has previously reached an impasse, as frameworks employing quantized visual tokens have lost prominence due to subpar performance and convergence in multimodal comprehension (compared to BLIP-2, etc.) or generation (compared to Stable Diffusion, etc.). Despite the limitations, we remain confident in its natural capacity to unify visual and textual representations, facilitating scalable multimodal training with LLM's original recipe. In this study, we identify two crucial principles for the architecture and training of SEED that effectively ease subsequent alignment with LLMs. (1) Image tokens should be independent of 2D physical patch positions and instead be produced with a 1D causal dependency, exhibiting intrinsic interdependence that aligns with the left-to-right autoregressive prediction mechanism in LLMs. (2) Image tokens should capture high-level semantics consistent with the degree of semantic abstraction in words, and be optimized for both discriminativeness and reconstruction during the tokenizer training phase. As a result, the off-the-shelf LLM is able to perform both image-to-text and text-to-image generation by incorporating our SEED through efficient LoRA tuning. Comprehensive multimodal pretraining and instruction tuning, which may yield improved results, are reserved for future investigation. This version of SEED was trained in 5.7 days using only 64 V100 GPUs and 5M publicly available image-text pairs. Our preliminary study emphasizes the great potential of discrete visual tokens in versatile multimodal LLMs and the importance of proper image tokenizers in broader research.

📄 PDF Abstract BibTeX arXiv:2307.08041

Code (1)

ailab-cvc/seed 공식 구현 pytorch

Tasks

Image GenerationImage to textLanguage ModelingLanguage ModellingLarge Language ModelText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Scheduling Planting Time Through Developing an Optimization Model and Analysis of Time Series Growing Degree Units

2022-07-02 · Javad Ansarifar, Faezeh Akhavizadegan, Lizhi Wang

Producing higher-quality crops within shortened breeding cycles ensures global food availability and security, but this improvement intensifies logistical and productivity challenges for seed industries in the year-round…

SchedulingTime SeriesTime Series Analysis

Maize Seedling Detection Dataset (MSDD): A Curated High-Resolution RGB Dataset for Seedling Maize Detection and Benchmarking with YOLOv9, YOLO11, YOLOv12 and Faster-RCNN

2025-09-18 · Dewi Endah Kharismawati, Toni Kazic arxiv

Accurate maize seedling detection is crucial for precision agriculture, yet curated datasets remain scarce. We introduce MSDD, a high-quality aerial image dataset for maize seedling stand counting, with applications in e…

Hierarchical Modeling of Seed Variety Yields and Decision Making for Future Planting Plans

2017-11-15 · Huaiyang Zhong, Xiaocheng Li, David Lobell, Stefano Ermon 외

Eradicating hunger and malnutrition is a key development goal of the 21st century. We address the problem of optimally identifying seed varieties to reliably increase crop yield within a risk-sensitive decision-making fr…

Decision MakingDecision Making Under UncertaintyWeather Forecasting

Automatic counting of mounds on UAV images: combining instance segmentation and patch-level correction

2022-09-06 · Majid Nikougoftar Nategh, Ahmed Zgaren, Wassim Bouachir, Nizar Bouguila

Site preparation by mounding is a commonly used silvicultural treatment that improves tree growth conditions by mechanically creating planting microsites called mounds. Following site preparation, the next critical step …

Instance Segmentationobject-detectionObject DetectionSemantic Segmentation

A random planting model

2024-09-24 · Julian Talbot, Pascal Viot, David Colliaux

The adoption of agroecological practices will be crucial to address the challenges of climate change and biodiversity loss. Such practices favor the cultivation of plants in complex mixtures with layouts differing from t…

model