paper-with-me

홈 › Papers

LLM as an Art Director (LaDi): Using LLMs to improve Text-to-Media Generators

2023-11-07 · Allen Roush, Emil Zakirov, Artemiy Shirokov, Polina Lunina, Jack Gane, Alexander Duffy, Charlie Basil, Aber Whitcomb, Jim Benedetto, Chris DeWolfe

Recent advancements in text-to-image generation have revolutionized numerous fields, including art and cinema, by automating the generation of high-quality, context-aware images and video. However, the utility of these technologies is often limited by the inadequacy of text prompts in guiding the generator to produce artistically coherent and subject-relevant images. In this paper, We describe the techniques that can be used to make Large Language Models (LLMs) act as Art Directors that enhance image and video generation. We describe our unified system for this called "LaDi". We explore how LaDi integrates multiple techniques for augmenting the capabilities of text-to-image generators (T2Is) and text-to-video generators (T2Vs), with a focus on constrained decoding, intelligent prompting, fine-tuning, and retrieval. LaDi and these techniques are being used today in apps and platforms developed by Plai Labs.

📄 PDF Abstract BibTeX arXiv:2311.03716

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationRetrievalText to Image GenerationText-to-Image GenerationVideo Generation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases

2025-09-25 · Sri Vatsa Vuddanti, Aarav Shah, Satwik Kumar Chittiprolu, Tony Song 외 arxiv

Tool-augmented language agents frequently fail in real-world deployment due to tool malfunctions--timeouts, API exceptions, or inconsistent outputs--triggering cascading reasoning errors and task abandonment. Existing ag…

LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning

2025-10-06 · Haoqiang Kang, Yizhe Zhang, Nikki Lijing Kuang, Nicklas Majamaki 외 arxiv

Large Language Models (LLMs) demonstrate their reasoning ability through chain-of-thought (CoT) generation. However, LLM's autoregressive decoding may limit the ability to revisit and refine earlier tokens in a holistic …

Mathematical ReasoningCode Generation

ALADIN:Attribute-Language Distillation Network for Person Re-Identification

2026-03-23 · Wang Zhou, Boran Duan, Haojun Ai, Ruiqi Lan 외 arxiv

Recent vision-language models such as CLIP provide strong cross-modal alignment, but current CLIP-guided ReID pipelines rely on global features and fixed prompts. This limits their ability to capture fine-grained attribu…

Person Re-IdentificationRepresentation Learning

Exploring NLP Benchmarks in an Extremely Low-Resource Setting

2025-09-04 · Ulin Nuha, Adam Jatowt arxiv

The effectiveness of Large Language Models (LLMs) diminishes for extremely low-resource languages, such as indigenous languages, primarily due to the lack of labeled data. Despite growing interest, the availability of hi…

Machine TranslationSentiment AnalysisQuestion Answering

Distributed Consensus Optimization with Consensus ALADIN

2025-03-21 · Xu Du, Jingzhe Wang

TThe paper proposes the Consensus Augmented Lagrange Alternating Direction Inexact Newton (Consensus ALADIN) algorithm, a novel approach for solving distributed consensus optimization problems (DC). Consensus ALADIN allo…

Computational Efficiency