paper-with-me

홈 › Papers

LLMs can see and hear without any training

2025-01-30 · Kumar Ashutosh, Yossi Gandelsman, Xinlei Chen, Ishan Misra, Rohit Girdhar

We present MILS: Multimodal Iterative LLM Solver, a surprisingly simple, training-free approach, to imbue multimodal capabilities into your favorite LLM. Leveraging their innate ability to perform multi-step reasoning, MILS prompts the LLM to generate candidate outputs, each of which are scored and fed back iteratively, eventually generating a solution to the task. This enables various applications that typically require training specialized models on task-specific data. In particular, we establish a new state-of-the-art on emergent zero-shot image, video and audio captioning. MILS seamlessly applies to media generation as well, discovering prompt rewrites to improve text-to-image generation, and even edit prompts for style transfer! Finally, being a gradient-free optimization approach, MILS can invert multimodal embeddings into text, enabling applications like cross-modal arithmetic.

📄 PDF Abstract BibTeX arXiv:2501.18096

Code (1)

facebookresearch/mils 공식 구현 pytorch

Tasks

Audio captioningImage GenerationStyle TransferText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

Hypergraph Enterprise Agentic Reasoner over Heterogeneous Business Systems

2026-05-14 · Ling Wang, Xin Liu, Songnan Liu, Jianan Wang 외 arxiv

Applying Large Language Models (LLMs) to heterogeneous enterprise systems is hindered by hallucinations and failures in multi-hop, n-ary reasoning. Existing paradigms (e.g., GraphRAG, NL2SQL) lack the semantic grounding …

Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized Rehearsal

2024-03-02 · Jianheng Huang, Leyang Cui, Ante Wang, Chengyi Yang 외

Large language models (LLMs) suffer from catastrophic forgetting during continual learning. Conventional rehearsal-based methods rely on previous training data to retain the model's ability, which may not be feasible in …

Continual LearningIn-Context Learning

Protecting Bystander Privacy via Selective Hearing in Audio LLMs

2025-12-06 · Xiao Zhan, Guangzhi Sun, Jose Such, Phil Woodland arxiv

Audio Large language models (LLMs) are increasingly deployed in the real world, where they inevitably capture speech from unintended nearby bystanders, raising privacy risks that existing benchmarks and defences did not …

Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

2023-10-10 · Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng, Danqi Chen

The popularity of LLaMA (Touvron et al., 2023a;b) and other recently emerged moderate-sized large language models (LLMs) highlights the potential of building smaller yet powerful LLMs. Regardless, the cost of training su…

Language ModelingLanguage ModellingQuestion AnsweringSentence Completion

MLLMs-Augmented Visual-Language Representation Learning

2023-11-30 · Yanqing Liu, Kai Wang, Wenqi Shao, Ping Luo 외

Visual-language pre-training has achieved remarkable success in many multi-modal tasks, largely attributed to the availability of large-scale image-text datasets. In this work, we demonstrate that Multi-modal Large Langu…

Image-text RetrievalRepresentation LearningRetrievalText Retrieval