paper-with-me

홈 › Papers

StoryTailor:A Zero-Shot Pipeline for Action-Rich Multi-Subject Visual Narratives

2026-02-24 · Jinghao Hu, Yuhe Zhang, GuoHua Geng, Kang Li, Han Zhang arxiv

Generating multi-frame, action-rich visual narratives without fine-tuning faces a threefold tension: action text faithfulness, subject identity fidelity, and cross-frame background continuity. We propose StoryTailor, a zero-shot pipeline that runs on a single RTX 4090 (24 GB) and produces temporally coherent, identity-preserving image sequences from a long narrative prompt, per-subject references, and grounding boxes. Three synergistic modules drive the system: Gaussian-Centered Attention (GCA) to dynamically focus on each subject core and ease grounding-box overlaps; Action-Boost Singular Value Reweighting (AB-SVR) to amplify action-related directions in the text embedding space; and Selective Forgetting Cache (SFC) that retains transferable background cues, forgets nonessential history, and selectively surfaces retained cues to build cross-scene semantic ties. Compared with baseline methods, experiments show that CLIP-T improves by up to 10-15%, with DreamSim lower than strong baselines, while CLIP-I stays in a visually acceptable, competitive range. With matched resolution and steps on a 24 GB GPU, inference is faster than FluxKontext. Qualitatively, StoryTailor delivers expressive interactions and evolving yet stable scenes.

📄 PDF Abstract BibTeX arXiv:2602.21273

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Reading between the Lines: Leveraging Large Language Models for Global Dementia and Depression Assessment from Clinical Interviews

2026-06-16 · Franziska Braun, Alea Rüggeberg, Thomas Ranzenberger, Hartmut Lehfeld 외 arxiv

Dementia and depression are the most prevalent neuropsychiatric disorders in geriatric populations, and their overlapping symptoms pose major challenges for differential diagnosis. In this study, we investigate open-weig…

A Monte Carlo Language Model Pipeline for Zero-Shot Sociopolitical Event Extraction

2023-05-24 · Erica Cai, Brendan O'Connor

Current social science efforts automatically populate event databases of "who did what to whom?" tuples, by applying event extraction (EE) to text such as news. The event databases are used to analyze sociopolitical dyna…

Computational EfficiencyEvent ExtractionInstruction FollowingLanguage Modeling+3

MASTE: A Multi-Agent Pipeline for Zero-Shot Aspect Sentiment Triplet Extraction

2026-07-09 · Ao Hong, Lehang Wang, Zhirun Yue, Mingxin Wang 외 arxiv

Aspect Sentiment Triplet Extraction (ASTE) requires jointly identifying (aspect, opinion, sentiment) triples from a given review sentence. While large language models (LLMs) achieve strong zero-shot performance on many N…

Aspect Sentiment Triplet Extraction

Ontology Enrichment for Effective Fine-grained Entity Typing

2023-10-11 · Siru Ouyang, Jiaxin Huang, Pranav Pillai, Yunyi Zhang 외

Fine-grained entity typing (FET) is the task of identifying specific entity types at a fine-grained level for entity mentions based on their contextual information. Conventional methods for FET require extensive human an…

Entity Typing

Zero-Splat TeleAssist: A Zero-Shot Pose Estimation Framework for Semantic Teleoperation

2025-12-09 · Srijan Dokania, Dharini Raghavan arxiv

We introduce Zero-Splat TeleAssist, a zero-shot sensor-fusion pipeline that transforms commodity CCTV streams into a shared, 6-DoF world model for multilateral teleoperation. By integrating vision-language segmentation, …

Pose Estimation