paper-with-me

홈 › Papers

Vision Bridge Transformer at Scale

2025-11-28 · Zhenxiong Tan, Zeqing Wang, Xingyi Yang, Songhua Liu, Xinchao Wang arxiv

We introduce Vision Bridge Transformer (ViBT), a large-scale instantiation of Brownian Bridge Models designed for conditional generation. Unlike traditional diffusion models that transform noise into data, Bridge Models directly model the trajectory between inputs and outputs, creating an efficient data-to-data translation paradigm. By scaling these models to 20B and 1.3B parameters, we demonstrate their effectiveness for image and video translation tasks. To support this scale, we adopt a Transformer architecture and propose a variance-stabilized velocity-matching objective for robust training. Together, these advances highlight the power of scaling Bridge Models for instruction-based image editing and complex video translation.

📄 PDF Abstract BibTeX arXiv:2511.23199

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers

2026-06-30 · Ying Fan, Anej Svete, Kangwook Lee arxiv

Language models typically reason via explicit chain-of-thought (CoT), generating intermediate steps token-by-token. Latent CoT offers an alternative: it performs multi-step reasoning in the model's hidden states, replaci…

HBFormer: A Hybrid-Bridge Transformer for Microtumor and Miniature Organ Segmentation

2025-12-03 · Fuchen Zheng, Xinyi Chen, Weixuan Li, Quanjun Li 외 arxiv

Medical image segmentation is a cornerstone of modern clinical diagnostics. While Vision Transformers that leverage shifted window-based self-attention have established new benchmarks in this field, they are often hamper…

Medical Image Segmentation

Point Bridge: 3D Representations for Cross Domain Policy Learning

2026-01-22 · Siddhant Haldar, Lars Johannsmeier, Lerrel Pinto, Abhishek Gupta 외 arxiv

Robot foundation models are beginning to deliver on the promise of generalist robotic agents, yet progress remains constrained by the scarcity of large-scale real-world manipulation datasets. Simulation and synthetic dat…

Synthetic Data Generation

Class-agnostic Object Detection with Multi-modal Transformer

2021-11-22 · Muhammad Maaz, Hanoona Rasheed, Salman Khan, Fahad Shahbaz Khan 외

What constitutes an object? This has been a long-standing question in computer vision. Towards this goal, numerous learning-free and learning-based approaches have been developed to score objectness. However, they genera…

Class-agnostic Object DetectionObjectobject-detectionObject Detection+2

Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding

2026-04-21 · Kadir Yilmaz, Adrian Kruse, Tristan Höfer, Daan de Geus 외 arxiv

Transformers have become a common foundation across deep learning, yet 3D scene understanding still relies on specialized backbones with strong domain priors. This keeps the field isolated from the broader Transformer ec…

3D Semantic Segmentation3D Instance SegmentationScene Understanding