paper-with-me

홈 › Papers

On the Difference of BERT-style and CLIP-style Text Encoders

2023-06-06 · Zhihong Chen, Guiming Hardy Chen, Shizhe Diao, Xiang Wan, Benyou Wang

Masked language modeling (MLM) has been one of the most popular pretraining recipes in natural language processing, e.g., BERT, one of the representative models. Recently, contrastive language-image pretraining (CLIP) has also attracted attention, especially its vision models that achieve excellent performance on a broad range of vision tasks. However, few studies are dedicated to studying the text encoders learned by CLIP. In this paper, we analyze the difference between BERT-style and CLIP-style text encoders from three experiments: (i) general text understanding, (ii) vision-centric text understanding, and (iii) text-to-image generation. Experimental analyses show that although CLIP-style text encoders underperform BERT-style ones for general text understanding tasks, they are equipped with a unique ability, i.e., synesthesia, for the cross-modal association, which is more similar to the senses of humans.

📄 PDF Abstract BibTeX arXiv:2306.03678

Code (1)

zhjohnchan/bert-clip-synesthesia 공식 구현 pytorch

Tasks

Image GenerationLanguage ModelingLanguage ModellingMasked Language ModelingText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Multi-Head Attention 설명 없음
Attention 설명 없음
Residual Connection 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Adam 설명 없음

Similar Papers 제목 키워드 기반

DeltaEdit: Exploring Text-free Training for Text-Driven Image Manipulation

2023-03-11 · CVPR 2023 1 · Yueming Lyu, Tianwei Lin, Fu Li, Dongliang He 외

Text-driven image manipulation remains challenging in training or inference flexibility. Conditional generative models depend heavily on expensive annotated training data. Meanwhile, recent frameworks, which leverage pre…

Image Manipulation

StyleCLIPDraw: Coupling Content and Style in Text-to-Drawing Translation

2022-02-24 · Peter Schaldenbrand, Zhixuan Liu, Jean Oh

Generating images that fit a given text description using machine learning has improved greatly with the release of technologies such as the CLIP image-text encoder model; however, current methods lack artistic control o…

Style TransferTranslation

StyleCLIPDraw: Coupling Content and Style in Text-to-Drawing Synthesis

2021-11-04 · Peter Schaldenbrand, Zhixuan Liu, Jean Oh

Generating images that fit a given text description using machine learning has improved greatly with the release of technologies such as the CLIP image-text encoder model; however, current methods lack artistic control o…

Style Transfer

SEM-CS: Semantic CLIPStyler for Text-Based Image Style Transfer

2023-03-11 · Chanda G Kamra, Indra Deep Mastan, Debayan Gupta

CLIPStyler demonstrated image style transfer with realistic textures using only the style text description (instead of requiring a reference style image). However, the ground semantics of objects in style transfer output…

Style Transfer

Sem-CS: Semantic CLIPStyler for Text-Based Image Style Transfer

2023-07-12 · Chanda Grover Kamra, Indra Deep Mastan, Debayan Gupta

CLIPStyler demonstrated image style transfer with realistic textures using only a style text description (instead of requiring a reference style image). However, the ground semantics of objects in the style transfer outp…

Style Transfer