paper-with-me

홈 › Papers

Can Language Models Laugh at YouTube Short-form Videos?

2023-10-22 · Dayoon Ko, Sangho Lee, Gunhee Kim

As short-form funny videos on social networks are gaining popularity, it becomes demanding for AI models to understand them for better communication with humans. Unfortunately, previous video humor datasets target specific domains, such as speeches or sitcoms, and mostly focus on verbal cues. We curate a user-generated dataset of 10K multimodal funny videos from YouTube, called ExFunTube. Using a video filtering pipeline with GPT-3.5, we verify both verbal and visual elements contributing to humor. After filtering, we annotate each video with timestamps and text explanations for funny moments. Our ExFunTube is unique over existing datasets in that our videos cover a wide range of domains with various types of humor that necessitate a multimodal understanding of the content. Also, we develop a zero-shot video-to-text prompting to maximize video humor understanding of large language models (LLMs). With three different evaluation methods using automatic scores, rationale quality experiments, and human evaluations, we show that our prompting significantly improves LLMs' ability for humor explanation.

📄 PDF Abstract BibTeX arXiv:2310.14159

Code (1)

dayoon-ko/exfuntube 공식 구현 pytorch

Tasks

Form

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음
Attention 설명 없음
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Adam 설명 없음

Similar Papers 제목 키워드 기반

Laugh, Relate, Engage: Stylized Comment Generation for Short Videos

2025-11-05 · Xuan Ouyang, Senan Wang, Bouzhou Wang, Siyuan Xiahou 외 arxiv

Short-video platforms have become a central medium in the modern Internet landscape, where efficient information delivery and strong interactivity are reshaping user engagement and cultural dissemination. Among the vario…

Video Segmentation

SMILE: Multimodal Dataset for Understanding Laughter in Video with Language Models

2023-12-15 · Lee Hyun, Kim Sung-Bin, Seungju Han, Youngjae Yu 외

Despite the recent advances of the artificial intelligence, building social intelligence remains a challenge. Among social signals, laughter is one of the distinctive expressions that occurs during social interactions be…

Video Understanding

FinCap: Topic-Aligned Captions for Short-Form Financial YouTube Videos

2025-09-30 · Siddhant Sukhani, Yash Bhardwaj, Riya Bhadani, Veer Kejriwal 외 arxiv

We evaluate multimodal large language models (MLLMs) for topic-aligned captioning in financial short-form videos (SVs) by testing joint reasoning over transcripts (T), audio (A), and video (V). Using 624 annotated YouTub…

Sentiment AnalysisVideo Captioning

YouTube Videos for Public Health Literacy? A Machine Learning Pipeline to Curate Covid-19 Videos

2023-11-21 · Yawen Guo, Xiao Liu, Anjana Susarla, Rema Padman

The COVID-19 pandemic has highlighted the dire necessity to improve public health literacy for societal resilience. YouTube, the largest video-sharing social media platform, provides a vast repository of user-generated h…

ManagementRetrieval

Content and Engagement Trends in COVID-19 YouTube Videos: Evidence from the Late Pandemic

2025-09-02 · Nirmalya Thakur, Madeline D Hartel, Lane Michael Boden, Dallas Enriquez 외 arxiv

This work investigated about 10,000 COVID-19-related YouTube videos published between January 2023 and October 2024 to evaluate how temporal, lexical, linguistic, and structural factors influenced engagement during the l…