paper-with-me

홈 › Papers

Bollywood Movie Corpus for Text, Images and Videos

2017-10-11 · Nishtha Madaan, Sameep Mehta, Mayank Saxena, Aditi Aggarwal, Taneea S Agrawaal, Vrinda Malhotra

In past few years, several data-sets have been released for text and images. We present an approach to create the data-set for use in detecting and removing gender bias from text. We also include a set of challenges we have faced while creating this corpora. In this work, we have worked with movie data from Wikipedia plots and movie trailers from YouTube. Our Bollywood Movie corpus contains 4000 movies extracted from Wikipedia and 880 trailers extracted from YouTube which were released from 1970-2017. The corpus contains csv files with the following data about each movie - Wikipedia title of movie, cast, plot text, co-referenced plot text, soundtrack information, link to movie poster, caption of movie poster, number of males in poster, number of females in poster. In addition to that, corresponding to each cast member the following data is available - cast name, cast gender, cast verbs, cast adjectives, cast relations, cast centrality, cast mentions. We present some preliminary results on the task of bias removal which suggest that the data-set is quite useful for performing such tasks.

📄 PDF Abstract BibTeX arXiv:1710.04142

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Nollywood: Let's Go to the Movies!

2024-07-02 · John E. Ortega, Ibrahim Said Ahmad, William Chen

Nollywood, based on the idea of Bollywood from India, is a series of outstanding movies that originate from Nigeria. Unfortunately, while the movies are in English, they are hard to understand for many native speakers du…

Learning Video Context as Interleaved Multimodal Sequences

2024-07-31 · Kevin Qinghong Lin, Pengchuan Zhang, Difei Gao, Xide Xia 외

Narrative videos, such as movies, pose significant challenges in video understanding due to their rich contexts (characters, dialogues, storylines) and diverse demands (identify who, relationship, and reason). In this pa…

Language ModelingLanguage ModellingQuestion AnsweringText Retrieval+5

All It Takes is 20 Questions!: A Knowledge Graph Based Approach

2019-11-12 · Alvin Dey, Harsh Kumar Jain, Vikash Kumar Pandey, Tanmoy Chakraborty

20 Questions (20Q) is a two-player game. One player is the answerer, and the other is a questioner. The answerer chooses an entity from a specified domain and does not reveal this to the other player. The questioner can …

AllQuestion GenerationQuestion-GenerationRecommendation Systems

MovieFactory: Automatic Movie Creation from Text using Large Generative Models for Language and Images

2023-06-12 · Junchen Zhu, Huan Yang, Huiguo He, Wenjing Wang 외

In this paper, we present MovieFactory, a powerful framework to generate cinematic-picture (3072$\times$1280), film-style (multi-scene), and multi-modality (sounding) movies on the demand of natural languages. As the fir…

Retrieval

Moviescope: Large-scale Analysis of Movies using Multiple Modalities

2019-08-08 · Paola Cascante-Bonilla, Kalpathy Sitaraman, Mengjia Luo, Vicente Ordonez

Film media is a rich form of artistic expression. Unlike photography, and short videos, movies contain a storyline that is deliberately complex and intricate in order to engage its audience. In this paper we present a la…