Point-of-Interest Type Prediction using Text and Images
Point-of-interest (POI) type prediction is the task of inferring the type of a place from where a social media post was shared. Inferring a POI's type is useful for studies in computational social science including sociolinguistics, geosemiotics, and cultural geography, and has applications in geosocial networking technologies such as recommendation and visualization systems. Prior efforts in POI type prediction focus solely on text, without taking visual information into account. However in reality, the variety of modalities, as well as their semiotic relationships with one another, shape communication and interactions in social media. This paper presents a study on POI type prediction using multimodal information from text and images available at posting time. For that purpose, we enrich a currently available data set for POI type prediction with the images that accompany the text messages. Our proposed method extracts relevant information from each modality to effectively capture interactions between text and image achieving a macro F1 of 47.21 across eight categories significantly outperforming the state-of-the-art method for POI type prediction based on text-only methods. Finally, we provide a detailed analysis to shed light on cross-modal interactions and the limitations of our best performing model.
Code (1)
Tasks
PredictionType predictionVocal Bursts Type PredictionSimilar Papers 제목 키워드 기반
A unified framework of predicting binary interestingness of images based on discriminant correlation analysis and multiple kernel learning
In the modern content-based image retrieval systems, there is an increasingly interest in constructing a computationally effective model to predict the interestingness of images since the measure of image interestingness…
Content-Based Image RetrievalImage RetrievalRetrievalPoint-of-Interest Type Inference from Social Media Text
Physical places help shape how we perceive the experiences we have there. For the first time, we study the relationship between social media text and the type of the place from where it was posted, whether a park, restau…
Recommendation SystemsVocal Bursts Type PredictionDocReal: Robust Document Dewarping of Real-Life Images via Attention-Enhanced Control Point Prediction
Document image dewarping is a crucial task in computer vision with numerous practical applications. The control point method, as a popular image dewarping approach, has attracted attention due to its simplicity and effic…
Optical Character Recognition (OCR)Popular LLMs Amplify Race and Gender Disparities in Human Mobility
As large language models (LLMs) are increasingly applied in areas influencing societal outcomes, it is critical to understand their tendency to perpetuate and amplify biases. This study investigates whether LLMs exhibit …
UnsuperPoint: End-to-end Unsupervised Interest Point Detector and Descriptor
It is hard to create consistent ground truth data for interest points in natural images, since interest points are hard to define clearly and consistently for a human annotator. This makes interest point detectors non-tr…
Homography Estimation