paper-with-me

홈 › Papers

Combining Visual and Textual Features for Semantic Segmentation of Historical Newspapers

2020-02-14 · Raphaël Barman, Maud Ehrmann, Simon Clematide, Sofia Ares Oliveira, Frédéric Kaplan

The massive amounts of digitized historical documents acquired over the last decades naturally lend themselves to automatic processing and exploration. Research work seeking to automatically process facsimiles and extract information thereby are multiplying with, as a first essential step, document layout analysis. If the identification and categorization of segments of interest in document images have seen significant progress over the last years thanks to deep learning techniques, many challenges remain with, among others, the use of finer-grained segmentation typologies and the consideration of complex, heterogeneous documents such as historical newspapers. Besides, most approaches consider visual features only, ignoring textual signal. In this context, we introduce a multimodal approach for the semantic segmentation of historical newspapers that combines visual and textual features. Based on a series of experiments on diachronic Swiss and Luxembourgish newspapers, we investigate, among others, the predictive power of visual and textual features and their capacity to generalize across time and sources. Results show consistent improvement of multimodal models in comparison to a strong visual baseline, as well as better robustness to high material variance.

📄 PDF Abstract BibTeX arXiv:2002.06144

Code (3)

dhlab-epfl/dhSegment-text 공식 구현 tf
dhlab-epfl/dhSegment-text-torch 공식 구현 pytorch
raphaelBarman/dhSegment 공식 구현 tf

Tasks

Document Layout AnalysisSemantic Segmentation

Similar Papers 제목 키워드 기반

DFR: A Decompose-Fuse-Reconstruct Framework for Multi-Modal Few-Shot Segmentation

2025-07-22 · Shuai Chen, Fanman Meng, Xiwei Zhang, Haoran Wei 외 arxiv

This paper presents DFR (Decompose, Fuse and Reconstruct), a novel framework that addresses the fundamental challenge of effectively utilizing multi-modal guidance in few-shot segmentation (FSS). While existing approache…

Contrastive Learning

A Multimodal Approach Combining Structural and Cross-domain Textual Guidance for Weakly Supervised OCT Segmentation

2024-11-19 · Jiaqi Yang, Nitish Mehta, Xiaoling Hu, Chao Chen 외

Accurate segmentation of Optical Coherence Tomography (OCT) images is crucial for diagnosing and monitoring retinal diseases. However, the labor-intensive nature of pixel-level annotation limits the scalability of superv…

DescriptiveDiagnosticSegmentationSemantic Segmentation+2

BEVANet: Bilateral Efficient Visual Attention Network for Real-Time Semantic Segmentation

2025-08-10 · Ping-Mao Huang, I-Tien Chao, Ping-Chia Huang, Jia-Wei Liao 외 arxiv

Real-time semantic segmentation presents the dual challenge of designing efficient architectures that capture large receptive fields for semantic understanding while also refining detailed contours. Vision transformers m…

Real-Time Semantic Segmentation

Lifting GIS Maps into Strong Geometric Context for Scene Understanding

2015-07-14 · Raúl Díaz, Minhaeng Lee, Jochen Schubert, Charless C. Fowlkes

Contextual information can have a substantial impact on the performance of visual tasks such as semantic segmentation, object detection, and geometric estimation. Data stored in Geographic Information Systems (GIS) offer…

Depth Estimationobject-detectionObject DetectionScene Understanding+2

DCP-CLIP:A Coarse-to-Fine Framework for Open-Vocabulary Semantic Segmentation with Dual Interaction

2026-03-14 · Jing Wang, Huimin Shi, Quan Zhou, Qibo Liu 외 arxiv

The recent years have witnessed the remarkable development for open-vocabulary semantic segmentation (OVSS) using visual-language foundation models, yet still suffer from following fundamental challenges: (1) insufficien…

Semantic Segmentation