paper-with-me

홈 › Papers

IndoSum: A New Benchmark Dataset for Indonesian Text Summarization

2018-10-12 · Kemal Kurniawan, Samuel Louvan

Automatic text summarization is generally considered as a challenging task in the NLP community. One of the challenges is the publicly available and large dataset that is relatively rare and difficult to construct. The problem is even worse for low-resource languages such as Indonesian. In this paper, we present IndoSum, a new benchmark dataset for Indonesian text summarization. The dataset consists of news articles and manually constructed summaries. Notably, the dataset is almost 200x larger than the previous Indonesian summarization dataset of the same domain. We evaluated various extractive summarization approaches and obtained encouraging results which demonstrate the usefulness of the dataset and provide baselines for future research. The code and the dataset are available online under permissive licenses.

📄 PDF Abstract BibTeX arXiv:1810.05334

Code (1)

kata-ai/indosum 공식 구현 tf

Tasks

ArticlesExtractive SummarizationText Summarization

Similar Papers 제목 키워드 기반

Investigating Text Shortening Strategy in BERT: Truncation vs Summarization

2024-03-19 · Mirza Alim Mutasodirin, Radityo Eko Prasojo

The parallelism of Transformer-based models comes at the cost of their input max-length. Some studies proposed methods to overcome this limitation, but none of them reported the effectiveness of summarization as an alter…

ArticlesDocument SummarizationExtractive Summarizationtext-classification+1

Liputan6: A Large-scale Indonesian Dataset for Text Summarization

2020-11-02 · Asian Chapter of the Association for Computational Linguistics 2020 · Fajri Koto, Jey Han Lau, Timothy Baldwin

In this paper, we introduce a large-scale Indonesian summarization dataset. We harvest articles from Liputan6.com, an online news portal, and obtain 215,827 document-summary pairs. We leverage pre-trained language models…

Abstractive Text SummarizationArticlesText Summarization

LipKey: A Large-Scale News Dataset with Abstractive Keyphrases and Their Benefits for Summarization

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Summaries, keyphrases, and titles are different ways of concisely capturing the content of a document. While most previous work has addressed them separately, in this work, we jointly use the three elements via multi-tas…

Document Summarization

MSVD-Indonesian: A Benchmark for Multimodal Video-Text Tasks in Indonesian

2023-06-20 · Willy Fitra Hendria

Multimodal learning on video and text data has been receiving growing attention from many researchers in various research tasks, including text-to-video retrieval, video-to-text retrieval, and video captioning. Although …

Cross-Lingual TransferRetrievalText RetrievalText to Video Retrieval+5

A Publicly Available Indonesian Corpora for Automatic Abstractive and Extractive Chat Summarization

2016-05-01 · LREC 2016 5 · Fajri Koto

In this paper we report our effort to construct the first ever Indonesian corpora for chat summarization. Specifically, we utilized documents of multi-participant chat from a well known online instant messaging applicati…