paper-with-me

홈 › Papers

ISNA-Set: A novel English Corpus of Iran NEWS

2018-08-21 · Mohammad Kamel, Hadi Sadoghi-Yazdi

News agencies publish news on their websites all over the world. Moreover, creating novel corpuses is necessary to bring natural processing to new domains. Textual processing of online news is challenging in terms of the strategy of collecting data, the complex structure of news websites, and selecting or designing suitable algorithms for processing these types of data. Despite the previous works which focus on creating corpuses for Iran news in Persian, in this paper, we introduce a new corpus for English news of a national news agency. ISNA-Set is a new dataset of English news of Iranian Students News Agency (ISNA), as one of the most famous news agencies in Iran. We statistically analyze the data and the sentiments of news, and also extract entities and part-of-speech tagging.

📄 PDF Abstract BibTeX arXiv:1808.07046

Code (0)

등록된 구현이 없습니다.

Tasks

Part-Of-Speech Tagging

Similar Papers 제목 키워드 기반

Bianet: A Parallel News Corpus in Turkish, Kurdish and English

2018-05-14 · Duygu Ataman

We present a new open-source parallel corpus consisting of news articles collected from the Bianet magazine, an online newspaper that publishes Turkish news, often along with their translations in English and Kurdish. In…

ArticlesMachine TranslationTranslation

Massively Multi-Lingual Event Understanding: Extraction, Visualization, and Search

2023-05-17 · Chris Jenkins, Shantanu Agarwal, Joel Barry, Steven Fincke 외

In this paper, we present ISI-Clear, a state-of-the-art, cross-lingual, zero-shot event extraction system and accompanying user interface for event visualization & search. Using only English training data, ISI-Clear make…

Event ExtractionNatural Language QueriesZero-shot Event Extraction

Neural Machine Translation System using a Content-equivalently Translated Parallel Corpus for the Newswire Translation Tasks at WAT 2019

2019-11-01 · WS 2019 11 · Hideya Mino, Hitoshi Ito, Isao Goto, Ichiro Yamada 외

This paper describes NHK and NHK Engineering System (NHK-ES){'}s submission to the newswire translation tasks of WAT 2019 in both directions of Japanese→English and English→Japanese. In addition to the JIJI Corpus that w…

Machine TranslationSentenceTranslation

Automatic Parallel Corpus Creation for Hindi-English News Translation Task

2019-01-24 · Aditya Kumar Pathak, Priyankit Acharya, Dilpreet Kaur, Rakesh Chandra Balabantaray

The parallel corpus for multilingual NLP tasks, deep learning applications like Statistical Machine Translation Systems is very important. The parallel corpus of Hindi-English language pair available for news translation…

Machine TranslationMultilingual NLPTranslation

MiRANews: Dataset and Benchmarks for Multi-Resource-Assisted News Summarization

2021-09-22 · Findings (EMNLP) 2021 11 · Xinnuo Xu, Ondřej Dušek, Shashi Narayan, Verena Rieser 외

One of the most challenging aspects of current single-document news summarization is that the summary often contains 'extrinsic hallucinations', i.e., facts that are not present in the source document, which are often de…

ArticlesDocument SummarizationMulti-Document SummarizationNews Summarization+1