paper-with-me

홈 › Papers

CWRCzech: 100M Query-Document Czech Click Dataset and Its Application to Web Relevance Ranking

2024-05-31 · Josef Vonášek, Milan Straka, Rostislav Krč, Lenka Lasoňová, Ekaterina Egorova, Jana Straková, Jakub Náplava

We present CWRCzech, Click Web Ranking dataset for Czech, a 100M query-document Czech click dataset for relevance ranking with user behavior data collected from search engine logs of Seznam$.$cz. To the best of our knowledge, CWRCzech is the largest click dataset with raw text published so far. It provides document positions in the search results as well as information about user behavior: 27.6M clicked documents and 10.8M dwell times. In addition, we also publish a manually annotated Czech test for the relevance task, containing nearly 50k query-document pairs, each annotated by at least 2 annotators. Finally, we analyze how the user behavior data improve relevance ranking and show that models trained on data automatically harnessed at sufficient scale can surpass the performance of models trained on human annotated data. CWRCzech is published under an academic non-commercial license and is available to the research community at https://github.com/seznam/CWRCzech.

📄 PDF Abstract BibTeX arXiv:2405.20994

Code (1)

seznam/cwrczech 공식 구현

Similar Papers 제목 키워드 기반

Siamese BERT-based Model for Web Search Relevance Ranking Evaluated on a New Czech Dataset

2021-12-03 · Matěj Kocián, Jakub Náplava, Daniel Štancl, Vladimír Kadlec

Web search engines focus on serving highly relevant results within hundreds of milliseconds. Pre-trained language transformer models such as BERT are therefore hard to use in this scenario due to their high computational…

Document RankingLanguage ModelingSmall Language Model

MIRA: Leveraging Multi-Intention Co-click Information in Web-scale Document Retrieval using Deep Neural Networks

2020-07-03 · Yusi Zhang, Chuanjie Liu, Angen Luo, Hui Xue 외

We study the problem of deep recall model in industrial web search, which is, given a user query, retrieve hundreds of most relevance documents from billions of candidates. The common framework is to train two encoding m…

Graph AttentionRetrieval

Query Suggestion for Click-Absent Queries in Enterprise Search

2021-12-28 · Gizem Gezici

Creating alternative queries, also known as query suggestion, has been proved to be helpful on improving users' search experience. Owing to the suggestions, users could retrieve their information need more quickly and ac…

ORCAS: 18 Million Clicked Query-Document Pairs for Analyzing Search

2020-06-09 · Nick Craswell, Daniel Campos, Bhaskar Mitra, Emine Yilmaz 외

Users of Web search engines reveal their information needs through queries and clicks, making click logs a useful asset for information retrieval. However, click logs have not been publicly released for academic use, bec…

Information RetrievalRetrieval

CzechDocs: A Multiway Parallel Dataset of Formatted Documents for Minority Languages in Czechia

2026-06-18 · Josef Jon, Ondřej Bojar arxiv

We present CzechDocs, a multiway parallel dataset of formatted documents (HTML, DOCX, and PDF) covering Czech and minority languages used in Czechia-primarily Ukrainian and English, with smaller portions of Vietnamese, R…

Machine Translation