paper-with-me

홈 › Papers

The Web Is Your Oyster -- Knowledge-Intensive NLP against a Very Large Web Corpus

2021-12-18 · Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Dmytro Okhonko, Samuel Broscheit, Gautier Izacard, Patrick Lewis, Barlas Oğuz, Edouard Grave, Wen-tau Yih, Sebastian Riedel

In order to address increasing demands of real-world applications, the research for knowledge-intensive NLP (KI-NLP) should advance by capturing the challenges of a truly open-domain environment: web-scale knowledge, lack of structure, inconsistent quality and noise. To this end, we propose a new setup for evaluating existing knowledge intensive tasks in which we generalize the background corpus to a universal web snapshot. We investigate a slate of NLP tasks which rely on knowledge - either factual or common sense, and ask systems to use a subset of CCNet - the Sphere corpus - as a knowledge source. In contrast to Wikipedia, otherwise a common background corpus in KI-NLP, Sphere is orders of magnitude larger and better reflects the full diversity of knowledge on the web. Despite potential gaps in coverage, challenges of scale, lack of structure and lower quality, we find that retrieval from Sphere enables a state of the art system to match and even outperform Wikipedia-based models on several tasks. We also observe that while a dense index can outperform a sparse BM25 baseline on Wikipedia, on Sphere this is not yet possible. To facilitate further research and minimise the community's reliance on proprietary, black-box search engines, we share our indices, evaluation metrics and infrastructure.

📄 PDF Abstract BibTeX arXiv:2112.09924

Code (2)

facebookresearch/distributed-faiss 공식 구현
facebookresearch/sphere 공식 구현 pytorch

Tasks

Common Sense ReasoningRetrieval

Methods 이 논문이 사용한 방법론

CCNet Criss-Cross Network (CCNet) aims to obtain full-image contextual information in an effective and efficient way. Concretely, for each pixel, a novel criss-cross attention…

Similar Papers 제목 키워드 기반

The Web Can Be Your Oyster for Improving Large Language Models

2023-05-18 · Junyi Li, Tianyi Tang, Wayne Xin Zhao, Jingyuan Wang 외

Large language models (LLMs) encode a large amount of world knowledge. However, as such knowledge is frozen at the time of model training, the models become static and limited by the training data at that time. In order …

RetrievalWorld Knowledge

OysterNet: Enhanced Oyster Detection Using Simulation

2022-09-16 · Xiaomin Lin, Nitin J. Sanket, Nare Karapetyan, Yiannis Aloimonos

Oysters play a pivotal role in the bay living ecosystem and are considered the living filters for the ocean. In recent years, oyster reefs have undergone major devastation caused by commercial over-harvesting, requiring …

Is AI currently capable of identifying wild oysters? A comparison of human annotators against the AI model, ODYSSEE

2025-05-06 · Brendan Campbell, Alan Williams, Kleio Baxevani, Alyssa Campbell 외

Oysters are ecologically and commercially important species that require frequent monitoring to track population demographics (e.g. abundance, growth, mortality). Current methods of monitoring oyster reefs often require …

Modeling Oyster Reef Reproductive Sustainability: Analyzing Gamete Viability, Hydrodynamics, and Reef Structure to Facilitate Restoration of $\textit{Crassostrea virginica}$

2021-01-13 · Justin Weissberg, Vinny Pagano

The eastern oyster is a keystone species and ecosystem engineer. However, restoration efforts of wild oysters are often unsuccessful, in that they do not produce a robust population of oysters that are able to successful…

ODYSSEE: Oyster Detection Yielded by Sensor Systems on Edge Electronics

2024-09-11 · Xiaomin Lin, Vivek Mange, Arjun Suresh, Bernhard Neuberger 외

Oysters are a vital keystone species in coastal ecosystems, providing significant economic, environmental, and cultural benefits. As the importance of oysters grows, so does the relevance of autonomous systems for their …

object-detectionObject Detection