paper-with-me

홈 › Papers

Billions of Parameters Are Worth More Than In-domain Training Data: A case study in the Legal Case Entailment Task

2022-05-30 · Guilherme Moraes Rosa, Luiz Bonifacio, Vitor Jeronymo, Hugo Abonizio, Roberto Lotufo, Rodrigo Nogueira

Recent work has shown that language models scaled to billions of parameters, such as GPT-3, perform remarkably well in zero-shot and few-shot scenarios. In this work, we experiment with zero-shot models in the legal case entailment task of the COLIEE 2022 competition. Our experiments show that scaling the number of parameters in a language model improves the F1 score of our previous zero-shot result by more than 6 points, suggesting that stronger zero-shot capability may be a characteristic of larger models, at least for this task. Our 3B-parameter zero-shot model outperforms all models, including ensembles, in the COLIEE 2021 test set and also achieves the best performance of a single model in the COLIEE 2022 competition, second only to the ensemble composed of the 3B model itself and a smaller version of the same model. Despite the challenges posed by large language models, mainly due to latency constraints in real-time applications, we provide a demonstration of our zero-shot monoT5-3b model being used in production as a search engine, including for legal documents. The code for our submission and the demo of our system are available at https://github.com/neuralmind-ai/coliee and https://neuralsearchx.neuralmind.ai, respectively.

📄 PDF Abstract BibTeX arXiv:2205.15172

Code (1)

neuralmind-ai/coliee 공식 구현

Tasks

Language Modelling

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
{Dispute@FaQ-s}How to file a dispute with Expedia? How to file a dispute with Expedia? To file a complaint against Expedia, first try contacting their customer service directly. You can reach them by phone at…
Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Adam 설명 없음
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

A Few More Examples May Be Worth Billions of Parameters

2021-10-08 · Yuval Kirstain, Patrick Lewis, Sebastian Riedel, Omer Levy

We investigate the dynamics of increasing the number of model parameters versus the number of labeled examples across a wide variety of tasks. Our exploration reveals that while scaling parameters consistently yields per…

Extractive Question-AnsweringMultiple-choiceOpen-Ended Question AnsweringQuestion Answering

XuanYuan 2.0: A Large Chinese Financial Chat Model with Hundreds of Billions Parameters

2023-05-19 · Xuanyu Zhang, Qing Yang, Dongliang Xu

In recent years, pre-trained language models have undergone rapid development with the emergence of large-scale models. However, there is a lack of open-sourced chat models specifically designed for the Chinese language,…

Multilingual Sentence-T5: Scalable Sentence Encoders for Multilingual Applications

2024-03-26 · Chihiro Yano, Akihiko Fukuchi, Shoko Fukasawa, Hideyuki Tachibana 외

Prior work on multilingual sentence embedding has demonstrated that the efficient use of natural language inference (NLI) data to build high-performance models can outperform conventional methods. However, the potential …

Natural Language InferenceSentenceSentence EmbeddingSentence-Embedding

Scaling Laws of Graph Neural Networks for Atomistic Materials Modeling

2025-04-10 · Chaojian Li, Zhifan Ye, Massimiliano Lupo Pasini, Jong Youl Choi 외

Atomistic materials modeling is a critical task with wide-ranging applications, from drug discovery to materials science, where accurate predictions of the target material property can lead to significant advancements in…

Drug Discoveryscientific discovery

Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations

2024-02-27 · Jiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang 외

Large-scale recommendation systems are characterized by their reliance on high cardinality, heterogeneous features and the need to handle tens of billions of user actions on a daily basis. Despite being trained on huge v…

Recommendation Systems