SEDAR: a Large Scale French-English Financial Domain Parallel Corpus
This paper describes the acquisition, preprocessing and characteristics of SEDAR, a large scale English-French parallel corpus for the financial domain. Our extensive experiments on machine translation show that SEDAR is essential to obtain good performance on finance. We observe a large gain in the performance of machine translation systems trained on SEDAR when tested on finance, which makes SEDAR suitable to study domain adaptation for neural machine translation. The first release of the corpus comprises 8.6 million high quality sentence pairs that are publicly available for research at https://github.com/autorite/sedar-bitext.
Code (1)
Tasks
Domain AdaptationMachine TranslationSentenceTranslationSimilar Papers 제목 키워드 기반
NUIG at the FinSBD Task: Sentence Boundary Detection for Noisy Financial PDFs in English and French
Discovering material information using hierarchical Reformer model on financial regulatory filings
Most applications of machine learning for finance are related to forecasting tasks for investment decisions. Instead, we aim to promote a better understanding of financial markets with machine learning techniques. Levera…
BIG-bench Machine LearningSentenceThe Financial Document Structure Extraction Shared Task (FinTOC 2022)
This paper describes the FinTOC-2022 Shared Task on the structure extraction from financial documents, its participants results and their findings. This shared task was organized as part of The 4th Financial Narrative Pr…
SedarEval: Automated Evaluation using Self-Adaptive Rubrics
The evaluation paradigm of LLM-as-judge gains popularity due to its significant reduction in human labor and time costs. This approach utilizes one or more large language models (LLMs) to assess the quality of outputs fr…
Logical ReasoningGREYC@FinTOC-2022: Handling Document Layout and Structure in Native PDF Bundle of Documents
n this paper, we present our contribution to the FinTOC-2022 Shared Task “Financial Document Structure Extraction”. We participated in the three tracks dedicated to English, French and Spanish document processing. Our ma…
Boundary Detection