paper-with-me

Papers

OpenEDGAR: Open Source Software for SEC EDGAR Analysis

2018-06-13 · Michael J Bommarito II, Daniel Martin Katz, Eric M Detterman

OpenEDGAR is an open source Python framework designed to rapidly construct research databases based on the Electronic Data Gathering, Analysis, and Retrieval (EDGAR) system operated by the US Securities and Exchange Commission (SEC). OpenEDGAR is built on the Django application framework, supports distributed compute across one or more servers, and includes functionality to (i) retrieve and parse index and filing data from EDGAR, (ii) build tables for key metadata like form type and filer, (iii) retrieve, parse, and update CIK to ticker and industry mappings, (iv) extract content and metadata from filing documents, and (v) search filing document contents. OpenEDGAR is designed for use in both academic research and industrial applications, and is distributed under MIT License at https://github.com/LexPredict/openedgar.

📄 PDF Abstract BibTeX arXiv:1806.04973

Code (1)

LexPredict/openedgar 공식 구현

Tasks

Retrieval

Similar Papers 제목 키워드 기반

vEDGAR -- Can CARLA Do HiL?

2025-12-09 · Nils Gehrke, David Brecht, Dominik Kulmer, Dheer Patel 외 arxiv

Simulation offers advantages throughout the development process of automated driving functions, both in research and product development. Common open-source simulators like CARLA are extensively used in training, evaluat…

EDGAR-CORPUS: Billions of Tokens Make The World Go Round

2021-09-29 · EMNLP (ECONLP) 2021 11 · Lefteris Loukas, Manos Fergadiotis, Ion Androutsopoulos, Prodromos Malakasiotis

We release EDGAR-CORPUS, a novel corpus comprising annual reports from all the publicly traded companies in the US spanning a period of more than 25 years. To the best of our knowledge, EDGAR-CORPUS is the largest financ…

Word Embeddings

The Stanford EDGAR Filings Dataset: Reconstructing U.S. Corporate and Financial Disclosures into Layout-Faithful and Token-Efficient Pretraining Data

2026-06-16 · Nick Bettencourt, Xiaowei Ding, Kay Giesecke arxiv

As high-quality public web corpora become increasingly exhausted, clean long-context documents have become a scarce and expensive source of training data for large language models (LLMs). Existing long-context corpora ar…

Explainable Enrichment-Driven GrAph Reasoner (EDGAR) for Large Knowledge Graphs with Applications in Drug Repurposing

2024-09-27 · Olawumi Olasunkanmi, Evan Morris, Yaphet Kebede, Harlin Lee 외

Knowledge graphs (KGs) represent connections and relationships between real-world entities. We propose a link prediction framework for KGs named Enrichment-Driven GrAph Reasoner (EDGAR), which infers new edges by mining …

Knowledge GraphsLink Prediction

KPI-EDGAR: A Novel Dataset and Accompanying Metric for Relation Extraction from Financial Documents

2022-10-17 · Tobias Deußer, Syed Musharraf Ali, Lars Hillebrand, Desiana Nurchalifah 외

We introduce KPI-EDGAR, a novel dataset for Joint Named Entity Recognition and Relation Extraction building on financial reports uploaded to the Electronic Data Gathering, Analysis, and Retrieval (EDGAR) system, where th…

BenchmarkingJoint Entity and Relation Extractionnamed-entity-recognitionNamed Entity Recognition+4