Company Similarity using Large Language Models
Identifying companies with similar profiles is a core task in finance with a wide range of applications in portfolio construction, asset pricing and risk attribution. When a rigorous definition of similarity is lacking, financial analysts usually resort to 'traditional' industry classifications such as Global Industry Classification System (GICS) which assign a unique category to each company at different levels of granularity. Due to their discrete nature, though, GICS classifications do not allow for ranking companies in terms of similarity. In this paper, we explore the ability of pre-trained and finetuned large language models (LLMs) to learn company embeddings based on the business descriptions reported in SEC filings. We show that we can reproduce GICS classifications using the embeddings as features. We also benchmark these embeddings on various machine learning and financial metrics and conclude that the companies that are similar according to the embeddings are also similar in terms of financial performance metrics including return correlation.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
CompanyKG: A Large-Scale Heterogeneous Graph for Company Similarity Quantification
In the investment industry, it is often essential to carry out fine-grained company similarity quantification for a range of purposes, including market mapping, competitor analysis, and mergers and acquisitions. We propo…
BenchmarkingRetrievalInterpretable Company Similarity with Sparse Autoencoders
Determining company similarity is a vital task in finance, underpinning risk management, hedging, and portfolio diversification. Practitioners often rely on sector and industry classifications such as SIC and GICS codes …
Large Language ModelSemantic SimilaritySemantic Textual SimilarityFrom words to connections: Word use similarity as an honest signal conducive to employees' digital communication
Bringing together considerations from three research trends (honest signals of collaboration, socio-semantic networks and homophily theory), we hypothesise that word use similarity and having similar social network posit…
Named entity recognition using GPT for identifying comparable companies
For both public and private firms, comparable companies' analysis is widely used as a method for company valuation. In particular, the method is of great value for valuation of private equity companies. The several appro…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NERWords are all you need? Language as an approximation for human similarity judgments
Human similarity judgments are a powerful supervision signal for machine learning applications based on techniques such as contrastive learning, information retrieval, and model alignment, but classical methods for colle…
AllContrastive LearningInformation RetrievalRetrieval+1