paper-with-me

Papers

GitHub Repository Complexity Leads to Diminished Web Archive Availability

2025-05-21 · David Calano, Michele C. Weigle, Michael L. Nelson

Software is often developed using versioned controlled software, such as Git, and hosted on centralized Web hosts, such as GitHub and GitLab. These Web hosted software repositories are made available to users in the form of traditional HTML Web pages for each source file and directory, as well as a presentational home page and various descriptive pages. We examined more than 12,000 Web hosted Git repository project home pages, primarily from GitHub, to measure how well their presentational components are preserved in the Internet Archive, as well as the source trees of the collected GitHub repositories to assess the extent to which their source code has been preserved. We found that more than 31% of the archived repository home pages examined exhibited some form of minor page damage and 1.6% exhibited major page damage. We also found that of the source trees analyzed, less than 5% of their source files were archived, on average, with the majority of repositories not having source files saved in the Internet Archive at all. The highest concentration of archived source files available were those linked directly from repositories' home pages at a rate of 14.89% across all available repositories and sharply dropping off at deeper levels of a repository's directory tree.

📄 PDF Abstract BibTeX arXiv:2505.15042

Code (0)

등록된 구현이 없습니다.

Tasks

Descriptive

Similar Papers 제목 키워드 기반

The Multiverse of Time Series Machine Learning: an Archive for Multivariate Time Series Classification

2026-03-20 · Matthew Middlehurst, Aiden Rushbrooke, Ali Ismail-Fawaz, Maxime Devanne 외 arxiv

Time series machine learning (TSML) is a growing research field that spans a wide range of tasks. The popularity of established tasks such as classification, clustering, and extrinsic regression has, in part, been driven…

Time Series Classification

Repo2Vec: A Comprehensive Embedding Approach for Determining Repository Similarity

2021-07-11 · Md Omar Faruk Rokon, Pei Yan, Risul Islam, Michalis Faloutsos

How can we identify similar repositories and clusters among a large online archive, such as GitHub? Determiningrepository similarity is an essential building block in studying the dynamics and the evolution of such softw…

Repository-Level Prompt Generation for Large Language Models of Code

2022-06-26 · Disha Shrivastava, Hugo Larochelle, Daniel Tarlow

With the success of large language models (LLMs) of code and their use as code assistants (e.g. Codex used in GitHub Copilot), techniques for introducing domain-specific knowledge in the prompt design process become impo…

A Comparison of Neuroelectrophysiology Databases

2023-06-26 · Priyanka Subash, Alex Gray, Misque Boswell, Samantha L. Cohen 외

As data sharing has become more prevalent, three pillars - archives, standards, and analysis tools - have emerged as critical components in facilitating effective data sharing and collaboration. This paper compares four …

Data Integration

NEMAR: An open access data, tools, and compute resource operating on NeuroElectroMagnetic data

2022-03-04 · Arnaud Delorme, Dung Truong, Choonhan Youn, Subha Sivagnanam 외

To take advantage of recent and ongoing advances in large-scale computational methods, and to preserve the scientific data created by publicly funded research projects, data archives must be created as well as standards …

EEGElectroencephalogram (EEG)