paper-with-me

홈 › Papers

An Effective Algorithm for Learning Single Occurrence Regular Expressions with Interleaving

2019-06-05 · Yeting Li, Haiming Chen, Xiaolan Zhang, Lingqi Zhang

The advantages offered by the presence of a schema are numerous. However, many XML documents in practice are not accompanied by a (valid) schema, making schema inference an attractive research problem. The fundamental task in XML schema learning is inferring restricted subclasses of regular expressions. Most previous work either lacks support for interleaving or only has limited support for interleaving. In this paper, we first propose a new subclass Single Occurrence Regular Expressions with Interleaving (SOIRE), which has unrestricted support for interleaving. Then, based on single occurrence automaton and maximum independent set, we propose an algorithm iSOIRE to infer SOIREs. Finally, we further conduct a series of experiments on real datasets to evaluate the effectiveness of our work, comparing with both ongoing learning algorithms in academia and industrial tools in real-world. The results reveal the practicability of SOIRE and the effectiveness of iSOIRE, showing the high preciseness and conciseness of our work.

📄 PDF Abstract BibTeX arXiv:1906.02074

Code (0)

등록된 구현이 없습니다.

Tasks

valid

Similar Papers 제목 키워드 기반

A Noise-tolerant Differentiable Learning Approach for Single Occurrence Regular Expression with Interleaving

2022-12-01 · Rongzhen Ye, Tianqu Zhuang, Hai Wan, Jianfeng Du 외

We study the problem of learning a single occurrence regular expression with interleaving (SOIRE) from a set of text strings possibly with noise. SOIRE fully supports interleaving and covers a large portion of regular ex…

Hamm-Grams: An Algorithm for Mining Regular Expressions of Bytes

2026-07-01 · Derek Everett, Edward Raff, James Holt arxiv

Malware poses a critical and ever-evolving threat, and robust and effective systems for detecting and classifying malware are of essential importance. $n$-grams features are among the common static features used in effec…

Malware Classification

UMDuluth-CS8761 at SemEval-2018 Task9: Hypernym Discovery using Hearst Patterns, Co-occurrence frequencies and Word Embeddings

2018-06-01 · SEMEVAL 2018 6 · Arshia Zernab Hassan, Manikya Swathi Vallabhajosyula, Ted Pedersen

Hypernym Discovery is the task of identifying potential hypernyms for a given term. A hypernym is a more generalized word that is super-ordinate to more specific words. This paper explores several approaches that rely on…

Hypernym DiscoveryWord Embeddings

UMDuluth-CS8761 at SemEval-2018 Task 9: Hypernym Discovery using Hearst Patterns, Co-occurrence frequencies and Word Embeddings

2018-05-25 · Arshia Z. Hassan, Manikya S. Vallabhajosyula, Ted Pedersen

Hypernym Discovery is the task of identifying potential hypernyms for a given term. A hypernym is a more generalized word that is super-ordinate to more specific words. This paper explores several approaches that rely on…

Hypernym DiscoveryWord Embeddings

Learning the Simplicity of Scattering Amplitudes

2024-08-08 · Clifford Cheung, Aurélien Dersy, Matthew D. Schwartz

The simplification and reorganization of complex expressions lies at the core of scientific progress, particularly in theoretical high-energy physics. This work explores the application of machine learning to a particula…

Contrastive LearningDecoder