A methodology for the extraction of information about the usage of formulaic expressions in scientific texts
In this paper, we present a methodology for the extraction of formulaic expressions, which goes beyond the mere extraction of candidate patterns. Using a pipeline we are able to extract information about the usage of formulaic expressions automatically from text corpora. According to Biber and Barbieri (2007) formulaic expressions are important building blocks of discourse in spoken and written registers. The automatic extraction procedure can help to investigate the usage and function of these recurrent patterns in different registers and domains. Formulaic expressions are commonplace not only in every- day language but also in scientific writing. Patterns such as 'in this paper', 'the number of', 'on the basis of' are often used by scientists to convey research interests, the theoretical basis of their studies, results of experiments, sci- entific findings as well as conclusions and are used as dis- course organizers. For Hyland (2008) they help to shape meanings in specific context and contribute to our sense of coherence in a text. We are interested in: (i) which and what type of formulaic expressions are used in scientific texts? (ii) the distribution of formulaic expression across different scien- tific disciplines, (iii) where do formulaic expressions occur within a text?
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Extraction and Evaluation of Formulaic Expressions Used in Scholarly Papers
Formulaic expressions, such as 'in this paper we propose', are helpful for authors of scholarly papers because they convey communicative functions; in the above, it is showing the aim of this paper'. Thus, resources of f…
DiversitySentenceSynergistic Formulaic Alpha Generation for Quantitative Trading based on Reinforcement Learning
Mining of formulaic alpha factors refers to the process of discovering and developing specific factors or indicators (referred to as alpha factors) for quantitative trading in stock market. To efficiently discover alpha …
reinforcement-learningReinforcement Learning (RL)An Information-Theoretic Approach to Identifying Formulaic Clusters in Textual Data
Texts, whether literary or historical, exhibit structural and stylistic patterns shaped by their purpose, authorship, and cultural context. Formulaic texts, characterized by repetition and constrained expression, tend to…
KPIs-Based Clustering and Visualization of HPC jobs: a Feature Reduction Approach
High-Performance Computing (HPC) systems need to be constantly monitored to ensure their stability. The monitoring systems collect a tremendous amount of data about different parameters or Key Performance Indicators (KPI…
ClusteringCPUManagementTime SeriesUsing CollGram to Compare Formulaic Language in Human and Neural Machine Translation
A comparison of formulaic sequences in human and neural machine translation of quality newspaper articles shows that neural machine translations contain less lower-frequency, but strongly-associated formulaic sequences, …
ArticlesMachine TranslationTranslation