paper-with-me

홈 › Papers

Textual Characteristics for Language Engineering

2012-05-01 · LREC 2012 5 · Mathias Bank, Robert Remus, Martin Schierle

Language statistics are widely used to characterize and better understand language. In parallel, the amount of text mining and information retrieval methods grew rapidly within the last decades, with many algorithms evaluated on standardized corpora, often drawn from newspapers. However, up to now there were almost no attempts to link the areas of natural language processing and language statistics in order to properly characterize those evaluation corpora, and to help others to pick the most appropriate algorithms for their particular corpus. We believe no results in the field of natural language processing should be published without quantitatively describing the used corpora. Only then the real value of proposed methods can be determined and the transferability to corpora originating from different genres or domains can be estimated. We lay ground for a language engineering process by gathering and defining a set of textual characteristics we consider valuable with respect to building natural language processing systems. We carry out a case study for the analysis of automotive repair orders and explicitly call upon the scientific community to provide feedback and help to establish a good practice of corpus-aware evaluations.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalRetrieval

Similar Papers 제목 키워드 기반

MMCR: Advancing Visual Language Model in Multimodal Multi-Turn Contextual Reasoning

2025-03-24 · Dawei Yan, Yang Li, Qing-Guo Chen, Weihua Luo 외

Compared to single-turn dialogue, multi-turn dialogue involving multiple images better aligns with the needs of real-world human-AI interactions. Additionally, as training data, it provides richer contextual reasoning in…

DiagnosticLanguage ModelingLanguage ModellingPrompt Engineering

Automated Essay Scoring with Discourse-Aware Neural Models

2019-08-01 · WS 2019 8 · Farah Nadeem, Huy Nguyen, Yang Liu, Mari Ostendorf

Automated essay scoring systems typically rely on hand-crafted features to predict essay quality, but such systems are limited by the cost of feature engineering. Neural networks offer an alternative to feature engineeri…

Automated Essay ScoringFeature Engineering

Leveraging Large Language Models with Chain-of-Thought and Prompt Engineering for Traffic Crash Severity Analysis and Inference

2024-08-04 · Hao Zhen, Yucheng Shi, Yongcan Huang, Jidong J. Yang 외

Harnessing the power of Large Language Models (LLMs), this study explores the use of three state-of-the-art LLMs, specifically GPT-3.5-turbo, LLaMA3-8B, and LLaMA3-70B, for crash severity inference, framing it as a class…

Logical ReasoningPrompt Engineering

Investigating how well contextual features are captured by bi-directional recurrent neural network models

2017-09-03 · WS 2017 12 · Kushal Chawla, Sunil Kumar Sahu, Ashish Anand

Learning algorithms for natural language processing (NLP) tasks traditionally rely on manually defined relevant contextual features. On the other hand, neural network models using an only distributional representation of…

Feature Engineering

Fountain -- an intelligent contextual assistant combining knowledge representation and language models for manufacturing risk identification

2023-08-01 · Saurabh Kumar, Daniel Fuchs, Klaus Spindler

Deviations from the approved design or processes during mass production can lead to unforeseen risks. However, these changes are sometimes necessary due to changes in the product design characteristics or an adaptation i…

CPUSemantic SimilaritySemantic Textual Similarity