paper-with-me

Papers

DataPro -- A Standardized Data Understanding and Processing Procedure: A Case Study of an Eco-Driving Project

2025-01-21 · Zhipeng Ma, Bo Nørregaard Jørgensen, Zheng Grace Ma

A systematic pipeline for data processing and knowledge discovery is essential to extracting knowledge from big data and making recommendations for operational decision-making. The CRISP-DM model is the de-facto standard for developing data-mining projects in practice. However, advancements in data processing technologies require enhancements to this framework. This paper presents the DataPro (a standardized data understanding and processing procedure) model, which extends CRISP-DM and emphasizes the link between data scientists and stakeholders by adding the "technical understanding" and "implementation" phases. Firstly, the "technical understanding" phase aligns business demands with technical requirements, ensuring the technical team's accurate comprehension of business goals. Next, the "implementation" phase focuses on the practical application of developed data science models, ensuring theoretical models are effectively applied in business contexts. Furthermore, clearly defining roles and responsibilities in each phase enhances management and communication among all participants. Afterward, a case study on an eco-driving data science project for fuel efficiency analysis in the Danish public transportation sector illustrates the application of the DataPro model. By following the proposed framework, the project identified key business objectives, translated them into technical requirements, and developed models that provided actionable insights for reducing fuel consumption. Finally, the model is evaluated qualitatively, demonstrating its superiority over other data science procedures.

📄 PDF Abstract BibTeX arXiv:2501.12176

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DataProphet: Demystifying Supervision Data Generalization in Multimodal LLMs

2026-03-20 · Xuan Qi, Luxi He, Dan Roth, Xingyu Fu arxiv

Conventional wisdom for selecting supervision data for multimodal large language models (MLLMs) is to prioritize datasets that appear similar to the target benchmark, such as text-intensive or vision-centric tasks. Howev…

New Approaches for Natural Language Understanding based on the Idea that Natural Language encodes both Information and its Processing Procedures

2020-10-24 · Limin Zhang

We must recognize that natural language is a way of information encoding, and it encodes not only the information but also the procedures for how information is processed. To understand natural language, the same as we c…

AttributeNatural Language Understanding

Optimizing Multimodal LLMs for Egocentric Video Understanding: A Solution for the HD-EPIC VQA Challenge

2026-01-15 · Sicheng Yang, Yukai Huang, Shitong Sun, Weitong Cai 외 arxiv

Multimodal Large Language Models (MLLMs) struggle with complex video QA benchmarks like HD-EPIC VQA due to ambiguous queries/options, poor long-range temporal reasoning, and non-standardized outputs. We propose a framewo…

Seeding the Singularity for A.I

2019-08-04 · Pavel Kraikivski

The singularity refers to an idea that once a machine having an artificial intelligence surpassing the human intelligence capacity is created, it will trigger explosive technological and intelligence growth. I propose to…

A Comprehensive Study of Multimodal Large Language Models for Image Quality Assessment

2024-03-16 · Tianhe Wu, Kede Ma, Jie Liang, Yujiu Yang 외

While Multimodal Large Language Models (MLLMs) have experienced significant advancement in visual understanding and reasoning, their potential to serve as powerful, flexible, interpretable, and text-driven models for Ima…

Image Quality Assessment