paper-with-me

홈 › Papers

SLOTH: Structured Learning and Task-based Optimization for Time Series Forecasting on Hierarchies

2023-02-11 · Fan Zhou, Chen Pan, Lintao Ma, Yu Liu, Shiyu Wang, James Zhang, Xinxin Zhu, Xuanwei Hu, Yunhua Hu, Yangfei Zheng, Lei Lei, Yun Hu

Multivariate time series forecasting with hierarchical structure is widely used in real-world applications, e.g., sales predictions for the geographical hierarchy formed by cities, states, and countries. The hierarchical time series (HTS) forecasting includes two sub-tasks, i.e., forecasting and reconciliation. In the previous works, hierarchical information is only integrated in the reconciliation step to maintain coherency, but not in forecasting step for accuracy improvement. In this paper, we propose two novel tree-based feature integration mechanisms, i.e., top-down convolution and bottom-up attention to leverage the information of the hierarchical structure to improve the forecasting performance. Moreover, unlike most previous reconciliation methods which either rely on strong assumptions or focus on coherent constraints only,we utilize deep neural optimization networks, which not only achieve coherency without any assumptions, but also allow more flexible and realistic constraints to achieve task-based targets, e.g., lower under-estimation penalty and meaningful decision-making loss to facilitate the subsequent downstream tasks. Experiments on real-world datasets demonstrate that our tree-based feature integration mechanism achieves superior performances on hierarchical forecasting tasks compared to the state-of-the-art methods, and our neural optimization networks can be applied to real-world tasks effectively without any additional effort under coherence and task-based constraints

📄 PDF Abstract BibTeX arXiv:2302.05650

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingMultivariate Time Series ForecastingTime SeriesTime Series AnalysisTime Series Forecasting

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

SlothSpeech: Denial-of-service Attack Against Speech Recognition Models

2023-06-01 · Mirazul Haque, Rutvij Shah, Simin Chen, Berrak Şişman 외

Deep Learning (DL) models have been popular nowadays to execute different speech-related tasks, including automatic speech recognition (ASR). As ASR is being used in different real-time scenarios, it is important that th…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

FlashSloth : Lightning Multimodal Large Language Models via Embedded Visual Compression

2025-01-01 · CVPR 2025 1 · Bo Tong, Bokai Lai, Yiyi Zhou, Gen Luo 외

Despite a big leap forward in capability, multimodal large language models (MLLMs) tend to behave like a sloth in practical use, i.e., slow response and large latency. Recent efforts are devoted to building tiny MLLM…

Descriptive

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression

2024-12-05 · Bo Tong, Bokai Lai, Yiyi Zhou, Gen Luo 외

Despite a big leap forward in capability, multimodal large language models (MLLMs) tend to behave like a sloth in practical use, i.e., slow response and large latency. Recent efforts are devoted to building tiny MLLMs fo…

DescriptiveVisual Question Answering

Chronicals: A High-Performance Framework for LLM Fine-Tuning with 3.51x Speedup over Unsloth

2026-01-06 · Arjun S. Nair arxiv

Large language model fine-tuning is bottlenecked by memory: a 7B parameter model requires 84GB--14GB for weights, 14GB for gradients, and 56GB for FP32 optimizer states--exceeding even A100-40GB capacity. We present Chro…

Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families

2024-12-09 · Felipe Maia Polo, Seamus Somerstep, Leshem Choshen, Yuekai Sun 외

Scaling laws for large language models (LLMs) predict model performance based on parameters like size and training data. However, differences in training configurations and data processing across model families lead to s…

Emotional IntelligenceInstruction Following