paper-with-me

Papers

ClimaEmpact: Domain-Aligned Small Language Models and Datasets for Extreme Weather Analytics

2025-04-27 · Deeksha Varshney, Keane Ong, Rui Mao, Erik Cambria, Gianmarco Mengaldo

Accurate assessments of extreme weather events are vital for research and policy, yet localized and granular data remain scarce in many parts of the world. This data gap limits our ability to analyze potential outcomes and implications of extreme weather events, hindering effective decision-making. Large Language Models (LLMs) can process vast amounts of unstructured text data, extract meaningful insights, and generate detailed assessments by synthesizing information from multiple sources. Furthermore, LLMs can seamlessly transfer their general language understanding to smaller models, enabling these models to retain key knowledge while being fine-tuned for specific tasks. In this paper, we propose Extreme Weather Reasoning-Aware Alignment (EWRA), a method that enhances small language models (SLMs) by incorporating structured reasoning paths derived from LLMs, and ExtremeWeatherNews, a large dataset of extreme weather event-related news articles. EWRA and ExtremeWeatherNews together form the overall framework, ClimaEmpact, that focuses on addressing three critical extreme-weather tasks: categorization of tangible vulnerabilities/impacts, topic labeling, and emotion analysis. By aligning SLMs with advanced reasoning strategies on ExtremeWeatherNews (and its derived dataset ExtremeAlign used specifically for SLM alignment), EWRA improves the SLMs' ability to generate well-grounded and domain-specific responses for extreme weather analytics. Our results show that the approach proposed guides SLMs to output domain-aligned responses, surpassing the performance of task-specific models and offering enhanced real-world applicability for extreme weather analytics.

📄 PDF Abstract BibTeX arXiv:2504.19066

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesEmotion Recognition

Similar Papers 제목 키워드 기반

Robust Zero-Shot Cross-Domain Slot Filling with Example Values

2019-06-17 · ACL 2019 7 · Darsh J Shah, Raghav Gupta, Amir A Fayazi, Dilek Hakkani-Tur

Task-oriented dialog systems increasingly rely on deep learning-based slot filling models, usually needing extensive labeled training data for target domains. Often, however, little to no target domain training data may …

slot-fillingSlot FillingZero-shot Slot Filling

Switching-Aligned-Words Data Augmentation for Neural Machine Translation

2021-01-01 · Fengshun Xiao, Zuchao Li, Hai Zhao

In neural machine translation (NMT), data augmentation methods such as back-translation make it possible to use extra monolingual data to help improve translation performance, while it needs extra training data and the i…

Data AugmentationMachine TranslationNMTPosition+1

In-Training Defenses against Emergent Misalignment in Language Models

2025-08-08 · David Kaczér, Magnus Jørgenvåg, Clemens Vetter, Esha Afzal 외 arxiv

Fine-tuning lets practitioners repurpose aligned large language models (LLMs) for new domains, yet recent work reveals emergent misalignment (EM): Even a small, domain-specific fine-tune can induce harmful behaviors far …

Creating an Aligned Russian Text Simplification Dataset from Language Learner Data

2021-04-01 · EACL (BSNLP) 2021 4 · Anna Dmitrieva, Jörg Tiedemann

Parallel language corpora where regular texts are aligned with their simplified versions can be used in both natural language processing and theoretical linguistic studies. They are essential for the task of automatic te…

Text Simplification

Person Re-ID in 2025: Supervised, Self-Supervised, and Language-Aligned. What Works?

2026-01-28 · Lakshman Balasubramanian arxiv

Person Re-Identification (ReID) remains a challenging problem in computer vision. This work reviews various training paradigm and evaluates the robustness of state-of-the-art ReID models in cross-domain applications and …

Person Re-Identification