paper-with-me

Papers

ClimateChat: Designing Data and Methods for Instruction Tuning LLMs to Answer Climate Change Queries

2025-06-12 · Zhou Chen, Xiao Wang, YuanHong Liao, Ming Lin, Yuqi Bai

As the issue of global climate change becomes increasingly severe, the demand for research in climate science continues to grow. Natural language processing technologies, represented by Large Language Models (LLMs), have been widely applied to climate change-specific research, providing essential information support for decision-makers and the public. Some studies have improved model performance on relevant tasks by constructing climate change-related instruction data and instruction-tuning LLMs. However, current research remains inadequate in efficiently producing large volumes of high-precision instruction data for climate change, which limits further development of climate change LLMs. This study introduces an automated method for constructing instruction data. The method generates instructions using facts and background knowledge from documents and enhances the diversity of the instruction data through web scraping and the collection of seed instructions. Using this method, we constructed a climate change instruction dataset, named ClimateChat-Corpus, which was used to fine-tune open-source LLMs, resulting in an LLM named ClimateChat. Evaluation results show that ClimateChat significantly improves performance on climate change question-and-answer tasks. Additionally, we evaluated the impact of different base models and instruction data on LLM performance and demonstrated its capability to adapt to a wide range of climate change scientific discovery tasks, emphasizing the importance of selecting an appropriate base model for instruction tuning. This research provides valuable references and empirical support for constructing climate change instruction data and training climate change-specific LLMs.

📄 PDF Abstract BibTeX arXiv:2506.13796

Code (1)

thu-esis/jiuzhou 공식 구현 pytorch

Tasks

scientific discovery

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

ClimateChat-300K: A Multi-Modal Facebook Dataset for Understanding Diverse Perspectives in Climate Communication

2026-05-22 · Wajdi Zaghouani, Md. Rafiul Biswas, Mabrouka Bessghaier, Shimaa Ibrahim 외 arxiv

We present ClimateChat-300K, a large-scale dataset of 299,329 public Facebook posts about climate change collected between May 2020 and May 2024 through the CrowdTangle platform. The dataset contains 41 metadata features…

Sentiment Analysis

The Flan Collection: Designing Data and Methods for Effective Instruction Tuning

2023-01-31 · Shayne Longpre, Le Hou, Tu Vu, Albert Webson 외

We study the design decisions of publicly available instruction tuning methods, and break down the development of Flan 2022 (Chung et al., 2022). Through careful ablation studies on the Flan Collection of tasks and metho…

COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning

2024-03-26 · Yuelin Bai, Xinrun Du, Yiming Liang, Yonggang Jin 외

Remarkable progress on English instruction tuning has facilitated the efficacy and reliability of large language models (LLMs). However, there remains a noticeable gap in instruction tuning for Chinese, where the complex…

All

Automatic Instruction Evolving for Large Language Models

2024-06-02 · Weihao Zeng, Can Xu, Yingxiu Zhao, Jian-Guang Lou 외

Fine-tuning large pre-trained language models with Evol-Instruct has achieved encouraging results across a wide range of tasks. However, designing effective evolving methods for instruction evolution requires substantial…

GSM8KHumanEval

Lius: Translation Model Based Instructional Lingustic Using Continual Instruction Tuning In Kupang Malay

2026-06-10 · Joanito Agili Lopo, Yunita Sari, Guntur Budi Herwanto arxiv

Large Language Models (LLMs) offer new potential for translation tasks but often experience performance degradation when handling low-resource languages. To address this limitation, we propose an approach for fine-tuning…

Machine Translation