paper-with-me

홈 › Papers

Climate Research Domain BERTs: Pretraining, Adaptation, and Evaluation

2025-05-19 · Preprint 2025 5 · Andrija Poleksić, Sanda Martinčić-Ipšić

Motivated by the pressing issue of climate change and the growing volume of data, we pretrain three new language models using climate change research papers published in top-tier journals. Adaptation of existing domain-specific models is utilized for CliSciBERT and SciClimateBERT and pretraining from scratch resulted in CliReBERT (Climate Research BERT). The performance assessment is performed on the climate change NLP benchmark ClimaBench. We evaluate SciBERT, ClimateBERT, BERT, RoBERTa and DistilRoBERTa - along with our new models - CliReBERT, CliSciBERT and SciClimateBERT - using five different random seeds on all seven ClimaBench datasets. CliReBERT achieves the highest overall performance with a macro-averaged F1 score of 65.45%, and outperforms all other models on three of the seven tasks. Additionally, CliReBERT demonstrates the most stable fine-tuning behavior, yielding the lowest average standard deviation across seeds (0.0118). The 5-fold stratified cross-validation on the SciDCC dataset showed that CliReBERT achieved the highest overall macro-average F1 score (53.75%), slightly outperforming RoBERTa and DistilRoBERTa, while the domain-adapted models underperformed their base counterparts. The superior performance of CliReBERT is accompanied by the lowest tokenizer fertility, suggesting appropriateness to model domain-specific vocabulary.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Text Classification

Methods 이 논문이 사용한 방법론

RoBERTa 설명 없음
BASE 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
Weight Decay 설명 없음
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
WordPiece 설명 없음

Similar Papers 제목 키워드 기반

Multimodal Pretraining Unmasked: A Meta-Analysis and a Unified Framework of Vision-and-Language BERTs

2020-11-30 · Emanuele Bugliarello, Ryan Cotterell, Naoaki Okazaki, Desmond Elliott

Large-scale pretraining and task-specific fine-tuning is now the standard methodology for many tasks in computer vision and natural language processing. Recently, a multitude of methods have been proposed for pretraining…

The Climate Change Knowledge Graph: Supporting Climate Services

2026-02-23 · Miguel Ceriani, Fiorela Ciroku, Alessandro Russo, Massimiliano Schembri 외 arxiv

Climate change impacts a broad spectrum of human resources and activities, necessitating the use of climate models to project long-term effects and inform mitigation and adaptation strategies. These models generate multi…

Climate, Agriculture and Food

2021-05-25 · Ariel Ortiz-Bobea

Agriculture is arguably the most climate-sensitive sector of the economy. Growing concerns about anthropogenic climate change have increased research interest in assessing its potential impact on the sector and in identi…

Domain-Adaptive Climate Downscaling Under Temporal Distribution Shift

2026-07-06 · Shuochen Wang, Nishant Yadav, Auroop R. Ganguly arxiv

Deep-learning-based climate downscaling aims to learn relationships from historical low-resolution (LR) and high-resolution (HR) climate data to generate HR climate projections. However, this setting faces a temporal out…

Domain Adaptation

Self-Training Vision Language BERTs with a Unified Conditional Model

2022-01-06 · Xiaofeng Yang, Fengmao Lv, Fayao Liu, Guosheng Lin

Natural language BERTs are trained with language corpus in a self-supervised manner. Unlike natural language BERTs, vision language BERTs need paired data to train, which restricts the scale of VL-BERT pretraining. We pr…