paper-with-me

홈 › Papers

CAPC-CG: A Large-Scale, Expert-Directed LLM-Annotated Corpus of Adaptive Policy Communication in China

2025-10-10 · Bolun Sun, Charles Chang, Yuen Yuen Ang, Ruotong Mu, Yuchen Xu, Zhengxin Zhang, Pingxu Hao arxiv

We introduce CAPC-CG, the Chinese Adaptive Policy Communication (Central Government) Corpus, the first open dataset of Chinese policy directives annotated with a five-color taxonomy of clear and ambiguous language categories, building on Ang's theory of adaptive policy communication. Spanning 1949-2023, this corpus includes national laws, administrative regulations, and ministerial rules issued by China's top authorities. Each document is segmented into paragraphs, producing a total of 3.3 million units. Alongside the corpus, we release comprehensive metadata, a two-round labeling framework, and a gold-standard annotation set developed by expert and trained coders. Inter-annotator agreement achieves a Fleiss's kappa of K = 0.86 on directive labels, indicating high reliability for supervised modeling. We provide baseline classification results with several large language models (LLMs), together with our annotation codebook, and describe patterns from the dataset. This release aims to support downstream tasks and multilingual NLP research in policy communication.

📄 PDF Abstract BibTeX arXiv:2510.08986

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MultiCapCLIP: Auto-Encoding Prompts for Zero-Shot Multilingual Visual Captioning

2023-08-25 · Bang Yang, Fenglin Liu, Xian Wu, YaoWei Wang 외

Supervised visual captioning models typically require a large scale of images or videos paired with descriptions in a specific language (i.e., the vision-caption pairs) for training. However, collecting and labeling larg…

Image CaptioningVideo Captioning

Identification and Estimation of Conditional Average Partial Causal Effects via Instrumental Variable

2024-01-20 · Yuta Kawakami, manabu kuroki, Jin Tian

There has been considerable recent interest in estimating heterogeneous causal effects. In this paper, we study conditional average partial causal effects (CAPCE) to reveal the heterogeneity of causal effects with contin…

CapCLIP: A Vision-Language Representation Alignment Approach for Wireless Capsule Endoscopy Analysis

2026-05-08 · Haroon Wahab, Irfan Mehmood, Hassan Ugail arxiv

Wireless capsule endoscopy (WCE) enables non-invasive visual assessment of the small bowel, but its clinical utility is constrained by the large volume of frames generated per examination and the difficulty of recognisin…

Representation LearningCross-Modal RetrievalText ClassificationImage Retrieval

Learning Transferable Facial Emotion Representations from Large-Scale Semantically Rich Captions

2025-07-28 · Licai Sun, Xingxun Jiang, Haoyu Chen, Yante Li 외 arxiv

Current facial emotion recognition systems are predominately trained to predict a fixed set of predefined categories or abstract dimensional values. This constrained form of supervision hinders generalization and applica…

Facial Emotion RecognitionRepresentation LearningContrastive Learning

Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Tests

2026-06-05 · Thanawat Lodkaew, Johannes Ackermann, Soichiro Nishimori, Nontawat Charoenphakdee 외 arxiv

A growing failure mode in agent evaluation and training is that models can achieve high evaluation scores by exploiting shortcuts instead of solving the intended task, producing deceptive performance. This makes evaluati…