paper-with-me

홈 › Papers

How Useful are LLMs for Grammar Engineering? Cantonese ParGram Resources and Controlled Experimental Evaluation with English Baselines

2026-08-24 · Chit-Fung Lam arxiv

This paper presents new Cantonese ParGram resources and evaluates LLMs for knowledge-driven grammar engineering within a controlled experimental paradigm. Using Cantonese ParGram resources as gold standards, with corresponding English baselines, we investigate whether OpenAI's gpt-oss-120b and GPT-5.4 can generate machine-processable grammars from sentences and target formal structures under systematically varied prompting conditions. GPT-5.4 outperformed gpt-oss-120b, while grammars generated from target formal structures generally outperformed those generated from sentences. Although both models could generate locally plausible phrase-structure rules, lexical entries, and templates, they often struggled to coordinate interacting formal constraints, especially in multi-construction settings. The results characterize both the capabilities and limitations of current LLMs for potential integration into AI-assisted expert workflows: LLMs may support intermediate stages of grammar development, but human linguistic expertise remains central to analysis, validation, and refinement. The study also contributes new Cantonese symbolic grammatical resources.

📄 PDF Abstract BibTeX arXiv:2608.23448

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Developing a Deep Grammar of Indonesian within the ParGram Framework: Theoretical and Implementational Challenges

2012-11-01 · PACLIC 2012 11 · I Wayan Arka

Using Meta-Morph Rules to develop Morphological Analysers: A case study concerning Tamil

2019-09-01 · WS 2019 9 · Kengatharaiyer Sarveswaran, Gihan Dias, Miriam Butt

This paper describes a new and larger coverage Finite-State Morphological Analyser (FSM) and Generator for the Dravidian language Tamil. The FSM has been developed in the context of computational grammar engineering, adh…

MORPHMorphological Analysis

ParGramBank: The ParGram Parallel Treebank

2013-08-01 · ACL 2013 8 · Sebastian Sulger, Miriam Butt, Tracy Holloway King, Paul Meurer 외

How Well Do LLMs Handle Cantonese? Benchmarking Cantonese Capabilities of Large Language Models

2024-08-29 · Jiyue Jiang, Pengan Chen, Liheng Chen, Sheng Wang 외

The rapid evolution of large language models (LLMs) has transformed the competitive landscape in natural language processing (NLP), particularly for English and other data-rich languages. However, underrepresented langua…

BenchmarkingGeneral Knowledge

ACE: Automatic Colloquialism, Typographical and Orthographic Errors Detection for Chinese Language

2016-12-01 · COLING 2016 12 · Shichao Dong, Gabriel Pui Cheong Fung, Binyang Li, Baolin Peng 외

We present a system called ACE for Automatic Colloquialism and Errors detection for written Chinese. ACE is based on the combination of N-gram model and rule-base model. Although it focuses on detecting colloquial Canton…

Language ModelingLanguage Modelling