paper-with-me

홈 › Papers

Towards Open-Ended Discovery for Low-Resource NLP

2025-09-22 · Bonaventure F. P. Dossou, Henri Aïdasso arxiv

Natural Language Processing (NLP) for low-resource languages remains fundamentally constrained by the lack of textual corpora, standardized orthographies, and scalable annotation pipelines. While recent advances in large language models have improved cross-lingual transfer, they remain inaccessible to underrepresented communities due to their reliance on massive, pre-collected data and centralized infrastructure. In this position paper, we argue for a paradigm shift toward open-ended, interactive language discovery, where AI systems learn new languages dynamically through dialogue rather than static datasets. We contend that the future of language technology, particularly for low-resource and under-documented languages, must move beyond static data collection pipelines toward interactive, uncertainty-driven discovery, where learning emerges dynamically from human-machine collaboration instead of being limited to pre-existing datasets. We propose a framework grounded in joint human-machine uncertainty, combining epistemic uncertainty from the model with hesitation cues and confidence signals from human speakers to guide interaction, query selection, and memory retention. This paper is a call to action: we advocate a rethinking of how AI engages with human knowledge in under-documented languages, moving from extractive data collection toward participatory, co-adaptive learning processes that respect and empower communities while discovering and preserving the world's linguistic diversity. This vision aligns with principles of human-centered AI, emphasizing interactive, cooperative model building between AI systems and speakers.

📄 PDF Abstract BibTeX arXiv:2510.01220

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Lingual Transfer

Similar Papers 제목 키워드 기반

CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery

2026-04-02 · Ao Qu, Han Zheng, Zijian Zhou, Yihao Yan 외 arxiv

Large language model (LLM)-based evolution is a promising approach for open-ended discovery, where progress requires sustained search and knowledge accumulation. Existing methods still rely heavily on fixed heuristics an…

OpenAgenet / OAN Yellow Paper: Technical Architecture for Trust-Governed Resource Identity and Discovery

2026-06-02 · Jinliang Xu arxiv

This yellow paper describes the technical architecture of OpenAgenet / OAN. OAN is a protocol-neutral trust layer for open Agent interconnection and discoverable AI resource products. It specifies the role architecture, …

Linghub2: Language Resource Discovery Tool for Language Technologies

2022-06-01 · LREC 2022 6 · Cécile Robin, Gautham Vadakkekara Suresh, Víctor Rodriguez-Doncel, John P. McCrae 외

Language resources are a key component of natural language processing and related research and applications. Users of language resources have different needs in terms of format, language, topics, etc. for the data they n…

Management

StatefulDiscovery: Evidence-Calibrated Claim Formation in Open-Ended Scientific Discovery

2026-06-10 · Jiayao Chen, Shi Liu, Linyi Yang arxiv

Open-ended scientific discovery asks agents to move beyond executing analyses for predefined questions. Across multiple rounds of exploration, a discovery agent must decide which phenomena warrant investigation while avo…

AlphaResearch: Accelerating New Algorithm Discovery with Language Models

2025-11-11 · Zhaojian Yu, Kaiyue Feng, Yilun Zhao, Shilin He 외 arxiv

LLMs have made significant progress in complex but easy-to-verify problems, yet they still struggle with discovering the unknown. In this paper, we present \textbf{AlphaResearch}, an autonomous research agent designed to…