paper-with-me

Papers

Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research

2025-09-19 · Richard Diehl Martinez, David Demitri Africa, Yuval Weiss, Suchir Salhan, Ryan Daniels, Paula Buttery arxiv

Building language models (LMs), especially small and medium ones, remains more art than science. While large LMs often improve by sheer scale, it is still unclear why many design choices work. For small LMs, this uncertainty is more limiting: tight parameter budgets make each decision critical, yet researchers still lack systematic, scientific ways to test and refine new ideas. We introduce Pico, a lightweight, modular framework that enables systematic, hypothesis-driven research for small and medium-scale language model development. Pico consists of two libraries that together provide a practical sandbox where researchers can make targeted changes to a model's architecture or training procedures and directly observe their effects on the model's behavior. To support reproducible experimentation, we also release a suite of baseline models, pico-decoder, trained under standardized conditions and open-sourced for the community. Case studies highlight how Pico can support iterative small LM design and analysis.

📄 PDF Abstract BibTeX arXiv:2509.16413

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Semi-Supervised Learning from Small Annotated Data and Large Unlabeled Data for Fine-grained PICO Entity Recognition

2024-12-26 · Fangyi Chen, Gongbo Zhang, Yilu Fang, Yifan Peng 외

Objective: Extracting PICO elements -- Participants, Intervention, Comparison, and Outcomes -- from clinical trial literature is essential for clinical evidence retrieval, appraisal, and synthesis. Existing approaches do…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1

Small Satellite Optical Communication Networks: Analytical Models

2018-08-29

Small satellites, especially picosatellites, appear poised to play an important role in the future of space systems. Due to their size, however, integrating them with high-throughput laser-based communication systems rem…

Real-Time Performance Benchmarking of TinyML Models in Embedded Systems (PICO: Performance of Inference, CPU, and Operations)

2025-09-05 · Abhishek Dey, Saurabh Srivastava, Gaurav Singh, Robert G. Pettit arxiv

This paper presents PICO-TINYML-BENCHMARK, a modular and platform-agnostic framework for benchmarking the real-time performance of TinyML models on resource-constrained embedded systems. Evaluating key metrics such as in…

Keyword Spotting

Towards the New XAI: A Hypothesis-Driven Approach to Decision Support Using Evidence

2024-02-02 · Thao Le, Tim Miller, Liz Sonenberg, Ronal Singh

Prior research on AI-assisted human decision-making has explored several different explainable AI (XAI) approaches. A recent paper has proposed a paradigm shift calling for hypothesis-driven XAI through a conceptual fram…

Decision Making

Unlocking the Power of Deep PICO Extraction: Step-wise Medical NER Identification

2020-04-30 · Tengteng Zhang, Yiqin Yu, Jing Mei, Zefang Tang 외

The PICO framework (Population, Intervention, Comparison, and Outcome) is usually used to formulate evidence in the medical domain. The major task of PICO extraction is to extract sentences from medical literature and cl…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+2