propella-1: Multi-Property Document Annotation for LLM Data Curation at Scale
Since FineWeb-Edu, data curation for LLM pretraining has predominantly relied on single scalar quality scores produced by small classifiers. A single score conflates multiple quality dimensions, prevents flexible filtering, and offers no interpretability. We introduce propella-1, a family of small multilingual LLMs (0.6B, 1.7B, 4B parameters) that annotate text documents across 18 properties organized into six categories: core content, classification, quality and value, audience and purpose, safety and compliance, and geographic relevance. The models support 57 languages and produce structured JSON annotations conforming to a predefined schema. Evaluated against a frontier commercial LLM as a reference annotator, the 4B model achieves higher agreement than much larger general-purpose models. We release propella-annotations, a dataset of over three billion document annotations covering major pretraining corpora including data from FineWeb-2, FinePDFs, HPLT 3.0, and Nemotron-CC. Using these annotations, we present a multi-dimensional compositional analysis of widely used pretraining datasets, revealing substantial differences in quality, reasoning depth, and content composition that single-score approaches cannot capture. All model weights and annotations are released under permissive, commercial-use licenses.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Nanosatellite Design Considerations for a Mission to Explore the Propellant Sloshing Problem
Sloshing Platform for In-Orbit Controller Experimentation is an ambitious, student run mission to design and fly a cubesat to study fluid sloshing in spacecraft. The project will examine zero-g propellant sloshing from a…
A Markov Decision Process Framework for Early Maneuver Decisions in Satellite Collision Avoidance
We develop a Markov decision process (MDP) framework to autonomously make guidance decisions for satellite collision avoidance maneuver (CAM) and a reinforcement learning policy gradient (RL-PG) algorithm to enable direc…
Reinforcement LearningCollision AvoidanceMultidisciplinary Design Optimization of Reusable Launch Vehicles for Different Propellants and Objectives
Identifying the optimal design of a new launch vehicle is most important since design decisions made in the early development phase limit the vehicles' later performance and determines the associated costs. Reusing the f…
Low-cost, Lightweight Electronic Flow Regulators for Throttling Liquid Rocket Engines
For small-scale liquid rockets, pressure-fed systems are commonly favoured due to their simplicity and low weight. In such systems, accurate regulation of both tank and injector pressures over a wide range of upstream pr…
Economics of In-Space Industry and Competitiveness of Lunar-Derived Rocket Propellant
Economic parameters are identified for an in-space industry where the capital is made on one planet, it is transported to and teleoperated on a second planet, and the product is transported off the second planet for cons…