Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
3 results for “GRPO”
Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL
Hugging Face announced an experimental asynchronous variant of GRPO (Generalized Reinforcement Learning from Preferences) using LoRA adapters across distributed training jobs, eliminating NCCL dependencies by introducing a custom bucket-and-proxy coordination layer.
Published Sep 10, 2026 · Analyzed Sep 14, 2026
Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps
Hugging Face announced a new fine-tuning method called GRPO (Guided Reinforcement Policy Optimization) that achieves improved structured output generation from a 350M-parameter model in just 100 optimization steps, positioning it as a computationally efficient alternative to standard RLHF.
Sep 3, 2026
GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity
Researchers prove three popular methods for training language models are actually different settings of one parameter.
Published Jul 2, 2026 · Analyzed Jul 5, 2026