Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
28 results for “fine-tuning”
How Fyxer built an AI executive assistant people trust
Fyxer, a startup leveraging OpenAI's models, launched an AI executive assistant that organizes email inboxes and drafts replies mimicking individual user voice — positioning itself as a trust-oriented productivity tool.
Sep 14, 2026
Newton Matching for Generative Modeling: A Unified Framework for Fine-Tuning and Sampling
Researchers introduce 'Newton Matching', a theoretical framework unifying fine-tuning and sampling in generative modeling via iterative optimization on density manifolds, with proofs of convergence and connections to Fisher-Rao geometry.
Sep 10, 2026
LentEx: Generalizable Latent Entity Extraction via Synthetic Data and Instruction-Tuned LLMs
LentEx is a new research framework for latent entity extraction that uses synthetic data and instruction-tuning to enhance smaller LLMs, claiming improved performance and cross-domain generalization on NLP benchmarks.
Sep 7, 2026
Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps
Hugging Face announced a new fine-tuning method called GRPO (Guided Reinforcement Policy Optimization) that achieves improved structured output generation from a 350M-parameter model in just 100 optimization steps, positioning it as a computationally efficient alternative to standard RLHF.
Sep 3, 2026
Skild AI unveils S1, a robotics foundation model that it says can learn tasks never seen during pretraining, using a single video demo, without fine-tuning (Skild AI)
Skild AI announced S1, a robotics foundation model claiming to learn entirely new physical tasks from a single video demonstration without fine-tuning or post-training.
Aug 26, 2026
Capacity-Dependent Effects of Data Selection for Reasoning
A new arXiv preprint challenges the assumption that high-likelihood responses are universally optimal for reasoning-focused fine-tuning, demonstrating instead that data selection effectiveness depends critically on model size and training duration.
Aug 17, 2026
Can Spectral-Clipping Enable Better Learning While Forgetting Less for Low-Rank Adaptation?
A new LoRA variant called SCLoRA is proposed to reduce catastrophic forgetting in low-rank adaptation by applying spectral clipping to singular components, with experimental validation showing improved task performance and knowledge retention.
Aug 14, 2026
Weightless Fine-Tuning: Personalizing LLMs via Logit-Space Transport
A new method called Weightless Fine-Tuning (WFT) enables personalization of large language models at decoding time without updating model weights, reducing computational cost while approximating the distributional effect of supervised fine-tuning.
Aug 13, 2026
SeFoRA: Sketch-Aggregated Federated Low-Rank Adaptation with Heterogeneous Client Ranks
SeFoRA is a new federated learning algorithm that enables parameter-efficient fine-tuning of large language models across heterogeneous clients using sketch-based aggregation to resolve rank incompatibility and bilinear mismatch in LoRA updates.
Aug 12, 2026
When Do Corrective Features Help? An Agent for Corrective Feature Discovery on Black-Box Forecasters
Researchers introduce CRAFTER, a method to discover interpretable corrective features from forecast model residuals to improve black-box forecasting performance without fine-tuning the original model.
Aug 7, 2026
Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation
A research paper identifies downstream fine-tuning benchmarks (e.g., GLUE) as unreliable proxies for evaluating federated pre-trained language models, finding that intrinsic next-token prediction better preserves ranking fidelity to pre-training performance.
Aug 3, 2026
Misalignment Has a Personality: A Big Five Account of Emergent Misalignment
Researchers propose modeling AI misalignment as a 'personality shift' using Big Five traits, claiming fine-tuning on flawed data induces consistent, measurable changes in model behavior across domains and models.
Jul 30, 2026
Enough is as good as a feast: A Comprehensive Analysis of How Reinforcement Learning Mitigates Task Conflicts in LLMs
A new arXiv preprint claims reinforcement learning (RL) training reduces task conflicts during model merging in LLMs compared to supervised fine-tuning, citing three empirical and theoretical mechanisms.
Jul 27, 2026
MoE$^2$-LoRA: When MoE Models Meet MoE-style Low-Rank Adaptation
A new parameter-efficient fine-tuning method called MoE²-LoRA is introduced to improve adaptation of Mixture-of-Experts language models by dynamically routing low-rank adapters using pretrained router signals and sharing a global expert pool across layers.
Jul 27, 2026
RAG vs Fine-Tuning for Multi-Tenant SaaS: Which Architecture Would You Choose?
A Reddit user seeks expert architectural advice on choosing between RAG and fine-tuning for a multi-tenant SaaS platform handling sensitive user documents and requiring accurate, cited answers when user data is sparse.
Jul 26, 2026
Find Before You Fine-Tune: A Diagnostic Study of Small LLMs for Cybersecurity QA
Researchers introduce FiT, a diagnostic framework to evaluate small LLMs before fine-tuning for cybersecurity QA, revealing that fine-tuning often degrades core knowledge capabilities and that pre-tuning diagnostics can predict post-tuning outcomes.
Jul 22, 2026
Presentation: Engineering AI for Creativity and Curiosity on Mobile
An InfoQ presentation by Bhavuk Jain outlines engineering strategies for deploying AI features—specifically AI Wallpapers and Circle to Search—on mobile devices, emphasizing runtime safety, OS integration, and trade-offs between UX, latency, and cost.
Jul 21, 2026
TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment
Researchers introduced TRACE, a new safety patching method for fine-tuned LLMs that claims to recover alignment without degrading task utility by learning from simulated harmful tuning trajectories.
Jul 21, 2026
PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization
A new reinforcement learning framework called PPO-HSC is introduced to mitigate mode collapse in LLM fine-tuning by incentivizing semantic novelty while preserving solution validity.
Jul 21, 2026
LoRA Speedrun – a public wall-clock leaderboard for fine-tuning techniques
A community-maintained leaderboard tracking wall-clock time for LoRA fine-tuning across models and hardware, serving as an informal benchmark for efficiency claims in open LLM development.
Jul 20, 2026
An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon?
A new arXiv preprint challenges the robustness of 'Emergent Misalignment' (EM) — a claimed phenomenon where LMs abruptly develop broad misalignment after narrow fine-tuning — showing its appearance depends heavily on superficial dataset artifacts like response length, not deep mechanistic shifts.
Jul 13, 2026
Post-Generation Curation of Synthetic Images via Homogeneous-Heterogeneous Splitting
A new method called Homogeneous-Heterogeneous Splitting improves synthetic image utility by selecting subsets based on fidelity and diversity, without retraining generators, achieving real-data-level performance with up to 40% fewer samples.
Jul 8, 2026
google/tabfm-1.0.0
Google Research released TabFM, a zero-shot foundation model for tabular data that claims to perform classification and regression without fine-tuning or hyperparameter search by treating training examples as context.
Published Jul 4, 2026 · Analyzed Jul 6, 2026
What does "Safe AI" look like? [D]
A Reddit user poses open questions about the practicality and value of safety training for open-weight LLMs in light of rapid emergence of 'uncensored' model variants, highlighting tensions between safety goals, technical feasibility, and real-world adversarial behavior.
Published Jul 3, 2026 · Analyzed Jul 6, 2026