synthetic
Status: Exploratory

Project Halflight
.

What synthetic data does to a model over time.

Current StatusExploratory
AttributionSleepers Research Laboratory
Published Date2026-03-15
Last Updated2026-07-20
Abstract

Overview

Studies the compounding effects of synthetic data on model behavior across successive fine-tuning generations, tracking alignment drift, style collapse, and cognitive degradation.

Problem & Scope

Research Focus

The Research Problem

Training language models recursively on model-generated synthetic datasets leads to variance reduction, loss of rare domain knowledge, and behavioral homogenization over repeated fine-tuning cycles.

Why It Matters

As web datasets become increasingly saturated with AI-generated text, maintaining data quality and preventing model collapse is essential for long-term AI safety.

Research Objective

Quantify behavioral drift, perplexity degradation, and safety alignment shifts across 10+ successive generations of synthetic self-training.

Methodology

Technical Approach

Key Methodology Vectors

  • Recursive fine-tuning loops using synthetic output datasets across multiple architecture variants.
  • Semantic diversity measurement and tail-knowledge evaluation across generations.
  • Benchmark testing against standard reasoning and safety evaluation suites.

Expected Outcomes & Milestones

  • Data curation heuristics and hybrid synthetic-human data ratios that prevent recursive model collapse.
  • Quantitative threshold indicators for synthetic data pollution in training pipelines.
Citations

References & Prior Work

  • The Curse of Recursion: Training on Generated Data Makes Models ForgetShumailov et al., Nature (2024).
  • Synthetic Data Compounding and Alignment StabilitySleepers Research Paper (2026).

Interested in this research direction?

Sleepers Research welcomes inquiries from teams exploring autonomous agent security and applied machine learning.

Contact the Laboratory