Detect early signs of model collapse by analyzing training data degradation over time using statistical, semantic, and temporal linguistic features.
Model collapse is a critical failure mode in modern AI systems — especially Large Language Models (LLMs). When a model is repeatedly trained on synthetic or degraded data (e.g., its own outputs), the training data loses diversity, repetition increases, and the model's outputs become nonsensical or circular. This is sometimes called "data poisoning by recursion".
This project builds a machine learning pipeline that:
- Detects which stage of collapse a text corpus is in (5 stages: Healthy → Collapse)
- Quantifies collapse severity with a continuous Collapse Risk Score (0.0 → 1.0)
- Validates that the Risk Score reliably correlates with actual model performance degradation
| Question | Answer this project provides |
|---|---|
| Can we detect collapse from raw text features alone? | ✅ Yes — using vocabulary, entropy, repetition, and n-gram diversity metrics |
| How early can we detect it? | ✅ Even at Stage 1 (Slight Drift), the Risk Score begins rising |
| Can we assign a continuous severity number? | ✅ Yes — via Autoencoder reconstruction error |
| Does the score reflect real performance loss? | ✅ Yes — validated against simulated NLP model accuracy & perplexity |
MLR/
├── README.md ← You are here
├── requirements.txt ← All Python dependencies
├── data/ ← Auto-created; stores all outputs
│ ├── synthetic_dataset.csv ← Raw generated text data (5 stages)
│ ├── features.csv ← Extracted linguistic features
│ ├── predictions.csv ← Model predictions + risk scores
│ ├── feature_summary.png ← Feature distribution by stage
│ ├── model_comparison.png ← Accuracy/F1 bar chart across models
│ ├── confusion_matrix.png ← XGBoost confusion matrix
│ ├── feature_importance.png ← Top predictive features (XGBoost)
│ ├── risk_score_plot.png ← Risk scores across collapse stages
│ └── validation_plot.png ← Risk score vs. NLP accuracy correlation
├── 01_data_generation.ipynb ← Step 1: Create synthetic datasets
├── 02_feature_extraction.ipynb ← Step 2: Extract linguistic features
├── 03_model_training.ipynb ← Step 3: Train all models + Autoencoder
├── 04_evaluation.ipynb ← Step 4: Compare models, visualize
└── 05_validation_experiment.ipynb ← Step 5: Validate the Risk Score
- Python 3.9+ (tested on 3.13)
- Jupyter Notebook or VS Code with Jupyter extension
pip install -r requirements.txtRun each notebook top to bottom, in sequence:
01 → 02 → 03 → 04 → 05
jupyter notebookEach notebook is self-contained and saved with outputs. Simply open and run all cells.
| Stage | Label | What It Looks Like | Why It Matters |
|---|---|---|---|
| 0 | Healthy | Diverse, rich vocabulary, no repetition | Baseline for "normal" data |
| 1 | Slight Drift | Minor repetition of common phrases | Earliest detectable signal |
| 2 | Moderate Degradation | Reduced vocabulary, rising repetition | Clear drift, model starts to suffer |
| 3 | Severe Degradation | Heavy repetitive loops, gibberish phrases | Strong performance loss |
| 4 | Collapse | Fully collapsed — monotone or nonsensical | Model unusable |
These match real-world patterns observed in LLMs trained repeatedly on their own outputs.
What it does:
Creates a synthetic dataset of 5,000 text samples (1,000 per stage). Each stage simulates how model-generated text degrades over time.
How it works:
- Stage 0 (Healthy): Randomly assembled sentences with diverse vocabulary
- Stage 1–4: Progressively introduces repetition, removes vocabulary variety, adds gibberish tokens
Why synthetic data?
Real collapse data is hard to obtain at scale. Synthetic data lets us precisely control degradation levels, test all five stages, and avoid copyright/license issues with real corpora.
Output: data/synthetic_dataset.csv — columns: text, label
What it does:
Converts raw text into 7 quantitative linguistic features that capture collapse signals.
Features extracted:
| Feature | What It Measures | Why It Detects Collapse |
|---|---|---|
| TTR (Type-Token Ratio) | Unique words ÷ total words | Drops sharply as repetition increases |
| Entropy | Information diversity (Shannon entropy) | Falls as text becomes more predictable |
| Average Word Frequency | How common words are on average | Rises as rare vocabulary disappears |
| Vocabulary Size | Count of unique words per sample | Shrinks during degradation |
| Bigram Diversity | Unique 2-word combinations ÷ total bigrams | Reflects structural repetition |
| Avg Sentence Length | Mean words per sentence | Becomes erratic or very short at collapse |
| Duplicate Ratio | Fraction of repeated sentences | Spikes to 1.0 at full collapse |
Why these features?
They are computationally cheap, interpretable, and proven effective for detecting text quality degradation. No neural network is required for feature extraction — making the pipeline fast and reproducible.
Output: data/features.csv, data/feature_summary.png
What it does:
Trains four distinct models on the extracted features to classify collapse stage, then builds an Autoencoder to generate a continuous Collapse Risk Score.
- Why: Fast, interpretable baseline — shows what's achievable with a linear model
- Input: 7 extracted features
- Output: Class prediction (0–4)
- Why: Handles non-linear relationships, robust to outliers, gives feature importance
- Input: 7 extracted features
- Output: Class prediction + probability
- Why: State-of-the-art gradient boosting — highest accuracy on tabular data
- Input: 7 extracted features
- Output: Class prediction + probability + feature importance
- Why: Collapse often develops over time across a sequence of batches; LSTM captures that temporal pattern
- Architecture: Embedding → LSTM (64 units) → Dense → Softmax
- Input: Feature vectors as sequences
- Output: Class prediction
- Why: Instead of classifying, the Autoencoder learns what "healthy" data looks like. When shown degraded data, it fails to reconstruct it accurately — and that reconstruction error becomes the Risk Score.
- Architecture: Encoder (7 → 4 → 2) → Decoder (2 → 4 → 7)
- Training: Trained only on Stage 0 (Healthy) samples
- Output: Risk Score per sample (0.0 = healthy, 1.0 = collapsed)
Why use an Autoencoder for the score?
Classification gives you a category (Stage 2), but the Autoencoder gives you a continuous, nuanced signal — Stage 2 samples close to Stage 1 will score lower than those near Stage 3. This is more useful for real-time monitoring.
Output: data/predictions.csv (all model predictions + risk scores)
What it does:
Loads predictions and generates four comprehensive visualizations to compare models and understand what drives collapse detection.
Visualizations generated:
| Plot | File | What It Shows |
|---|---|---|
| Risk Score Distribution | risk_score_plot.png |
How risk scores separate the 5 stages |
| Model Comparison | model_comparison.png |
Accuracy + F1 for all 4 classifiers |
| Confusion Matrix | confusion_matrix.png |
Where XGBoost mistakes occur between stages |
| Feature Importance | feature_importance.png |
Which features XGBoost relies on most |
Key finding: XGBoost achieves the highest classification accuracy. The Autoencoder risk scores cleanly separate all 5 stages with minimal overlap.
What it does:
Answers the critical research question: "Does a higher Risk Score actually mean worse model performance?"
Experiment design:
- For each collapse stage, a simulated NLP classification task is run on the text data
- Accuracy drops and perplexity (a language model metric — lower = better) increases as stage increases
- The average Collapse Risk Score for each stage is plotted against simulated accuracy and perplexity
Why perplexity?
Perplexity is the standard measure of how "surprised" a language model is by text. Collapsed text is highly predictable (low perplexity is actually bad here — it means the model learned repetition, not language).
Result:
A clear, monotonic correlation is observed: as Risk Score increases from Stage 0 → 4, NLP accuracy falls and perplexity rises. This validates that the Risk Score is a meaningful early warning signal.
Output: data/validation_plot.png, summary metrics printed in notebook
| Library | Version | What It Does | Why Used |
|---|---|---|---|
numpy |
any | Numerical arrays & math | Core of all ML computations |
pandas |
any | Tabular data manipulation | Load/save CSVs, feature tables |
matplotlib |
any | Plotting | All visualizations |
seaborn |
any | Statistical plots | Prettier distribution/heatmap plots |
scikit-learn |
any | Logistic Regression, Random Forest, metrics | Industry-standard ML library |
xgboost |
any | Gradient boosted trees | Best-in-class tabular classifier |
nltk |
any | Tokenization, n-gram analysis | Text preprocessing for feature extraction |
torch (PyTorch) |
any | LSTM + Autoencoder | Deep learning framework — flexible and well-documented |
scipy |
any | KL divergence, statistical tests | Used in validation metrics |
tqdm |
any | Progress bars | Better UX during long loops |
transformers |
any | HuggingFace tokenizers (optional) | Available for semantic embeddings if extended |
| Model | Accuracy | F1 Score |
|---|---|---|
| Logistic Regression | ~72% | ~0.71 |
| Random Forest | ~88% | ~0.87 |
| XGBoost | ~93% | ~0.92 |
| LSTM | ~80% | ~0.79 |
Collapse Risk Score (Autoencoder):
- Stage 0 (Healthy): 0.05 – 0.15
- Stage 2 (Moderate): 0.45 – 0.60
- Stage 4 (Collapse): 0.85 – 1.00
Binary (Collapsed / Not Collapsed) is too coarse for early warning. A 5-stage taxonomy lets researchers act at Stage 1–2, before irreversible collapse.
Classifier probabilities are unreliable for out-of-distribution samples. An Autoencoder trained only on healthy data provides a principled anomaly detection score — no labeled collapse data needed at inference time.
Gives precise control over degradation severity, ensures reproducibility, and avoids the complexity of sourcing/licensing real collapse datasets.
- Real data integration — replace synthetic data with actual Reddit/WikiText corpora with artificial degradation applied
- Embedding-based features — add BERT sentence embeddings for semantic drift detection
- Real-time dashboard — feed the Risk Score into a Streamlit or Gradio monitoring UI
- Temporal LSTM pipeline — feed batches sequentially to detect drift across training iterations
- Alert thresholds — define configurable thresholds (e.g., Risk Score > 0.6 → raise alert)
- Model Collapse in LLMs — Shumailov et al., 2024 — Original paper coining "model collapse"
- Shannon Entropy & Text Diversity
- XGBoost Documentation
- PyTorch Autoencoder Tutorial
- Self-BLEU for Diversity Evaluation
Built as part of a research project on AI System Safety & Language Model Robustness.
"The collapse of a language model is not a sudden event — it is a slow, measurable drift. This project measures it."