Skip to content

About

Early Detection of Model Collapse in AI Systems using Statistical and Semantic Analysis

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

🧠 Model Collapse Detector for AI Systems

Detect early signs of model collapse by analyzing training data degradation over time using statistical, semantic, and temporal linguistic features.


📌 What Is This Project?

Model collapse is a critical failure mode in modern AI systems — especially Large Language Models (LLMs). When a model is repeatedly trained on synthetic or degraded data (e.g., its own outputs), the training data loses diversity, repetition increases, and the model's outputs become nonsensical or circular. This is sometimes called "data poisoning by recursion".

This project builds a machine learning pipeline that:

  1. Detects which stage of collapse a text corpus is in (5 stages: Healthy → Collapse)
  2. Quantifies collapse severity with a continuous Collapse Risk Score (0.0 → 1.0)
  3. Validates that the Risk Score reliably correlates with actual model performance degradation

🎯 Research Problem

Question Answer this project provides
Can we detect collapse from raw text features alone? ✅ Yes — using vocabulary, entropy, repetition, and n-gram diversity metrics
How early can we detect it? ✅ Even at Stage 1 (Slight Drift), the Risk Score begins rising
Can we assign a continuous severity number? ✅ Yes — via Autoencoder reconstruction error
Does the score reflect real performance loss? ✅ Yes — validated against simulated NLP model accuracy & perplexity

📁 Project Structure

MLR/
├── README.md                         ← You are here
├── requirements.txt                  ← All Python dependencies
├── data/                             ← Auto-created; stores all outputs
│   ├── synthetic_dataset.csv         ← Raw generated text data (5 stages)
│   ├── features.csv                  ← Extracted linguistic features
│   ├── predictions.csv               ← Model predictions + risk scores
│   ├── feature_summary.png           ← Feature distribution by stage
│   ├── model_comparison.png          ← Accuracy/F1 bar chart across models
│   ├── confusion_matrix.png          ← XGBoost confusion matrix
│   ├── feature_importance.png        ← Top predictive features (XGBoost)
│   ├── risk_score_plot.png           ← Risk scores across collapse stages
│   └── validation_plot.png           ← Risk score vs. NLP accuracy correlation
├── 01_data_generation.ipynb          ← Step 1: Create synthetic datasets
├── 02_feature_extraction.ipynb       ← Step 2: Extract linguistic features
├── 03_model_training.ipynb           ← Step 3: Train all models + Autoencoder
├── 04_evaluation.ipynb               ← Step 4: Compare models, visualize
└── 05_validation_experiment.ipynb    ← Step 5: Validate the Risk Score

🚀 How to Run

Prerequisites

  • Python 3.9+ (tested on 3.13)
  • Jupyter Notebook or VS Code with Jupyter extension

Step 1 — Install Dependencies

pip install -r requirements.txt

Step 2 — Run Notebooks in Order

Run each notebook top to bottom, in sequence:

01 → 02 → 03 → 04 → 05
jupyter notebook

Each notebook is self-contained and saved with outputs. Simply open and run all cells.


📊 The 5 Collapse Stages

Stage Label What It Looks Like Why It Matters
0 Healthy Diverse, rich vocabulary, no repetition Baseline for "normal" data
1 Slight Drift Minor repetition of common phrases Earliest detectable signal
2 Moderate Degradation Reduced vocabulary, rising repetition Clear drift, model starts to suffer
3 Severe Degradation Heavy repetitive loops, gibberish phrases Strong performance loss
4 Collapse Fully collapsed — monotone or nonsensical Model unusable

These match real-world patterns observed in LLMs trained repeatedly on their own outputs.


🔬 Notebook-by-Notebook Breakdown


📓 01_data_generation.ipynb — Data Generation

What it does:
Creates a synthetic dataset of 5,000 text samples (1,000 per stage). Each stage simulates how model-generated text degrades over time.

How it works:

  • Stage 0 (Healthy): Randomly assembled sentences with diverse vocabulary
  • Stage 1–4: Progressively introduces repetition, removes vocabulary variety, adds gibberish tokens

Why synthetic data?
Real collapse data is hard to obtain at scale. Synthetic data lets us precisely control degradation levels, test all five stages, and avoid copyright/license issues with real corpora.

Output: data/synthetic_dataset.csv — columns: text, label


📓 02_feature_extraction.ipynb — Feature Engineering

What it does:
Converts raw text into 7 quantitative linguistic features that capture collapse signals.

Features extracted:

Feature What It Measures Why It Detects Collapse
TTR (Type-Token Ratio) Unique words ÷ total words Drops sharply as repetition increases
Entropy Information diversity (Shannon entropy) Falls as text becomes more predictable
Average Word Frequency How common words are on average Rises as rare vocabulary disappears
Vocabulary Size Count of unique words per sample Shrinks during degradation
Bigram Diversity Unique 2-word combinations ÷ total bigrams Reflects structural repetition
Avg Sentence Length Mean words per sentence Becomes erratic or very short at collapse
Duplicate Ratio Fraction of repeated sentences Spikes to 1.0 at full collapse

Why these features?
They are computationally cheap, interpretable, and proven effective for detecting text quality degradation. No neural network is required for feature extraction — making the pipeline fast and reproducible.

Output: data/features.csv, data/feature_summary.png


📓 03_model_training.ipynb — Model Training

What it does:
Trains four distinct models on the extracted features to classify collapse stage, then builds an Autoencoder to generate a continuous Collapse Risk Score.

🔵 Logistic Regression (Baseline)

  • Why: Fast, interpretable baseline — shows what's achievable with a linear model
  • Input: 7 extracted features
  • Output: Class prediction (0–4)

🟢 Random Forest

  • Why: Handles non-linear relationships, robust to outliers, gives feature importance
  • Input: 7 extracted features
  • Output: Class prediction + probability

🟡 XGBoost (Best Classifier)

  • Why: State-of-the-art gradient boosting — highest accuracy on tabular data
  • Input: 7 extracted features
  • Output: Class prediction + probability + feature importance

🔴 LSTM (Deep Learning — Temporal)

  • Why: Collapse often develops over time across a sequence of batches; LSTM captures that temporal pattern
  • Architecture: Embedding → LSTM (64 units) → Dense → Softmax
  • Input: Feature vectors as sequences
  • Output: Class prediction

🟣 Autoencoder (Risk Score Generator)

  • Why: Instead of classifying, the Autoencoder learns what "healthy" data looks like. When shown degraded data, it fails to reconstruct it accurately — and that reconstruction error becomes the Risk Score.
  • Architecture: Encoder (7 → 4 → 2) → Decoder (2 → 4 → 7)
  • Training: Trained only on Stage 0 (Healthy) samples
  • Output: Risk Score per sample (0.0 = healthy, 1.0 = collapsed)

Why use an Autoencoder for the score?
Classification gives you a category (Stage 2), but the Autoencoder gives you a continuous, nuanced signal — Stage 2 samples close to Stage 1 will score lower than those near Stage 3. This is more useful for real-time monitoring.

Output: data/predictions.csv (all model predictions + risk scores)


📓 04_evaluation.ipynb — Model Evaluation

What it does:
Loads predictions and generates four comprehensive visualizations to compare models and understand what drives collapse detection.

Visualizations generated:

Plot File What It Shows
Risk Score Distribution risk_score_plot.png How risk scores separate the 5 stages
Model Comparison model_comparison.png Accuracy + F1 for all 4 classifiers
Confusion Matrix confusion_matrix.png Where XGBoost mistakes occur between stages
Feature Importance feature_importance.png Which features XGBoost relies on most

Key finding: XGBoost achieves the highest classification accuracy. The Autoencoder risk scores cleanly separate all 5 stages with minimal overlap.


📓 05_validation_experiment.ipynb — Validation Experiment

What it does:
Answers the critical research question: "Does a higher Risk Score actually mean worse model performance?"

Experiment design:

  1. For each collapse stage, a simulated NLP classification task is run on the text data
  2. Accuracy drops and perplexity (a language model metric — lower = better) increases as stage increases
  3. The average Collapse Risk Score for each stage is plotted against simulated accuracy and perplexity

Why perplexity?
Perplexity is the standard measure of how "surprised" a language model is by text. Collapsed text is highly predictable (low perplexity is actually bad here — it means the model learned repetition, not language).

Result:
A clear, monotonic correlation is observed: as Risk Score increases from Stage 0 → 4, NLP accuracy falls and perplexity rises. This validates that the Risk Score is a meaningful early warning signal.

Output: data/validation_plot.png, summary metrics printed in notebook


🛠 Tools & Libraries — What, Why, How

Library Version What It Does Why Used
numpy any Numerical arrays & math Core of all ML computations
pandas any Tabular data manipulation Load/save CSVs, feature tables
matplotlib any Plotting All visualizations
seaborn any Statistical plots Prettier distribution/heatmap plots
scikit-learn any Logistic Regression, Random Forest, metrics Industry-standard ML library
xgboost any Gradient boosted trees Best-in-class tabular classifier
nltk any Tokenization, n-gram analysis Text preprocessing for feature extraction
torch (PyTorch) any LSTM + Autoencoder Deep learning framework — flexible and well-documented
scipy any KL divergence, statistical tests Used in validation metrics
tqdm any Progress bars Better UX during long loops
transformers any HuggingFace tokenizers (optional) Available for semantic embeddings if extended

📈 Results Summary

Model Accuracy F1 Score
Logistic Regression ~72% ~0.71
Random Forest ~88% ~0.87
XGBoost ~93% ~0.92
LSTM ~80% ~0.79

Collapse Risk Score (Autoencoder):

  • Stage 0 (Healthy): 0.05 – 0.15
  • Stage 2 (Moderate): 0.45 – 0.60
  • Stage 4 (Collapse): 0.85 – 1.00

🧪 Key Design Decisions

Why 5 Stages Instead of Binary?

Binary (Collapsed / Not Collapsed) is too coarse for early warning. A 5-stage taxonomy lets researchers act at Stage 1–2, before irreversible collapse.

Why Autoencoder for Risk Score (Not Just Probability)?

Classifier probabilities are unreliable for out-of-distribution samples. An Autoencoder trained only on healthy data provides a principled anomaly detection score — no labeled collapse data needed at inference time.

Why Synthetic Data?

Gives precise control over degradation severity, ensures reproducibility, and avoids the complexity of sourcing/licensing real collapse datasets.


🔭 Future Extensions

  • Real data integration — replace synthetic data with actual Reddit/WikiText corpora with artificial degradation applied
  • Embedding-based features — add BERT sentence embeddings for semantic drift detection
  • Real-time dashboard — feed the Risk Score into a Streamlit or Gradio monitoring UI
  • Temporal LSTM pipeline — feed batches sequentially to detect drift across training iterations
  • Alert thresholds — define configurable thresholds (e.g., Risk Score > 0.6 → raise alert)

📚 References & Background


👤 Author

Built as part of a research project on AI System Safety & Language Model Robustness.

"The collapse of a language model is not a sudden event — it is a slow, measurable drift. This project measures it."

About

Early Detection of Model Collapse in AI Systems using Statistical and Semantic Analysis

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages