interpret
Status: Exploratory

Project Fathom
.

Looking inside small models.

Current StatusExploratory
AttributionSleepers Research Laboratory
Published Date2026-02-15
Last Updated2026-07-20
Abstract

Overview

Interpretability research focused on fine-tuned small language models to examine internal weight matrix shifts and activation changes during LoRA/QLoRA adaptation.

Problem & Scope

Research Focus

The Research Problem

Adapting base models using parameter-efficient fine-tuning (LoRA/QLoRA) is widely used, but the internal mechanistic changes remain opaque. Unintended capability drift, hidden safety regressions, or token distribution anomalies often pass undetected.

Why It Matters

Edge-deployed agents rely heavily on fine-tuned small models (1B-7B parameters). Certification for mission-critical deployment requires empirical verification of internal representation safety.

Research Objective

Map internal weight shifts, attention head reallocation, and activation trajectories in fine-tuned small language models to identify where adaptation introduces capability loss or safety drift.

Methodology

Technical Approach

Key Methodology Vectors

  • Singular value decomposition (SVD) and rank analysis on LoRA adapter weight matrices.
  • Layer-wise activation patching and probing on safety-critical token paths.
  • Comparative probing between base model representations and fine-tuned checkpoints.

Expected Outcomes & Milestones

  • Diagnostic toolkits for detecting hidden alignment regressions in fine-tuned edge models.
  • Quantifiable interpretability metrics for parameter-efficient adaptation techniques.
Citations

References & Prior Work

  • LoRA: Low-Rank Adaptation of Large Language ModelsHu et al., ICLR (2022).
  • Mechanistic Interpretability of Parameter-Efficient Fine-TuningSleepers Research Working Paper (2026).

Interested in this research direction?

Sleepers Research welcomes inquiries from teams exploring autonomous agent security and applied machine learning.

Contact the Laboratory