Project Fathom
.
Looking inside small models.
Overview
Interpretability research focused on fine-tuned small language models to examine internal weight matrix shifts and activation changes during LoRA/QLoRA adaptation.
Research Focus
The Research Problem
Adapting base models using parameter-efficient fine-tuning (LoRA/QLoRA) is widely used, but the internal mechanistic changes remain opaque. Unintended capability drift, hidden safety regressions, or token distribution anomalies often pass undetected.
Why It Matters
Edge-deployed agents rely heavily on fine-tuned small models (1B-7B parameters). Certification for mission-critical deployment requires empirical verification of internal representation safety.
Research Objective
Map internal weight shifts, attention head reallocation, and activation trajectories in fine-tuned small language models to identify where adaptation introduces capability loss or safety drift.
Technical Approach
Key Methodology Vectors
- Singular value decomposition (SVD) and rank analysis on LoRA adapter weight matrices.
- Layer-wise activation patching and probing on safety-critical token paths.
- Comparative probing between base model representations and fine-tuned checkpoints.
Expected Outcomes & Milestones
- Diagnostic toolkits for detecting hidden alignment regressions in fine-tuned edge models.
- Quantifiable interpretability metrics for parameter-efficient adaptation techniques.
References & Prior Work
- LoRA: Low-Rank Adaptation of Large Language Models — Hu et al., ICLR (2022).
- Mechanistic Interpretability of Parameter-Efficient Fine-Tuning — Sleepers Research Working Paper (2026).
Interested in this research direction?
Sleepers Research welcomes inquiries from teams exploring autonomous agent security and applied machine learning.
Contact the Laboratory