Project Marrow
.
How much intelligence survives compression.
Overview
Efficiency research measuring how much reasoning capability and safety alignment survives model quantization, pruning, and knowledge distillation.
Research Focus
The Research Problem
Aggressive model compression (4-bit, 2-bit quantization, structural pruning) reduces hardware footprints but often non-linearly degrades complex reasoning, tool schema adherence, and safety guardrails.
Why It Matters
Deploying high-performance agent middleware on edge or cost-constrained hardware requires knowing the exact precision thresholds where safety policies remain uncompromised.
Research Objective
Map Pareto frontiers comparing model size, quantization format (AWQ, GPTQ, GGUF), and retention of tool-use accuracy and safety alignment.
Technical Approach
Key Methodology Vectors
- Systematic quantization of 3B to 70B parameter models across INT8, INT4, and FP4 formats.
- Evaluation against complex tool-call schemas and adversarial safety test prompts.
- Knowledge distillation benchmarking comparing teacher-student capability retention.
Expected Outcomes & Milestones
- Empirical guidelines for optimal quantization precision without sacrificing agentic tool reliability.
- Compression recipes optimized for security middleware proxies.
Related Projects
References & Prior Work
- AWQ: Activation-aware Weight Quantization for LLM Compression — Lin et al., MLSys (2024).
- Safety Policy Stability Under Model Quantization — Sleepers Research Benchmark (2026).
Interested in this research direction?
Sleepers Research welcomes inquiries from teams exploring autonomous agent security and applied machine learning.
Contact the Laboratory