compress
Status: Exploratory

Project Marrow
.

How much intelligence survives compression.

Current StatusExploratory
AttributionSleepers Research Laboratory
Published Date2026-04-01
Last Updated2026-07-20
Abstract

Overview

Efficiency research measuring how much reasoning capability and safety alignment survives model quantization, pruning, and knowledge distillation.

Problem & Scope

Research Focus

The Research Problem

Aggressive model compression (4-bit, 2-bit quantization, structural pruning) reduces hardware footprints but often non-linearly degrades complex reasoning, tool schema adherence, and safety guardrails.

Why It Matters

Deploying high-performance agent middleware on edge or cost-constrained hardware requires knowing the exact precision thresholds where safety policies remain uncompromised.

Research Objective

Map Pareto frontiers comparing model size, quantization format (AWQ, GPTQ, GGUF), and retention of tool-use accuracy and safety alignment.

Methodology

Technical Approach

Key Methodology Vectors

  • Systematic quantization of 3B to 70B parameter models across INT8, INT4, and FP4 formats.
  • Evaluation against complex tool-call schemas and adversarial safety test prompts.
  • Knowledge distillation benchmarking comparing teacher-student capability retention.

Expected Outcomes & Milestones

  • Empirical guidelines for optimal quantization precision without sacrificing agentic tool reliability.
  • Compression recipes optimized for security middleware proxies.
Citations

References & Prior Work

  • AWQ: Activation-aware Weight Quantization for LLM CompressionLin et al., MLSys (2024).
  • Safety Policy Stability Under Model QuantizationSleepers Research Benchmark (2026).

Interested in this research direction?

Sleepers Research welcomes inquiries from teams exploring autonomous agent security and applied machine learning.

Contact the Laboratory