Project Nocturne
.
Intelligence that doesn't need the cloud.
Overview
On-device and local inference research studying latency, resource tradeoffs, and behavioral consistency when running LLMs entirely locally without cloud dependencies.
Research Focus
The Research Problem
Hosted AI cloud APIs introduce network latency, privacy concerns, and availability risks. However, local inference engines struggle with hardware resource contention, variable token generation rates, and state drift.
Why It Matters
Air-gapped enterprise environments and high-security defense applications require local, fully autonomous agent execution nodes that operate reliably without external cloud connectivity.
Research Objective
Evaluate execution stability, latency bounds, and tool execution reliability of local inference runtimes under constrained hardware conditions.
Technical Approach
Key Methodology Vectors
- Benchmarking local inference runtimes (vLLM, llama.cpp, TensorRT-LLM) under dynamic memory loads.
- Latency and jitter testing during concurrent multi-agent tool execution.
- Behavioral consistency evaluation between cloud-hosted base models and local quantized counterparts.
Expected Outcomes & Milestones
- Reference architectures for fully air-gapped, zero-cloud dependency local agent security nodes.
- Latency-optimized local inference configurations for MCP tool proxies.
References & Prior Work
- vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention — Kwon et al., SOSP (2023).
- Air-Gapped Agent Runtimes & Local Execution Security — Sleepers Research Special Report (2026).
Interested in this research direction?
Sleepers Research welcomes inquiries from teams exploring autonomous agent security and applied machine learning.
Contact the Laboratory