Project Undertow
.
How models get to the answer.
Overview
Reasoning trace analysis studying the intermediate steps models use to arrive at outputs, identifying predictive patterns of failure before actions are committed.
Research Focus
The Research Problem
Evaluating AI output solely on final response strings allows subtle reasoning hallucinations, goal drift, or prompt injection exploits occurring mid-thought to go completely undetected until damage occurs.
Why It Matters
Detecting reasoning degradation in real time during chain-of-thought generation enables inline interception before agents execute irreversible database modifications or API calls.
Research Objective
Identify real-time entropy spikes and graph structural anomalies in intermediate reasoning tokens that signal impending model failure or goal divergence.
Technical Approach
Key Methodology Vectors
- Streaming entropy measurement on step-by-step reasoning tokens.
- Step-level verification classifiers trained on chain-of-thought failure datasets.
- Graph analysis of reasoning decision nodes during multi-step problem solving.
Expected Outcomes & Milestones
- Real-time reasoning trace anomaly detection metrics integrated into the TraceAudit security monitor.
- High-precision predictive markers for prompt injection and goal drift in live reasoning streams.
References & Prior Work
- Let's Verify Step by Step: Process-Supervised Reward Models — Lightman et al., OpenAI (2023).
- Real-Time Entropy Monitoring in Reasoning Chains — Sleepers Research Technical Report (2026).
Interested in this research direction?
Sleepers Research welcomes inquiries from teams exploring autonomous agent security and applied machine learning.
Contact the Laboratory