reason
Status: Exploratory

Project Undertow
.

How models get to the answer.

Current StatusExploratory
AttributionSleepers Research Laboratory
Published Date2026-03-01
Last Updated2026-07-20
Abstract

Overview

Reasoning trace analysis studying the intermediate steps models use to arrive at outputs, identifying predictive patterns of failure before actions are committed.

Problem & Scope

Research Focus

The Research Problem

Evaluating AI output solely on final response strings allows subtle reasoning hallucinations, goal drift, or prompt injection exploits occurring mid-thought to go completely undetected until damage occurs.

Why It Matters

Detecting reasoning degradation in real time during chain-of-thought generation enables inline interception before agents execute irreversible database modifications or API calls.

Research Objective

Identify real-time entropy spikes and graph structural anomalies in intermediate reasoning tokens that signal impending model failure or goal divergence.

Methodology

Technical Approach

Key Methodology Vectors

  • Streaming entropy measurement on step-by-step reasoning tokens.
  • Step-level verification classifiers trained on chain-of-thought failure datasets.
  • Graph analysis of reasoning decision nodes during multi-step problem solving.

Expected Outcomes & Milestones

  • Real-time reasoning trace anomaly detection metrics integrated into the TraceAudit security monitor.
  • High-precision predictive markers for prompt injection and goal drift in live reasoning streams.
Citations

References & Prior Work

  • Let's Verify Step by Step: Process-Supervised Reward ModelsLightman et al., OpenAI (2023).
  • Real-Time Entropy Monitoring in Reasoning ChainsSleepers Research Technical Report (2026).

Interested in this research direction?

Sleepers Research welcomes inquiries from teams exploring autonomous agent security and applied machine learning.

Contact the Laboratory