Towards Measuring and Detecting Unverbalized Evaluation Awareness
ICML 2026 MI Workshop posterMeasuring cases where models stop verbalizing evaluation awareness while evaluation-conditioned behavior and internal eval/deploy signals persist.
I am an AI safety researcher and ML engineer studying monitorability, situational awareness, and world models in advanced AI systems. I work on cases where behavior, visible reasoning, and internal representations come apart.
Pivotal Research Fellow · MATS 9.0 Exploration Phase (Nanda stream) · MSCS, Georgia Tech
Measuring cases where models stop verbalizing evaluation awareness while evaluation-conditioned behavior and internal eval/deploy signals persist.
Testing whether LLM agents build target-specific models of others, rely on generic social priors, or project from their own policy.
A benchmark for long-horizon planning in tool-calling LLMs using deterministic games with external environment simulation.
Studying whether pretrained full-attention Transformers can warm-start TTT-E2E models to reduce long-context training cost while preserving quality.