Luis M. Montoya

I am an AI safety researcher and ML engineer studying monitorability, situational awareness, and world models in advanced AI systems. I work on cases where behavior, visible reasoning, and internal representations come apart.

Pivotal Research Fellow · MATS 9.0 Exploration Phase (Nanda stream) · MSCS, Georgia Tech

Selected work

Selected work

View all

GamesBench

Open-source benchmark

A benchmark for long-horizon planning in tool-calling LLMs using deterministic games with external environment simulation.