Understand to align, for me, the order matters. Monitoring, oversight, and evaluation all assume we understand what a model's outputs reveal about its policy. Where that assumption fails, where behavior, visible reasoning, and internal representations come apart, safety measurements stop measuring. Understanding is the input that alignment decisions are made of, so I want to make it measurable. I try to turn safety concerns into verifiable behaviors, evaluation harnesses, controls, and results that can update my beliefs. What follows is what that looks like in practice.
Measurement
Given the choice between a demonstration and a clean measurement, I take the measurement. A behavior is a good experimental target when a deterministic evaluator can score it without human judgment, because once scoring requires judgment, disagreements about the result collapse into disagreements about the scorer. This is why I have studied evaluation awareness through type-hint use rather than a safety-adjacent behavior, and long-horizon planning through games with exact win conditions rather than rubric-graded tasks. The safety-flavored version is always more alarming and less believable. This exchange rate between "advertising" results and rigor is real, and I tend to price it deliberately.
Intervention
An observed correlation between two things a model does is a hypothesis about their relationship, and it stays a hypothesis until something is intervened on. Observation tells you the world is compatible with your claim. Intervention tells you the claim survives contact with a counterfactual. In practice this means treating a suggestive split in the data as a reason to build the manipulation, the way an observed gap in non-verbalizing responses became an on-policy resampling experiment rather than a finding on its own. The observational result suggested the claim, while the interventional one grounded it.
Controls
I want for every alternative explanation to have a control that could have caught it, enumerated before the experiment runs rather than added after. The shape repeats across projects. Probes get masked cues, contrastive behaviors that flip across scenarios, disjoint training domains, and frozen weights before transfer. Forecasts of another agent's actions get read against self-prediction, anonymized-context, shuffled-history, and self-projection baselines, all paired on identical frozen environment snapshots. I value really highly the habit of asking what else could produce this result and building the version of the experiment that would tell me.
Disclosure
Results should be reported against interest. If a check can only confirm what I already believe, it is not a check, so the ones that should be focused on are the ones that can shrink the result. When our stricter second auditor flagged 15 of 100 suppressed samples that the primary autorater had passed, that number went in the paper, because the headline claim is only as strong as the strongest attack it has survived, and readers deserve to see it.
Calibration
Every claim should be stated at exactly the strength the evidence supports. A probe is a readout of a context direction, so I never describe it as evidence of what a model believes. A confounded comparison gets its confound named in the same sentence as the hypothesis. The discipline is cheap sentence by sentence, and it compounds, because precision about what a result does not show is what makes the rest of the sentence trustworthy.
Infrastructure
Measurement at the scale these questions need is an engineering problem before it is a research problem. I build tend to build the scalable harness first. Infrastructure must not be looked as overhead on the research. It is what makes a result mean the same thing at scale.