What to watch
- Interpretability methods that scale to frontier models
- Evaluations for autonomous and dangerous capabilities
- Policy and disclosure norms for frontier releases
The state of play
Every assistant model is shaped by alignment techniques: RLHF taught models to follow instructions, and Constitutional AI showed principles could steer training through AI feedback. As models gain autonomy, the questions shift from “is the answer polite?” to “will an agent pursue the goal we meant, under pressure, over many steps?”
This is also where the long-range conversation about superintelligence lives: how to keep systems that may exceed human ability honest, controllable and beneficial.
What works today
- Refusal and policy behaviour trained into models.
- Pre-release evaluations and red-teaming at major labs.
- Early interpretability results that identify human-readable features inside models.
What doesn’t yet
- Robust defence against prompt injection and jailbreaks.
- Reliable ways to verify that a model’s stated reasoning reflects its actual computation.
Open questions
Can oversight scale as fast as capability?
Evidence