01 / The Researchers
Scientists
- Alignment failure — systems that optimize for proxies of human values, not values themselves.
- Loss of interpretability as models scale past human-legible reasoning.
- Recursive self-improvement collapsing the window for course-correction.
- Concentration of capability in a handful of labs with asymmetric incentives.