Professor, JHU SGP
An Overview of the Machine Intelligence and Normative Theory Lab’s work
AGI Alignment, Governance, and Adaptation
250-year-old liberal democratic institutions are sliding into a basin of mutually reinforcing crises. Into this comes math that wakes sand up.
The variance of possible outcomes is very high. By the end of the 2030s we could no longer be able to use networked computers; AI-enabled authoritarianism might dominate the world; even worse outcomes can’t be ruled out. Or AI may be our best hope for escaping the civilisational basin we’re currently in.
Entirely unclear whether we’ll make it through the transition with our values and institutions intact. But a reckoning is also about figuring out who owes what to whom. AGI will change that too.
People tie themselves up in knots over terminology and definitions. I say: just spell it out up front. But there are some useful distinctions to draw.
≥ Human level at all cognitive functions, e.g. Language, perception, planning, reasoning, memory, commonsense, social/moral cognition, metacognition etc
“Jagged frontier” but progress on all fronts
Focus on behaviours not cognitive capacities. Domain-general, ≥expert human level, autonomous (able to succeed at its goals w/o human supervision)
More about what AGI can do than what it is
Alternative approach looks not at what individual systems can do but at the overall impact on society, e.g. Transformative AI
Maybe too ambiguous?
What about “normal technology”. Not remotely plausible to describe AI itself as an intrinsically normal technology. Nothing else has ever come close. We are investing dumb matter with spirit. But will the societal impacts of AI, like any other technology, be determined by social and human factors, as well as the raw power of the technology? Yes for sure.
Let’s actually take stock. If you showed anyone back in 2022 Codex or Claude Code, they’d call it AGI. We’ve come a very long way.
Behavioural AGI is very close. Most capable models + harnesses are able to do significant proportion of computer-based work. Note also multi-agent progress.
Key ongoing defects: memory and creativityEssentially a solved problem — Needs only good RL environments and reward signals. There's still a threshold in competence/insight/judgment, but (a) that threshold is higher than most entry-level employees, and (b) every reason to think it will be crossed.
Within-paradigm progress from AI in many different fields (AI research, mathematics, biology) already feasible. Paradigm-shifting not yet.
Creativity again the key bottleneckEarly signs — AI for math is the canary in the coal mine. Modest RSI is already happening — Anthropic's engineers barely write code themselves — but as yet it is a “lossy” form of self-improvement. Source
Autonomous superintelligence. Plausibly beyond current architectures, but don't bet against it.
Surprisingly few Rubicons remainBetween AI-for-science and ASI there may be surprisingly few remaining Rubicons. Embodied AI (robotics) is a separate path, also worth watching. Worth factoring into models but should not be the focus.
Where we are now: GPT-4 was the first inflection point. o1 the next. Opus 4.5 the third. Perhaps the Hugging Face incident was the fourth, for multi-agent systems? I’ve left robotics out here because a software singularity would be enough to motivate everything else in this talk. But physical autonomous systems will also be part of this story. Is *verifiability* the bottleneck? Massive progress is being made in areas where only “soft” verification is possible (e.g. moral and legal reasoning). But this is an open question.
On the governance side, preventing this outcome is mostly about implementation and political will. Totally feasible; likely to be stuffed up. Not so interesting for a philosopher (technical AGI alignment, on the other hand, very interesting)
A safe superintelligence will, necessarily, be able to robustly understand and abide by the moral reasons that apply to it. Can we get there by leveraging existing LLMs’ analytical moral understanding?
The transition to powerful AI involves many risks; addressing those risks incentivises centralising authority; if we do so we might successfully avoid AGI’s worst risks, at the expense of acutely undermining core liberal democratic values.
Our key projects in this area: technical and policy interventions that support building agent advocates, and avoiding platform agents
Applied philosophy usually = taking one’s philosophical understanding and applying it to some practical problem.
We do some of that. But philosophy is also about understanding; the best way to use, test, and advance your understanding of X is to build a better X.
AGI itself enables new kinds of agents, and new kinds of social institutions. No better test for your theories of morality or of justice than the challenge of steering safely into the AGI transition.
Seth Lazar · slazar@jhu.edu
JHU SGP · MINT Research Lab