JHU School of Government & Policy

Seth Lazar

Navigating the AGI Reckoning

Professor, JHU SGP

An Overview of the Machine Intelligence and Normative Theory Lab’s work

AGI Alignment, Governance, and Adaptation

01 — Background Assumptions

Background Assumptions

Institutional decline
meets transformative
AI

250-year-old liberal democratic institutions are sliding into a basin of mutually reinforcing crises. Into this comes math that wakes sand up.

The variance of possible outcomes is very high. By the end of the 2030s we could no longer be able to use networked computers; AI-enabled authoritarianism might dominate the world; even worse outcomes can’t be ruled out. Or AI may be our best hope for escaping the civilisational basin we’re currently in.

Key Framing

This is a reckoning for humanity, literally and figuratively

Entirely unclear whether we’ll make it through the transition with our values and institutions intact. But a reckoning is also about figuring out who owes what to whom. AGI will change that too.

02 — Defining AGI

Definitions of AGI

People tie themselves up in knots over terminology and definitions. I say: just spell it out up front. But there are some useful distinctions to draw.

Hassabis’ version

Mechanistic/Cognitive

≥ Human level at all cognitive functions, e.g. Language, perception, planning, reasoning, memory, commonsense, social/moral cognition, metacognition etc

“Jagged frontier” but progress on all fronts

Legg and co

Capabilities

Focus on behaviours not cognitive capacities. Domain-general, ≥expert human level, autonomous (able to succeed at its goals w/o human supervision)

More about what AGI can do than what it is

Outcomes

Outcomes

Alternative approach looks not at what individual systems can do but at the overall impact on society, e.g. Transformative AI

Maybe too ambiguous?

What about “normal technology”. Not remotely plausible to describe AI itself as an intrinsically normal technology. Nothing else has ever come close. We are investing dumb matter with spirit. But will the societal impacts of AI, like any other technology, be determined by social and human factors, as well as the raw power of the technology? Yes for sure.

03 — Concrete Waypoints

Concrete Waypoints

Let’s actually take stock. If you showed anyone back in 2022 Codex or Claude Code, they’d call it AGI. We’ve come a very long way.

Locked in

Agents

Behavioural AGI is very close. Most capable models + harnesses are able to do significant proportion of computer-based work. Note also multi-agent progress.

Key ongoing defects: memory and creativityEssentially a solved problem — Needs only good RL environments and reward signals. There's still a threshold in competence/insight/judgment, but (a) that threshold is higher than most entry-level employees, and (b) every reason to think it will be crossed.

On the horizon

Science

Within-paradigm progress from AI in many different fields (AI research, mathematics, biology) already feasible. Paradigm-shifting not yet.

Creativity again the key bottleneckEarly signs — AI for math is the canary in the coal mine. Modest RSI is already happening — Anthropic's engineers barely write code themselves — but as yet it is a “lossy” form of self-improvement. Source

Can’t rule out

Independent ASI

Autonomous superintelligence. Plausibly beyond current architectures, but don't bet against it.

Surprisingly few Rubicons remainBetween AI-for-science and ASI there may be surprisingly few remaining Rubicons. Embodied AI (robotics) is a separate path, also worth watching. Worth factoring into models but should not be the focus.

Where we are now: GPT-4 was the first inflection point. o1 the next. Opus 4.5 the third. Perhaps the Hugging Face incident was the fourth, for multi-agent systems? I’ve left robotics out here because a software singularity would be enough to motivate everything else in this talk. But physical autonomous systems will also be part of this story. Is *verifiability* the bottleneck? Massive progress is being made in areas where only “soft” verification is possible (e.g. moral and legal reasoning). But this is an open question.

04 — Core Question

MINT Lab’s Research

What if AGI doesn’t kill us all?

AGI might kill us all. Do not downplay that possibility because it seems sci-fi. We are living in sci-fi

On the governance side, preventing this outcome is mostly about implementation and political will. Totally feasible; likely to be stuffed up. Not so interesting for a philosopher (technical AGI alignment, on the other hand, very interesting)

05 — Steering the Transition

MINT Lab’s Research

Steering the Transition

1
Alignment
What reasons should an AI agent comply with? How can we train models that both understand and are motivated to act on the reasons that apply to them?
2
Governance
How can laws and other societal institutions that take AI as their object be changed so as to preserve liberal democratic values in the age of AGI?
3
Adaptation
AGI will affect everything; we’ll need to adapt many institutions (etc) that *don’t* take AI as their explicit object. What principles should guide that adaptation? How can we get started in practice?
06 — Moral Competence
Alignment

Moral Competence for AGI Alignment

A safe superintelligence will, necessarily, be able to robustly understand and abide by the moral reasons that apply to it. Can we get there by leveraging existing LLMs’ analytical moral understanding?

What should we align to?!
The most basic question in AI alignment still needs more work (it always will). What is alignment? I think: AI systems that respond appropriately to reasons that apply to them are aligned. This needs defence. And what reasons apply to AI systems? A whole new field of normative ethics beckons.
Today’s LLMs have a profound understanding of morality
We’ve done philosophy—driven LLM evaluations on this; it’s undeniable. Give an LLM any moral problem and its moral analysis of that problem will be human expert level or better.
But local ≠ global; analytical ≠ behavioural; individual ≠ collective alignment
Responses to some particular problem are one thing. But to trust an agent operating in open-ended environments, their judgments need to fit together sensibly. That coherence is currently lacking. Understanding morality is not the same as behaving morally! Agents demonstrably depart from norms they understand. And individually competent agents might be jointly incompetent. All that needs to change
Training for coherence and knowhow
We can do something about both. Coherence can be a target of optimisation; analytical competence can potentially be translated into practical competence with the right training regime; and we have plans to leverage political philosophy, cognitive science, and cultural evolution to support collective alignment
07 — Decentralisation
Governance

Building a Safe, Decentralised, AI Agent Economy

The transition to powerful AI involves many risks; addressing those risks incentivises centralising authority; if we do so we might successfully avoid AGI’s worst risks, at the expense of acutely undermining core liberal democratic values.

I don’t do much straight-up AI policy. I *do* care a lot about AI governance that preserves core liberal values
First big question: ought we build it at all? Obvious but under-examined; less so now the backlash has begun
How can we advance safely towards AGI without excessively concentrating power?
AI models themselves must be designed to attend to this; they should act as genuine advocates for users asserting their core liberal rights.

Our key projects in this area: technical and policy interventions that support building agent advocates, and avoiding platform agents

08 — Societal Adaptation

Societal Adaptation

Post-AGI Political Philosophy
AGI reckoning = accounting for who owes what to whom. AGI will change both contents and grounds of our mutual obligations. Big questions: should we build it? How ought it behave? How should we treat it? How does what we owe to each other change?
Obstacle-Mapping
Institutional sclerosis has asymmetric effects: it won't slow the most serious risks, but it will block many of the benefits. Can we identify those obstacles ahead of time and clear the path for powerful AI to actually benefit us?
Vulnerability-Scanning
We’re deploying powerful new agents into systems that were designed for humans, with a specific set of capabilities and limitations. We know systemic risks inevitable; where will they land first?
Capacity-Building
If AI were used to massively improve bureaucracies and core service delivery, many democratic concerns about AI would dissipate. Promising avenues: evals and RL environments for public service agents; agentic support for participatory decision-making

Societal Adaptation / Alignment’s Complement

09 — Applied Philosophy

A different kind of applied philosophy.

Applied philosophy usually = taking one’s philosophical understanding and applying it to some practical problem.

We do some of that. But philosophy is also about understanding; the best way to use, test, and advance your understanding of X is to build a better X.

AGI itself enables new kinds of agents, and new kinds of social institutions. No better test for your theories of morality or of justice than the challenge of steering safely into the AGI transition.

Seth Lazar · slazar@jhu.edu
JHU SGP · MINT Research Lab