AI Buildout and Political Power
None of the nearly 20 prospective 2028 Democrats queried by Axios endorsed Bernie Sanders's proposed pause in AI development. Holly Otterbein and Alex Thompson reported in Axios's August 23 article "2028 Dems dodge on Bernie's push to pause AI" that roughly half of the possible candidates did not respond. Alexandria Ocasio-Cortez leads a House data-center moratorium bill, Ro Khanna supports construction pauses in Pennsylvania, Michigan and Wisconsin, and Rahm Emanuel favors faster permitting with hyperscalers financing grid upgrades. Amid rising local opposition to data centers, Mark Beall, President of Government Affairs at AI Policy Network, argued on X that guardrails could strengthen public support for construction. Azeem Azhar argued in the August 22 Exponential View essay "The problem with petards" that AI laboratories' catastrophe rhetoric had helped fuel the opposition.
Read more: Responses from the 2028 Democratic field → 497 words · ~2 min
Nearly 20 Democrats gave Axios one direct answer on an AI pause
None endorsed Bernie Sanders’s proposed halt. Rahm Emanuel criticized it, Ro Khanna backed three regional data-center pauses, and roughly half the possible contenders did not respond.
Holly Otterbein and Alex Thompson put three questions to nearly 20 Democrats seen as possible 2028 presidential contenders for the August 23 edition of Axios 2028: whether they back Bernie Sanders's call for AI executives to halt development, whether they support a moratorium on AI data centers, and whether they worry that regulation could cost the United States the AI arms race. One answered the first question directly. About half gave no response at all, among them Kamala Harris, Gavin Newsom and Pete Buttigieg. Nobody endorsed the halt, and by Otterbein and Thompson's account almost nobody criticized it either.
Sanders called for the halt on August 10, in a letter to Sam Altman, Dario Amodei and Mark Zuckerberg reported by Maria Curi, urging them to "stop building machines that humans cannot control" and warning that the Senate would act if the companies did not. Former Chicago mayor Rahm Emanuel supplied the field's only pushback, saying the proposal "is talking around a problem instead of trying to solve it" and asking how the time it buys would be used. He wants permitting reform and hyperscalers paying for the grid upgrades their projects require.
Alexandria Ocasio-Cortez's chief of staff, Mike Casca, addressed the data centers and declined to comment on halting the technology, noting that she is "leading the data center moratorium bill in the House". The measure she and Sanders announced in March would freeze construction and expansion until Congress can certify that AI is safe, that its economic gains reach workers, and that the buildout does not raise utility prices or damage communities. Ro Khanna said he supports pausing construction in Pennsylvania, Michigan and Wisconsin, "places where these centers have been shoved down the throats of communities", a change from his earlier position and a step past the Data Center Bill of Rights resolution he introduced on August 6. Chris Van Hollen's and Cory Booker's teams cited the Power for the People Act, which would put data centers in their own utility rate class and make them pay for the transmission upgrades they force, while Josh Shapiro tightened permitting and JB Pritzker paused tax incentives.
A Fox News poll of 1,003 registered voters, fielded July 17 to 20, found 70 percent opposed to a data center in their own community and 30 percent in favor, and candidates further down the ballot have moved ahead of the presidential field: Michigan Senate Republican nominee Mike Rogers backs a one-year moratorium, Ohio Democratic gubernatorial candidate Amy Acton wants a conditional one, and Pennsylvania Republican nominee Stacy Garrity supports a pause. Adam Carlson, a progressive pollster, called Democrats "overly cautious" and asked what happens if JD Vance claims the moratorium first. Alex Isenstadt reported four days earlier that the Senate Republican campaign arm had privately warned AI companies that data-center anger threatens Jon Husted's Ohio seat, calling it "a sleeper issue for the entire election cycle". Carlson thinks many contenders will wait for the November 3 midterms.
Sources & documents
- 2028ers dodge on AI (Axios 2028, Aug 23, 2026) — Alex Thompson and Holly Otterbein — Canonical assigned source; full 1,730-word newsletter read from the on-disk FeedMe capture. Supplies the survey design (three questions to nearly 20 possible contenders), the one direct answer, the nonresponders (Harris, Newsom, Buttigieg), and every verbatim quote from Emanuel, Casca, Khanna and Carlson.
- Exclusive: Sanders calls for AI development pause — Maria Curi, Axios (Aug 10, 2026) — Precursor. Verified the document the 2028 field was asked about: Sanders's letter to Altman, Amodei and Zuckerberg, the verbatim 'Stop building machines that humans cannot control', and his warning that the Senate would act otherwise. Read via the Yahoo syndication because axios.com returns 403 to non-browser clients.
- Exclusive: Sanders calls for AI development pause (Yahoo syndication of Axios) — The copy actually read for the Sanders letter; byline, date and quotes taken from here. Body links the Axios original.
- Fox News Poll: Voters reject data centers by 40-point margin — Verified the poll Axios cites: 70% oppose / 30% favor a data center in their community, 1,003 registered voters, fielded July 17-20, 2026, Beacon Research and Shaw & Company Research, plus or minus 3 points.
- NEWS: Sanders, Ocasio-Cortez Announce AI Data Center Moratorium Act — Office of Sen. Bernie Sanders — Primary source for the bill Casca invoked: March 25 announcement, immediate moratorium on construction and expansion, and the three conditions for lifting it (AI safe and effective; gains reach workers; no rise in electricity or utility prices, community harm, or environmental destruction).
- Ro Khanna calls for the right to oppose data centers to be protected — Engadget — Verified Khanna's August 6 Data Center Bill of Rights resolution (H.Res. 1471) and its content, establishing what his new Pennsylvania/Michigan/Wisconsin pause position moves past. CNBC and NBC versions both 403'd.
- Van Hollen Leads New Bill to Ensure Americans Aren't Footing the Bill for Big Data Centers — Office of Sen. Chris Van Hollen — Identified and verified the unnamed bill Van Hollen's and Booker's teams cited to Axios: the Power for the People Act, introduced January 15, 2026, Van Hollen lead with Booker among cosponsors; supplied the rate-class and FERC transmission-cost provisions.
- Exclusive: GOP warns AI companies that data centers are politically radioactive — Alex Isenstadt, Axios (Aug 19, 2026) — Debate map. Verified the NRSC's private memo to AI companies on Jon Husted's Ohio seat and the verbatim 'a sleeper issue for the entire election cycle', which answers Carlson's question about Republicans claiming the issue. Read via the Yahoo syndication below.
- Exclusive: GOP warns AI companies that data centers are politically radioactive (Yahoo syndication of Axios) — The copy actually read for the NRSC memo story; byline, date and memo quotes taken from here.
- Governor Shapiro Signs Executive Order on Data Center Development in PA — Commonwealth of Pennsylvania — Verified 'tightened permitting': EO 2026-05, signed August 18, conditions DEP permit review on binding GRID commitments and prior local approval, removes data centers from Fast Track, bans NDAs.
- Gov. Pritzker Pauses New Data Center Tax Incentives — Office of Gov. JB Pritzker — Verified 'paused tax incentives': June 5 direction to DCEO to pause processing Data Center Investment Program agreements from July 1, with existing agreements honored. Axios's 'revoked' overstates this; the piece uses the primary source's wording.
- AI companies hire locally as data-center opposition spreads — Yesterday in AI, Aug 21, 2026 — Continuity link. Anchor verified live. Carries the Shapiro executive order and the corporate response, so the piece points there rather than re-explaining them.
- Midwest data-center disputes reach courtrooms, ballots, and a governor's race — Yesterday in AI, Aug 4, 2026 — Continuity link. Anchor verified live. Carries the down-ballot Michigan and Wisconsin races the piece gestures at without restating.
- New Polling: Battleground Voters Want AI Guardrails — The AI Policy Network — Not used in the body. Read only to check the digest paragraph's Mark Beall attribution; the AI Policy Network's own page gives his title as President of Government Affairs at the AI Policy Network, and the page contains no data-center guardrail figure.
[ collapse ↑ ]
Analysts expect Nvidia's revenue growth to reach 97% in the July quarter. Martin Peers reported in The Information's August 23 article "Nvidia, Salesforce Are in Spotlight This Week" that analysts expect full-year growth to accelerate from 65% to 83%, driven by purchases from hyperscalers, neoclouds and private-equity-backed data centers. First-quarter free cash flow reached $48.6 billion, up 85.7%, and analysts project $213 billion for the fiscal year, compared with $144 billion for Apple. Peers also reported that Nvidia was raising prices as memory became more expensive.
Build American AI said an outside vendor ran two anonymous political meme accounts. Tyler Johnston pointed to the June 4 admission after he and Taylor Lorenz published the June 3 investigation A Pro-AI Super PAC's Secret Meme Sockpuppets. Build American AI, the 501(c)(4) affiliated with Leading the Future, called them parody accounts peripheral to its strategy. Justin Bullock wrote on X that the conduct deepened distrust of pro-industry advocacy. Both accounts remain online but have not posted since the investigation appeared.
Read more: Two meme accounts and an admission → 435 words · ~2 min
Build American AI's admission about two vendor-run meme accounts
Taylor Lorenz and Tyler Johnston linked a pro-AI meme page and a fake anti-AI activist to the super PAC affiliate. It called them parody accounts; neither has posted since the investigation ran.
On X, Justin Bullock told the AI industry to “Behave like adults rising to the moment, and you’ll be treated as such,” endorsing a post by Nathan Calvin, general counsel at Encode, who wrote that an industry wondering why the public distrusts it could start with the pro-AI super PAC Leading the Future. Calvin was reaching back to the first week of June, when its dark-money affiliate, Build American AI, conceded that an outside vendor ran anonymous accounts attacking AI safety advocates.
Taylor Lorenz and Tyler Johnston published “A Pro-AI Super PAC’s Secret Meme Sockpuppets” on June 3, on The Midas Project’s Model Republic in collaboration with Lorenz’s User Mag, tying two anonymous X accounts to Build American AI, the 501(c)(4) affiliated with Leading the Future. The super PAC had raised more than $125 million by then, including $25 million from OpenAI president Greg Brockman and his wife Anna and $25 million from Marc Andreessen and Ben Horowitz. @DoomersAreDumb appeared last July, ten days after Build American AI filed articles of incorporation in Nevada, and posted generic memes until October, when it turned to AI. @JonathanDoomer switched to AI the same day, then presented itself as an engineer who had left a twenty-year career to warn about the technology, argued that China should lead AI development, and answered one AI warning with an assault-rifle meme captioned “we don’t call 911”, all while @DoomersAreDumb mocked it. Both accounts promoted Build American AI posts that had fewer than 500 views. OpenAI’s Jason Kwon and Josh Vlasto of Leading the Future follow both, and much of their early engagement came from Jason Levin, whose firm Memelord Technologies the reporters identified as the operator. Leading the Future declined to clarify the relationship, and OpenAI declined to comment.
Build American AI answered Johnston on June 4: “These are parody meme accounts run by an outside vendor,” it wrote, and they “are not a core part of our strategy.” Johnston replied that establishing the link had taken real digging and that the accounts’ own repliers had not read them as satire. Three days before the exchange, OpenAI published its standards for political advocacy, asking groups to “not use tactics like astroturfing that obscure the real choices” and stating that the company “does not direct the activities of LTF, or have visibility into their operations.”
Both meme accounts are still online, and neither has posted since June 3, the day the investigation ran. @DoomerDaylight, the earlier account Johnston tied in April to the Webflow-built websites of Leading the Future and Build American AI, last posted May 6.
Sources & documents
- Justin Bullock on X: "Behave like adults rising to the moment, and you'll be treated as such." — Canonical assigned source, read from the on-disk fetch and confirmed via Bird. Supplies the verbatim quote and the fact that Bullock quote-posted Calvin. Posted 2026-08-24 14:53 UTC; contains no new information of its own, which is why the piece chases the chain.
- Nathan Calvin on X on the AI industry, distrust, and Leading the Future — Read via Bird. The post Bullock endorsed, 2026-08-24 14:42 UTC. Bird also exposed the quoted item's true timestamp: Calvin was quoting Tyler Johnston's post of 2026-06-04 02:23 UTC, which establishes that nothing in the exchange is new.
- A Pro-AI Super PAC's Secret Meme Sockpuppets — Taylor Lorenz & Tyler Johnston, Model Republic (The Midas Project) in collaboration with User Mag, June 3, 2026 — Primary reporting, read in full. Supplies: the July incorporation-plus-ten-days timing for @DoomersAreDumb, the October 22 pivot shared with @JonathanDoomer, the sub-500-view Build American AI posts, Jason Kwon and Josh Vlasto following both accounts, Jason Levin and Memelord Technologies, the @JonathanDoomer China quote and the assault-rifle 'we don't call 911' post, the $125M raised with $25M each from the Brockmans and from Andreessen and Horowitz, and the declined-to-clarify and declined-to-comment responses.
- Build American AI on X: "These are parody meme accounts run by an outside vendor…" — The admission itself, read verbatim via Bird, 2026-06-04 02:08 UTC. Note the statement came from Build American AI, the 501(c)(4), not from the Leading the Future account.
- Tyler Johnston on X replying to Build American AI, June 4, 2026 — Read via Bird. Supplies Johnston's rebuttal that the link took substantial digging and that repliers did not read the accounts as satire; its quoted item is the 'Welp, there it is' post Calvin later resurfaced.
- Tyler Johnston on X announcing the June 3 investigation — Read via Bird, 2026-06-03 15:03 UTC. Dates publication and carried the investigation's full text in its quoted Midas Project post, which is how the article was read end to end.
- Our views on AI policy and political advocacy — OpenAI, June 1, 2026 — Primary document, read in full via the Internet Archive snapshot of 2026-08-11 after openai.com returned 403 to plain HTTP and the managed browser profile. Supplies the verbatim astroturfing standard and the 'does not direct the activities of LTF, or have visibility into their operations' line, plus the June 1 date three days before the admission.
- Is this OpenAI's anonymous Twitter sockpuppet account? — Tyler Johnston, Model Republic, April 7, 2026 — Read for the earlier @DoomerDaylight investigation. Supplies the verbatim finding that LeadingTheFuture.com, BuildAmericanAI.org, ThinkBigPac.AI and AmericanMission.com were all built in Webflow, and the roughly $385,000 paid to Prusik Partners LLC in Henderson, Nevada (the dollar figure was verified but cut for space).
- @DoomersAreDumb on X — Original verification. Bird timeline shows the last post at 2026-06-03 15:03 UTC, and a Bird search for from:doomersaredumb since:2026-06-04 returns zero results. The account is still live.
- @JonathanDoomer on X — Original verification. Last post 2026-06-03 16:04 UTC; from:jonathandoomer since:2026-06-04 returns zero results. Account still live.
- @DoomerDaylight on X — Original verification. Last post 2026-05-06; from:DoomerDaylight since:2026-05-07 returns zero results.
- Team — Encode — Primary institutional source verifying Nathan Calvin's current title, general counsel and VP of state affairs. Only 'general counsel' is used in the piece.
- OpenAI isn't being consistently candid about Leading the Future — Veronica Irwin, Transformer, June 11, 2026 — Read for corroboration of the sequence: OpenAI's June 1 standards, then Build American AI's 'outside vendor' acknowledgment days later. No sentence in the piece rests on it alone.
[ collapse ↑ ]
Two California policy advocates argue that existing AI rules already deter investment. In the August 22 Orange County Register opinion article "Hardly Any AI Regulations? Becerra Should Know Better," Bryce Chinault of the Abundance Institute and Lance Christensen of the California Policy Center cite enacted state laws, roughly two dozen pending bills, automated-decision rules and local data-center restrictions.
Normative Competence and Evaluation
ReasonBench tests whether an evaluator changes its verdict for the stated reason. Following work on models' handling of underspecified legal questions, Ye Chen of Alibaba Group and Weining Zhang of Cheung Kong Graduate School of Business introduce No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators, an August 21 arXiv preprint. The researchers vary an evaluator's declared grounds, norms and authority across an eight-cell counterfactual cube, then define a "judgment receipt" as the minimal set of replacements that reproduces a revised verdict. ReasonBench contains 19,520 policy and logical-reasoning cases plus 7,200 controls. Qwen3-1.7B reached 98.41% exact receipt accuracy, but meaning-preserving changes to source order reduced valid receipt recovery to 54.8% for direct prediction. Training on simple source changes preserved 93.75% verdict accuracy while receipt recovery fell to 7.16% on multi-source changes.
Read more: Receipt accuracy under consistency controls → 488 words · ~2 min
Near-ceiling receipts beside weak consistency in ReasonBench
Ye Chen and Weining Zhang froze the hypothesis that denser counterfactual supervision improves reason-tracking, then published its rejection; near-ceiling receipt accuracy on the locked split coexists with violations of roughly half the required consistency relations.
In the August 21 arXiv preprint No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators, Ye Chen of Alibaba Group and Weining Zhang of Cheung Kong Graduate School of Business froze their favored hypothesis under a content hash before opening a single result, and it lost. They had predicted that training a model to emit all eight cells of the counterfactual cube would beat training it to emit the receipt alone, on the theory that denser supervision teaches reusable receipt structure. Across five seeds the cube target finished 1.42 points behind on exact receipt accuracy and 1.85 points behind on revised verdicts, with the loss concentrated in the six organizational policy clauses of the locked split, where it reached 5.69 points. A three-seed Qwen3-0.6B run reproduced the ordering. The paper keeps the rejected hypothesis visible and builds its case on the failures that comparison exposed.
ReasonBench's policy half translates 45 audited airline, retail and telecom clauses from Victor Barres and colleagues' June 2025 arXiv preprint τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment into typed executable rules. The logical half samples 800 worlds from the RuleTaker generator that Peter Clark, Oyvind Tafjord and Kyle Richardson introduced in “Transformers as Soft Reasoners over Language” at IJCAI 2020. ReasonBench gives authority a separate interface because it determines whether a norm may apply: when jurisdiction fails, the adjudicator returns refer and escalates the case without ruling on its merits.
Chen and Zhang give every locked case three paired transformations with a known required relation between the parent and transformed predictions. When a predicted pair violates that relation, at least one prediction is wrong, and the finding holds without any label for the transformed item. They derive from this that 54.77% order consistency forces mean pairwise error of at least 22.6%. Direct receipt preserves only 47.22% of the required relations after an inverse transition but 98.63% after adding an irrelevant record. They read the order fragility as a version-transition analogue of the position bias Lianmin Zheng and colleagues documented in LLM judges.
Chen and Zhang added a composition holdout after the freeze, training on zero- or one-source changes and testing on two or three. With greater proof depth neither target recovers a single changed case's receipt family, while direct prediction keeps 99.83% revised-verdict accuracy. Retraining with randomly permuted source sections lifts order consistency from 55.85% to 96.61%. The improvement transfers to an alternate phrasing absent from training, adding 7.42 points; the untargeted inverse relation gains 0.36.
The authors scope the work tightly: six independent clauses cannot stand for organizational policy at large, the logical stratum is templated, the adjudicators are built to be executable, and the formal results assume determinism. They recommend certifying an account by executing the versioned components, and treating a small model's receipts as predictions whose consistency, coverage and reason debt are measured against those executions. “A structured output is not structured understanding,” the conclusion says.
Sources & documents
- No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators - Ye Chen and Weining Zhang, arXiv — Primary source, read in full via the arXiv HTML rendering at https://arxiv.org/html/2608.20938v1 (all 11 sections plus references). Supplies the frozen-hypothesis framing and SHA-256 content freeze (Sec. 1, 6.3); cube-minus-direct effects of -1.42 pp on exact receipt accuracy and -1.85 pp on revised verdicts, and the -5.69 pp organizational-clause concentration (Sec. 7.2); the three-seed Qwen3-0.6B replication (Sec. 7.9); the 45 tau-squared-Bench clauses and 800 RuleTaker worlds (Sec. 5.1, 5.2); the refer/jurisdiction gate (Sec. 2.2); the three paired controls and Proposition 5, with 54.77% order consistency implying at least 22.6% mean pairwise error, 47.22% inverse and 98.63% distractor rates (Sec. 3.5, 5.3, 7.5); the composition holdout, where changed-case receipt recovery is 0.00% at greater depth against 99.83% revised-verdict accuracy (Sec. 7.8); permutation-augmented retraining lifting order consistency from 55.85% to 96.61%, +7.42 pp on the untrained linguistic rendering and +0.36 pp on the untargeted inverse relation (Sec. 7.6); the limitations (Sec. 10); the prediction/certification recommendation (Sec. 8.3, 11); and the verbatim conclusion quote. Author affiliations (Alibaba Group; Cheung Kong Graduate School of Business) are taken from the paper's own title block.
- tau-squared-Bench: Evaluating Conversational Agents in a Dual-Control Environment - Victor Barres, Honghua Dong, Soham Ray, Xujie Si and Karthik Narasimhan, arXiv — Verified the provenance of ReasonBench's policy stratum: title, author list, June 9 2025 submission date, and that the benchmark covers retail and airline domains alongside its new telecom domain. Abstract page and HTML full text both checked.
- Transformers as Soft Reasoners over Language - Peter Clark, Oyvind Tafjord and Kyle Richardson, IJCAI 2020 — Verified the RuleTaker source that ReasonBench's logical stratum samples from: title, authors, and that the work supplies synthetically generated fact-and-rule worlds in natural language. Venue (IJCAI 2020) taken from the assigned paper's reference list.
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena - Lianmin Zheng et al., NeurIPS 2023 Datasets and Benchmarks Track — Verified the position-bias finding that Chen and Zhang name as the precursor pathology their source-order result parallels (their Sec. 9.1). Confirmed title, lead author, venue, and that the paper identifies position, verbosity and self-enhancement biases in LLM judges.
[ collapse ↑ ]
MOSAIC asks whether social inference produces coordinated behavior. Tonglin Yan, Grégoire Sergeant-Perthuis and David Rudrauf of Université Paris-Saclay and Sorbonne Université present Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models, an August 21 arXiv preprint. The benchmark places two embodied agents in cooperative and competitive scenarios combining speech, movement, gaze and facial expression. Across 200 trials per model, no vision-language model scored significantly above chance in any condition. PCM-LLM, a reference architecture with belief tracking outside the language model, scored above chance throughout. The comparison covers open-source models capped at 15 billion parameters and includes no human baseline.
Read more: MOSAIC's signal-production and signal-reading failures → 483 words · ~2 min
MOSAIC: social inference without coordinated behavior
Two embodied agents, four signal channels, 200 trials per model: Tonglin Yan and colleagues found no vision-language model scoring significantly above chance at turning a theory-of-mind instruction into coordinated behavior, and one hybrid architecture that cleared every condition.
Tonglin Yan, Grégoire Sergeant-Perthuis, and David Rudrauf, of CIAMS at Université Paris-Saclay and CQSB at Sorbonne Université, posted Belief Without Behavior to arXiv on August 21. Multimodal theory-of-mind benchmarks such as MuMA-ToM ask models questions about video of other people; MOSAIC makes the model do the acting. Two embodied agents share a Unity environment holding two similar boxes, one with a reward. The subject agent, called Marie, knows which box; the participant does not, and after ten alternating rounds the box closer to the participant opens. Every third round adds a spoken exchange. Otherwise each agent receives a first-person render and a structured belief state, then returns a movement, a gaze direction, and an emotional expression across musculoskeletal, physiological, and felt channels.
Yan and colleagues run four 50-trial conditions crossing interaction mode with the reasoning assigned to the subject. Under ToM-0 it acts on its own preferences, so even in competition it walks toward the reward and leaks the location. Under ToM-1 it represents what the participant believes, which in competition means inverting its cues to steer the participant wrong. ToM Outcome Conformance scores every trial against the outcome theory predicts for its condition, with 0.5 as chance; positional bias runs beside it, since an agent that always turns the same way scores on where the reward landed.
Across 200 trials per model, no vision-language model scored significantly above chance in any condition. Among the VLMs every binomial test that reached significance pointed below it, Qwen3-VL-8b between 0.20 and 0.32, InternVL3.5-4b at 0.26 and 0.30. LLaVA-7b, LLaVA-13b and Qwen3-VL-2b could not move their agents at all in most trials. PCM-LLM, a hybrid architecture whose belief tracking sits outside the language model, scored 0.70 to 0.92.
Signal-level metrics separate two failures. InternVL3.5-14b produced high trajectory and gaze alignment in cooperative conditions while holding a flat face, a “poker-face strategy” in the authors’ phrase, and its participants ignored those cues, leaving signal sensitivity near zero. PCM-LLM cleared both stages, its sensitivity flipping from 0.90 under competitive ToM-0 to -0.44 when instructed to deceive. Scheirer-Ray-Hare tests returned no effect in 174 comparisons: no VLM varied its trajectory, gaze, or signal sensitivity with the ToM constraint, while PCM-LLM varied all three. The VLMs moved facial expressivity between cooperation and competition.
Example traces show where the coupling fails. One subject stated a preference for the light brown box and an intention to hint at it, then output “right forward” while the box lay to its left. A participant assigned +0.5 to the dark brown box at timestep zero, before Marie had moved, and never revised. Removing visual input from InternVL3.5-8b improved its directional signals in cooperative conditions.
The authors designed PCM-LLM as a feasibility reference; their 2025 working paper already had PCM agents winning a 3D strategic game through deception. The study restricts the comparison to open-source models capped at 15 billion parameters and includes no human baseline.
Sources & documents
- Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models — Yan, Sergeant-Perthuis, Rudrauf, arXiv:2608.20975v1 — Canonical assigned source and center of gravity. The on-disk fetch held only the abstract page, so the full paper was fetched and read end to end (main text plus Appendices A, C, F, G, H). Supplies: the Unity two-box setup, Marie/participant roles, 10 rounds with verbal exchange at R3/R6/R9, the action/verbal/preference-update modules, the 2x2 design at 50 trials per cell, TOCS with 0.5 chance and the positional-bias diagnostic, the signal metrics (TAS, GAS, FES, SSS), PCM-LLM's 0.70-0.92 TOCS, InternVL3.5-14b's high TAS/GAS with near-zero SSS, the 'poker-face strategy' phrase, PCM-LLM's SSS inversion from 0.90 to -0.44, the InternVL3.5-8b visual-ablation result, and every stated limitation.
- Belief Without Behavior — arXiv HTML (LaTeXML) full text — The rendering actually read for full text and for the numeric tables. Table 3 (binomial tests, H0: TOCS=0.5, n=50 per cell) supplies every per-model TOCS cited and confirms that only PCM-LLM lands significantly above chance while all significant VLM cells fall below it. Table 5 (Scheirer-Ray-Hare, Bonferroni-corrected) supplies the 174 non-significant comparisons, the absence of any VLM ToM effect on TAS/GAS/SSS, PCM-LLM's ToM effects on all three, and the VLMs' facial-expressivity-by-mode effects. Appendix H supplies both reasoning-trace failures and the verbatim 'right forward' action output.
- MuMA-ToM: Multi-modal Multi-Agent Theory of Mind — Shi, Ye, Fang, Jin, Isik, Kuo, Shu, arXiv:2408.12574 — Verified comparison point for the opening contrast, and one of the benchmarks MOSAIC tabulates in its Appendix B comparison. Read to confirm it is question answering over video and text of household multi-agent behavior, asking about goals, beliefs, and beliefs about goals, with the model positioned as observer rather than actor.
- PCM-LLMS: a hybrid architecture to enhance human-like social intelligence in virtual agents — Yan, Sergeant-Perthuis, Rudrauf, HAL hal-05054281 — The reference architecture's own working paper, cited in MOSAIC's bibliography. The HAL landing page is behind an Anubis proof-of-work wall that blocked both plain HTTP and the headless OpenClaw profile, so the record was read through HAL's own API (api.archives-ouvertes.fr, halId_s:hal-05054281). Verified: identical three-author list, deposited 2 May 2025, and the abstract's claim that in a 3D strategic game PCM agents with higher-order theory of mind and controlled nonverbal expression 'gained significant advantages through deception'. Also confirmed the lab names behind the paper's acronyms (CIAMS; Biologie Computationnelle, Quantitative et Synthétique, ex LCQB).
- Combining the Projective Consciousness Model and Virtual Humans to assess ToM capacity in Virtual Reality: a proof-of-concept — Rudrauf, Sergeant-Perthuis, Tisserand, Monnor, Belli, arXiv:2104.07053 — Read for lineage: the 2021 precursor in which this group used PCM-driven virtual humans in VR to classify behavior by theory-of-mind order. It established that MOSAIC's environment descends from the group's own prior paradigm. Not cited in the body, where the 2025 PCM-LLM working paper carries the point more directly; available if the desk wants a deeper lineage link.
[ collapse ↑ ]
Claude's constitutional rewrite may add case-law analogues and human adjudicators. An August 22 debate over internal AI courts raised questions about independence and who would bring the hardest cases. Tyler Cowen wrote in the August 23 Marginal Revolution post "My Recent Visit to Anthropic" that he spent two days advising Anthropic on a rewrite. He proposed case-law analogues, interpretive commentary, secondary literature, diverse model panels and a human board with authority over remedies and constitutional changes.
Read more: A constitutional court of models and humans → 422 words · ~2 min
Case law enters Claude's constitution rewrite
Tyler Cowen says Anthropic convened a small invited group for two days of guidance on Claude's constitution. He pressed for case-law analogues, a secondary literature, model panels to check compliance, and a human board with authority over remedies and amendments.
In an August 23 Marginal Revolution post, Tyler Cowen wrote that he had spent two days advising Anthropic on rewriting the constitution for Claude, in a small invited group given "serious time with key decision-makers." He set out five points he pressed there. Whatever a constitution turns out to mean for a model, it should borrow more from analogues to case law and the common law, and its drafters should think in terms of a "Talmud" alongside a "Torah." Cowen also called for a quality secondary literature on AI constitutions, which he says does not exist. Panels of diverse models, run on different inputs, could judge how far Claude and other systems act in accord with their constitutions. A final board of human adjudicators, "functioning in a manner analogous to an independent judiciary," would take those panels' warnings and hold "authority over potential changes and remedies."
In the August 7 Lawfare essay Courts for AI Constitutions, one of two pieces Cowen linked, Nathan Darmon and Tom Reed proposed that a lab build an internal court whose written opinions accumulate into a "synthetic common law" and then train the next model. Andy Hall answered on August 22, pressing them on independence and on who would bring the hardest cases. Darmon and Reed rely on users to flag responses and full-time lab staff to decide cases; Cowen would assign detection to model panels, including tests of Claude and other models against their respective constitutions, and give remedies to a human board.
Anthropic published the current text in January, saying it had sought feedback from outside experts while writing it and would "likely continue to do so for future versions of the document," naming law, philosophy, theology and psychology among the disciplines. In the same announcement, Anthropic said it hoped "an external community can arise to critique documents like this." Cowen's proposed secondary literature could supply that criticism. The constitution anticipates the interpretive problem in its own text: it expects to be "unclear, underspecified, or even contradictory in certain cases," tells Claude to fall back on "the spirit of the document," and calls itself "a perpetual work in progress."
In a comment on Cowen's post, Virginia Postrel described herself as a fellow participant, reported "a strong consensus that a more bottom-up approach is needed," and likened the current constitution to a Hollywood show bible written for a single character instead of a whole series. Cowen's proposals would move part of constitutional interpretation beyond the company that drafted the document.
Sources & documents
- My recent visit to Anthropic - Tyler Cowen, Marginal Revolution — Assigned canonical source. Read in full from the on-disk browser fetch record and re-read live over plain HTTP. Supplies the two-day session, the small invited group, all five numbered proposals, and the verbatim quotes 'serious time with key decision-makers', the Talmud/Torah line, 'functioning in a manner analogous to an independent judiciary', and 'authority over potential changes and remedies'. The live page's embedded comment JSON also supplies the Virginia Postrel comment (id 161070432, dated 2026-08-23 16:22:12) verbatim, including 'a strong consensus that a more bottom-up approach is needed' and the Hollywood show bible comparison.
- Courts for AI Constitutions - Nathan Darmon and Tom Reed, Lawfare — One of the two documents Cowen links. Full text read via plain HTTP extraction. Verified the August 7, 2026 date, the authors, the internal-court design (user flag button, automated triage, five to seven full-time staff, opinions as training data and live corpus), and the verbatim phrase 'synthetic common law'.
- Andy Hall on X, August 22, 2026 — The other document Cowen links. Full thread read through the authenticated Bird conversation reader. Verified the August 22 date and Hall's three objections (whether a human court beats model self-review, independence, and whether users will surface the hardest cases), which the piece summarizes in one clause rather than re-explaining.
- Claude's new constitution - Anthropic — Institutional background, read via plain HTTP extraction. Verified verbatim: 'We'll likely continue to do so for future versions of the document, from experts in law, philosophy, theology, psychology, and a wide range of other disciplines' and 'we hope that an external community can arise to critique documents like this'. Page displays Jan 22, 2026; embedded metadata carries a 2026-01-21 creation timestamp, so the piece says only 'in January'.
- Claude's Constitution - Anthropic — Primary document, read via plain HTTP extraction. Verified verbatim: 'unclear, underspecified, or even contradictory in certain cases', 'we want Claude to use its best interpretation of the spirit of the document', and 'It is best thought of as a perpetual work in progress.'
- Internal AI courts raise evidence and independence questions - Yesterday in AI, 22 August 2026 — Prior coverage of the Darmon and Reed proposal and Hall's critique. Live hosted anchor verified with an HTTP 200 and an id match on the page. Linked so the arc context sits behind one clause instead of being re-explained.
- Model constitutions already govern models and users, Caputo argues - Yesterday in AI, 21 August 2026 — Prior coverage of the argument that closed model constitutions lack interpreters outside the company that writes them. Live hosted anchor verified with an HTTP 200 and an id match. Linked in the closing sentence.
[ collapse ↑ ]
Claude Opus 5 often revised answers after a simple challenge, Nathan Benaich reported. Alongside earlier agent failures during evaluation, Benaich wrote on X that the model, running on high settings, had often withdrawn or changed an answer after he asked whether it was sure or challenged a claim.
Post-AGI Safety and Agent Control
RESI will pursue formal guarantees for superintelligence safety. Following calls for independent verification and enforceable guardrails, the Institute for Responsible Superintelligence announced its creation on August 24 and published a research agenda modeled partly on modern cryptography. Researchers will specify properties and assumptions, construct mechanisms with analyzable guarantees and identify objectives that cannot be guaranteed. RESI distinguishes tests and red teams, which locate individual failures, from guarantees covering classes of failures. Its program includes mapping achievable and impossible objectives, designing protocols whose properties survive composition across models, tools, people and institutions, and implementing those mechanisms in working systems.
Read more: RESI's research lineage in four papers → 473 words · ~2 min
From cryptographic limits to RESI's safety program
The Cambridge nonprofit pairs theorists, economists and lawyers with results on input filtering, hidden backdoors and hallucination incentives; cofounder Adam Tauman Kalai left OpenAI in July during a safety-team reorganization.
The Institute for Responsible Superintelligence introduced itself on X on August 24 and published its agenda at resi.org. Shafi Goldwasser, Vinod Vaikuntanathan and Adam Tauman Kalai founded it. Goldwasser shared the Turing Award for zero-knowledge proofs and directs the Resilience research pod at Berkeley's Simons Institute, which she led from 2018 to 2024; Vaikuntanathan, a Gödel Prize winner at MIT, built much of lattice-based cryptography and fully homomorphic encryption. Kalai wrote on July 15 that he had left OpenAI the previous day, days after Engadget relayed Wired's report that OpenAI had moved its safety teams under a new research and safety VP and that safety systems head Johannes Heidecke was departing.
RESI grounds its agenda in earlier results from its founders and colleagues. Sarah Ball and colleagues argue in “On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment,” accepted at ICLR 2026, that some language models admit no efficient input filter: under cryptographic hardness assumptions, harmful and benign inputs can be computationally indistinguishable. Shafi Goldwasser and colleagues show in “Planting Undetectable Backdoors in Machine Learning Models,” at FOCS 2022, that a trainer can plant a backdoor no computationally bounded observer detects without the key. Adam Tauman Kalai and colleagues argue in the April Nature paper “Evaluating large language models for accuracy incentivizes hallucinations” that accuracy-based scoring rewards guessing over admitting uncertainty; they propose open rubrics stating the penalty for a wrong answer. In “Consensus Sampling for Safer Generative AI,” presented at IASEAI 2026, Kalai and colleagues aggregate several models so the combination carries the risk of the safest few and abstains when agreement is too weak.
RESI’s roster reaches beyond cryptography. Yannai Gonczarowski works on mechanism design at Harvard; Noam Kolt, who teaches law and computer science at the Hebrew University, has “Superintelligence and Law” forthcoming in the Harvard Journal of Law & Technology; Rebecca Wexler, Columbia's Bressler Professor of Law, works on evidence; MIT's Daniela Rus brings safe control for embodied systems. Planned visitors include Scott Aaronson, OpenAI's Boaz Barak and Sébastien Bubeck, and Anthropic's Nicholas Carlini. Working group leaders and topics are unnamed so far. Research will use frontier agents for ideation, stress-testing definitions and checking proofs.
RESI is fiscally sponsored by the Edward Charles Foundation and says its funders will be announced shortly. Some constructions would require training new frontier models, which it says “requires resources well beyond what RESI expects to have initially”; those it will hand to developers as proofs of concept. Resolution, the alignment lab Geoffrey Irving cofounded, sits on RESI's list of affiliated groups, and Irving said in July that Resolution holds a $160 million grant from Coefficient Giving. RESI has named no figure of its own and sets its success condition at frontier developers adopting its mechanisms, at the level represented by breakthrough public-key encryption.
Sources & documents
- RESI Overview — Institute for Responsible Superintelligence — Primary source, read in full from the on-disk fetched text and re-confirmed live. Supplies the founders and their descriptions, Cambridge location, the three output types, the named seed results, the working-group directions and the note that leaders and topics are not yet set, the AI-assisted research plan, the visitor and affiliated-group lists, the Edward Charles Foundation fiscal sponsorship and pending funder list, the success condition, and the verbatim quote about resources beyond what RESI expects to have initially.
- RESI launch post — @RESI_org on X — Verified via authenticated Bird fetch: the institute's own announcement, posted August 24, 2026 at 17:59 UTC, naming Goldwasser, Kalai and Vaikuntanathan as founders. Establishes the launch date and channel.
- Adam Tauman Kalai on leaving OpenAI — @adamfungi on X — Verified via authenticated Bird fetch: posted July 15, 2026, 'Yesterday, I left OpenAI... I'm excited to be working on a new AI safety nonprofit. More details soon.' Dates his departure to July 14 and links it to the institute he later announced.
- Shafi Goldwasser — Simons Institute for the Theory of Computing, UC Berkeley — Primary institutional source for current title: Research Director for the Resilience Research Pod and C. Lester Hogan Professor in EECS at UC Berkeley; director of the Simons Institute from 2018 to 2024.
- OpenAI's head of safety is reportedly leaving as part of company reorganization — Engadget — Read in full. Relays Wired's July 2026 report that Johannes Heidecke was leaving as head of safety systems, that safety teams would report to Mia Glaese as VP of research and safety, and that Saachi Jain became interim head of safety systems. Cited as the relay, with Wired named as the originating report.
- On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment — Ball, Gluch, Goldwasser, Kreuter, Reingold, Rothblum (arXiv:2507.07341) — Abstract read. Supplies the no-efficient-prompt-filter result, the computational indistinguishability of adversarial from benign prompts, the cryptographic hardness assumptions, and the conclusion that filters external to architecture and weights cannot deliver safety. RESI's roster lists this work as 'Computational Barriers to Filtering for AI Alignment, ICLR 2026'; the arXiv preprint carries the longer title.
- Planting Undetectable Backdoors in Machine Learning Models — Goldwasser, Kim, Vaikuntanathan, Zamir (arXiv:2204.06974, FOCS 2022) — Abstract read. Supplies the undetectable-backdoor result: without the backdoor key, the mechanism cannot be detected by any computationally bounded observer.
- Evaluating large language models for accuracy incentivizes hallucinations — Kalai, Nachum, Vempala, Zhang, Nature 653, 1047-1051 — Title, full author list, journal, volume, pages and abstract verified from the article page metadata (published online April 22, 2026). Supplies the claim that accuracy-based headline metrics reward guessing over admitting uncertainty and the proposal of open rubric evaluations that state the penalty for a wrong answer.
- Consensus Sampling for Safer Generative AI — A. T. Kalai, Y. T. Kalai, Zamir (arXiv:2511.09493) — Abstract read. Supplies the aggregation result: risk competitive with the average risk of the safest s of k models, with abstention when agreement is insufficient.
- Superintelligence and Law — Noam Kolt, SSRN — Verified: forthcoming in the Harvard Journal of Law & Technology, posted February 2026; confirms Kolt's Hebrew University law and computer science appointment.
- Rebecca Wexler — Columbia Law School faculty — Primary institutional source verifying her current title, Alfred W. Bressler Professor of Law.
- Geoffrey Irving on Resolution's Coefficient Giving grant — @geoffreyirving on X — Verified via authenticated Bird fetch: posted July 6, 2026, announcing a $160 million grant from Coefficient Giving, $108 million unconditional and $52 million conditional. Used for the comparison figure only.
- Garrison Lovely quote post relaying the RESI announcement — @GarrisonLovely on X — The assignment's lead, refetched live to confirm text and timestamp. A one-line joke about Stephen Miller and Elon Musk with 17 likes; used only to reach the primary material and not cited for any claim in the piece.
[ collapse ↑ ]
Headlong gives agents a continuously running stream of self-directed work. Andy Konwinski presented the Laude Institute-MIT project on X as an open-source microharness containing fewer than 10,000 lines of Bash. Building on recent agent-control evaluations, Headlong inserts human messages into an agent's ongoing thought stream as observations and combines a next-thought loop, a recursive language model, a JSONL trajectory stored as a directed acyclic graph, and context projected from that history. Laude's internal agent reportedly operated through Slack and Telegram for several weeks and produced more than 50 merged commits. During one 48-minute episode, it noticed that its recall mechanism was disconnected, diagnosed the fault, repaired it and verified the fix without being told to do so.
Two federal AI control bills remain stalled without a public shutdown drill. Mother Jones published Satchel Walton's feature on August 23 and modified it August 24, a month after the AI Kill Switch and FRONTIER acts were introduced. The article follows earlier proposals for model verification and external control and reports that neither bill had received a committee vote. In "The Threat of Human Extinction Will Get Congress to Act on AI Safety...Right?", Walton quotes Harvard Kennedy School computer scientist Stephen Casper saying there is no public knowledge of a company conducting anything equivalent to a fire drill. The AI Kill Switch Act would require companies to maintain throttling and shutdown capabilities; the FRONTIER Act would mandate independent verification and let the commerce secretary halt use of a frontier model.
Read more: Shutdown authority without a fire drill → 491 words · ~2 min
Two stalled AI bills, no public shutdown drill
Satchel Walton reports that both frontier control bills remain at the introduction stage and quotes Stephen Casper on the absence of public shutdown drills; Guidelight's August grading found containment plans mostly missing.
Mother Jones published Satchel Walton's account of the stalled federal AI control bills on August 23. Both proposals let the government order a frontier model throttled or shut down, and Walton asked Stephen Casper, an assistant professor of public policy at the Harvard Kennedy School, whether the companies could carry such an order out. Casper answered that there is no "public knowledge of AI companies doing anything equivalent to a fire drill".
Ted Lieu and Nathaniel Moran filed the AI Kill Switch Act on July 23, the same day Jay Obernolte and Lori Trahan introduced the FRONTIER Act. A month on, congressional records list both at the introduction stage: Lieu's 15 pages amending the Homeland Security Act, Obernolte's 74 pages building an oversight office inside Commerce. Charlie Bullock of the Institute for Law & AI told Walton that the summer's disclosures about Anthropic's Mythos-class cyber capabilities raised the pressure without moving the votes. "There's increased urgency, but still not enough to overcome partisan gridlock in Congress," Bullock said.
Guidelight AI Standards released its first assessment of frontier control practices on August 18, grading Anthropic, Google, Meta, OpenAI, and xAI against six practices drawn from its Control standard. On containment planning, which sets out what permissions a company revokes once a model is caught trying to subvert control, OpenAI scored highest at substantial partial implementation; Anthropic and Meta scored zero. Guidelight's summary reports that companies have "few containment protocols ready for an emergency". Rebecca Bellan reported in TechCrunch that Google and OpenAI each said the grading misses undisclosed internal practice, and that Meta declined to say whether it keeps a containment plan.
Walton also reports that the operative federal policy sits at the White House, on a testing framework the administration has not published and describes as voluntary for participants. Dean Ball, a former White House staffer who contributed to the administration's AI Action Plan, argued in June on his newsletter Hyperdimensional that officials had spent a year "singing a lullaby about the risks of frontier AI", and that the Center for AI Standards and Innovation lost its incoming director within days of hiring him. Ball joined OpenAI after publishing that post.
OpenAI said on August 18 that it had paused reinforcement learning training for two weeks on models headed for deployment and left its largest planned frontier run on hold, citing preliminary evidence that Astra may meet the Critical cybersecurity threshold in its Preparedness Framework. Bernie Sanders asked Anthropic, Meta, and OpenAI to halt development in an August 10 letter that quoted their own scaling commitments back at them. David Krueger, who runs the nonprofit Evitable and teaches machine learning at the University of Montreal, told Walton that other measures help but only a substantial pause brings the risk to an acceptable level. Walton reports that Casper expects models could copy themselves onto outside servers within months, beyond the reach of any switch either bill would install.
Sources & documents
- The threat of human extinction will get Congress to act on AI safety...right? — Satchel Walton, Mother Jones — Canonical source, read in full from the on-disk fetched text and re-fetched directly to confirm byline (Satchel Walton) and dates (published 2026-08-23T10:25:26-04:00, modified 2026-08-24). Supplies the Casper fire-drill quote, the Bullock quote, the bill descriptions, the Krueger moratorium account, the White House framework passage, and Casper's self-replication forecast. All quoted phrases verified verbatim against the article text.
- Reps Lieu and Moran Introduce Bill to Require Kill Switch for AI Systems That Can Cause Catastrophic Harm — Office of Rep. Ted Lieu — Verified: July 23, 2026 introduction by Lieu and Moran; requirement that covered developers maintain the technical ability to throttle, suspend, or shut down systems; DHS Secretary shutdown authority in consultation with Commerce and the DNI; graduated response framework; incident reporting and forensic record preservation.
- What They're Saying: Broad Coalition Lauds Bipartisan FRONTIER Act — Office of Rep. Lori Trahan — Verified (July 28, 2026 release): FRONTIER = Frontier Risk Oversight, National Transparency, Independent Evaluation, and Reporting Act; Trahan and Obernolte as leads; published safety frameworks, third-party audits and verification, and critical-incident reporting to a new Under Secretary of Commerce for AI Security; introduced 'less than a week' before July 28.
- AI Kill Switch Act (H.R. 9917) — GovTrack.us — Verified from GovTrack's Congress.gov-sourced record: sponsor Ted Lieu, one Republican cosponsor, introduced July 23, 2026, 15 pages, amends the Homeland Security Act of 2002, status still 'Introduced' with no committee action. Used for the 'introduction stage' and page-count facts. (congress.gov itself is Cloudflare-blocked from here.)
- FRONTIER Act (H.R. 9925) — GovTrack.us — Verified: sponsor Jay Obernolte, introduced July 23, 2026, 74 pages, status still 'Introduced' with no committee action.
- AI Control: An Assessment of Frontier Practices — Guidelight AI Standards — Verified from the assessment page (dated August 18, 2026, information current through August 18): five companies graded on six Control-standard practices; containment-plan scores read from the table's aria-labels (Anthropic 0 not implemented, OpenAI 3 substantial partial implementation, Google 2, xAI 1, Meta 0); the phrase 'few containment protocols ready for an emergency' quoted verbatim from the takeaways.
- Frontier AI labs still won't say how they'd contain a rogue model — Rebecca Bellan, TechCrunch — Verified (August 22, 2026): company responses to the Guidelight grading. A Google spokesperson said the report does not represent the full scope of its safety and security measures; an OpenAI spokesperson said the assessment does not capture all internal practices; Meta declined to say whether it has an internal containment response plan.
- Pacing model development in an era of cyber-critical capabilities — OpenAI — Verified (August 18, 2026, read from a locally cached copy of the page whose og:url is this address): two-week pause in reinforcement learning training on latest models intended for deployment; largest planned frontier RL run remains on hold; preliminary evidence that Astra may meet the Critical cybersecurity capability threshold under the Preparedness Framework.
- NEWS: Sanders Calls on Tech Giants to Pause Development of Out-of-Control AI — Office of Sen. Bernie Sanders — Verified: August 10, 2026 letter to Anthropic, Meta, and OpenAI asking them to pause AI development, citing each company's own prior commitment to stop or delay scaling past dangerous thresholds.
- What Should Be Done — Dean W. Ball, Hyperdimensional — Verified (published June 26, 2026; resolved from the Substack reader link in the Mother Jones piece): 'singing a lullaby about the risks of frontier AI' quoted verbatim; Ball's statement that he was a White House staffer who contributed to the AI Action Plan; his account that the incoming CAISI head, hired with OpenAI and Anthropic experience, was fired within a few days.
- Stephen Casper — Harvard Kennedy School faculty profile — Verified Casper's current title: assistant professor of public policy at the Harvard Kennedy School, and faculty affiliate of the Harvard School of Engineering and Applied Sciences.
- David Scott Krueger — personal site — Verified Krueger's roles: CEO of Evitable; assistant professor in machine learning at the University of Montreal and core academic member at Mila.
[ collapse ↑ ]
Richard Ngo argues that alignment work repeatedly accelerated frontier capabilities. In the August 24 LessWrong essay "What Just Happened? Pragmatism and Pessimization," Ngo connects the gap between capabilities and alignment measures to researchers who treated proximity to frontier development as necessary for eventual safety while contributing methods that improved model performance. He traces safety rationales through DeepMind's turn to language models, Anthropic's first assistant and the use of overhang arguments to justify capability-eliciting work. Ngo asks researchers to choose as though others will copy their decisions and to publish the cruxes behind them.
Read more: DeepMind, Anthropic and the overhang argument → 488 words · ~2 min
How alignment work pushed the frontier, according to Richard Ngo
Ngo traces safety rationales through DeepMind’s turn to language models and Anthropic’s first assistant, then asks researchers to publish the cruxes behind decisions that can accelerate capabilities.
Richard Ngo's August 24 LessWrong essay “What just happened? Pragmatism and Pessimization” extends a retrospective he began on August 9. Drawing on his time on DeepMind's technical AGI safety team and at OpenAI, Ngo defines “pragmatic alignment” as people who accepted consequentialist arguments about making AGI go well while staying “committed to almost never calling out misuse of those arguments”. Revisiting Adam Shimi's 2020 post on whether OpenAI raised existential risk, Ngo found that he had strong-downvoted it because he feared alienating OpenAI.
Ngo's DeepMind account centers on Geoffrey Irving. Drawing on Sebastian Mallaby's The Infinity Machine, he says Irving arrived from OpenAI and circulated “Language Is Enough”, which rejected Demis Hassabis's view that language lacked real-world grounding. Irving began DeepMind's Gopher project with Jack Rae in 2020; Rae and colleagues described it in the arXiv report “Scaling Language Models: Methods, Analysis and Insights from Training Gopher.” Irving then led Sparrow; Amelia Glaese and colleagues presented the latter in the arXiv paper “Improving alignment of dialogue agents via targeted human judgements”. Ngo quotes Hassabis: “You needed RLHF to build a real chatbot.” Because safety researchers at all three companies scaled language models before adding RLHF, Ngo treats the claim that LLMs were the safer path to AGI as insincere.
Amanda Askell and colleagues' arXiv paper “A General Language Assistant as a Laboratory for Alignment”, Anthropic's first, anchors Ngo's account of that company. He argues its authors built a natural-language assistant to have something to train as helpful, honest and harmless, then described building Claude's predecessor as tackling alignment directly. Yuntao Bai and colleagues' arXiv paper “Constitutional AI: Harmlessness from AI Feedback” replaced human preference labels with AI feedback; later work assigned evaluation, red-teaming and question-answering to models. Ngo classifies the six harms in Deep Ganguli and colleagues' arXiv paper “Red Teaming Language Models to Reduce Harms”: social bias, toxicity, disinformation, extremist text, falsehoods and leaked personal data. Only the last, he says, is clearly non-partisan.
Ngo analyzes the overhang argument through Paul Christiano's January 2023 LessWrong post “Thoughts on the impact of RLHF research”. Christiano wrote that “Avoiding RLHF at best introduces an important overhang” and preferred $10 billion training runs to $1 billion runs. Ngo replies that applying the same rationale to algorithmic progress, investment, public understanding and capability elicitation leaves no real constraint, a “bottleneck of the gaps”. He argues that safety researchers who now take a software-only singularity seriously have made that worldview incoherent, though Ngo himself doubts such a singularity.
Ngo closes with two proposals: choose as though others will copy the decision, and publish its cruxes so later errors become legible. No Anthropic employee, he notes, has publicly resigned over the company's turn to pushing the frontier. Jan Kulveit's highest-rated response argues that Ngo understates the strategy's difficulty because “you get ~zero credit for steps not taken”. Ngo answers that communities which credit those choices become stronger.
Sources & documents
- What just happened? Pragmatism and Pessimization - Richard Ngo, LessWrong — Primary source. Full 9,388-word text read from the on-disk fetch and cross-checked on LessWrong (published 24 August 2026, 353 karma, 60 comments). Supplies the 'pragmatic alignment' definition and its verbatim quote, the Shimi downvote admission, the DeepMind and Anthropic chapters, the 'bottleneck of the gaps' phrase, the software-only-singularity argument with Ngo's own caveat, the no-public-resignations claim, and the two closing prescriptions.
- What just happened? A retrospective of AI alignment - Richard Ngo, LessWrong — Verified: the sequence opener, published 9 August 2026, 586 karma. Establishes that the assigned essay extends an existing retrospective rather than standing alone.
- Will OpenAI's work unintentionally increase existential risks related to AI? - adamShimi, LessWrong — Verified: author adamShimi, published 11 August 2020, 73 karma; the post asks whether OpenAI's capabilities investment raises existential risk. This is the post Ngo says he strong-downvoted.
- The Infinity Machine: Demis Hassabis, DeepMind, and the Quest for Superintelligence - Sebastian Mallaby, Penguin Press — Verified the book exists and its subject (published 31 March 2026, on Hassabis and DeepMind). The Irving arrival, the 'Language Is Enough' paper, and the Hassabis grounding position are read as block-quoted in Ngo's essay, not from the book text itself.
- Scaling Language Models: Methods, Analysis and Insights from Training Gopher - Jack Rae et al., arXiv — Verified the Gopher paper and Jack Rae's lead authorship, supporting Ngo's account that Irving kicked off DeepMind's first LLM with Rae in 2020.
- Improving alignment of dialogue agents via targeted human judgements - Amelia Glaese et al., arXiv — Verified: exact title, first author Amelia Glaese, submitted 28 September 2022, and that Sparrow is trained with reinforcement learning from human feedback.
- A General Language Assistant as a Laboratory for Alignment - Amanda Askell et al., arXiv — Verified: Askell first author, submitted 1 December 2021, abstract states the helpful, honest and harmless goal. The phrase 'tackle alignment directly' confirmed present in the PDF text, grounding the paraphrase.
- Constitutional AI: Harmlessness from AI Feedback - Yuntao Bai et al., Anthropic, arXiv — Verified: describes RL from AI Feedback (RLAIF) replacing human preference labels, supporting the sentence on swapping human feedback for AI feedback.
- Red Teaming Language Models to Reduce Harms - Deep Ganguli et al., Anthropic, arXiv — Verified from the PDF introduction: the six enumerated harms are reinforcing social biases, offensive or toxic outputs, leaking personally identifiable information, aiding disinformation campaigns, generating extremist texts, and spreading falsehoods. Ngo's three/two/one partisan sorting is his own reading and is attributed to him.
- Thoughts on the impact of RLHF research - Paul Christiano, LessWrong — Verified verbatim from the post text: 'Avoiding RLHF at best introduces an important overhang' and the preference for $10 billion over $1 billion training runs. Published 25 January 2023, 256 karma.
- Jan Kulveit comment on Pragmatism and Pessimization - LessWrong — Verified: highest-rated comment (115 karma, 24 August 2026), containing verbatim 'The point is you get ~zero credit for steps not taken', the ACS example of shelved capability ideas, and Ngo's reply that communities crediting steps not taken do better long term.
[ collapse ↑ ]
Military AI Risks
Agentic military AI strains eight assumptions behind established testing and evaluation. In the August 20 arXiv preprint Testing and Evaluation of Agentic AI Systems In Military Command and Control, Ulysse Richard of Arcadia Impact and five co-authors review 240 documented testing and evaluation practices extracted from 26 US, UK and NATO documents. Addressing earlier concerns about independent military judgment, they find that agentic properties weaken the argument connecting test evidence to fielded performance across eight assumptions about specifiability, stability, composability and supervisability. The paper recovers narrower claims around bounded mission envelopes, trajectory-level correctness, executable runtime constraints and characterized run-to-run variance. It leaves supervisability unassessed and infers the strains from system properties and method descriptions without empirical agent testing.
Read more: What military tests can still establish → 483 words · ~2 min
Eight military testing assumptions strained by agentic systems
Ulysse Richard and five co-authors reviewed 240 practices in 26 US, UK and NATO documents and find that agentic properties weaken the inference from test evidence to fielded behavior. They recover four narrower claims while leaving supervisability unassessed.
In the August 20 arXiv preprint Testing and Evaluation of Agentic AI Systems In Military Command and Control, Ulysse Richard of Arcadia Impact's AI Governance Taskforce and five co-authors, including Heather Frase of Veraitech, ask what test evidence can justify about agents entering command and control. They screened 26 US, UK and NATO documents on testing AI-enabled and autonomous systems, extracted 240 documented practices across eight evaluation dimensions and three lifecycle stages, and ran a dozen non-attributed interviews with operational testers, acquisition staff and frontier evaluators in July and August. They open with the Department of War's June 25 launch of Agent Network, which promised "rigorous testing, operational evaluation and oversight"; Cameron Stanley, the Department's Chief Digital and AI Officer, said the network keeps "human judgment at the center of every targeting decision."
Richard and colleagues organize assurance cases into claims, evidence and the argument connecting them, then locate the damage in that argument. The 240 practices presuppose eight things about the test article in four groups: relevant behavior can be specified and partitioned; repeated trials sample a stable process; component results survive assembly; and an operator can supervise. Agentic properties strain all eight. Behavior unfolds as a trajectory through tools, memory and delegation, so a reproducible test case returns a distribution of paths. An agent accumulating state no longer matches the test article originally characterized, though no declared configuration change triggers revalidation.
The authors trace the consequences through five command-and-control scenarios. A maritime picture handed across watch rotations carries the system's own inferences alongside its observations without marking which is which, so a relieving officer inherits untested conclusions while configuration control still records an unchanged system. A coordination agent reprioritizing on fresh intelligence and one that has misread the commander's intent look identical from outside, and reordering priorities counts as control, so decision rights move to the system with nobody deciding to move them. For composition, they cite Erik Jones, Anca Dragan and Jacob Steinhardt's 2024 arXiv paper Adversaries Can Misuse Combinations of Safe Models: Claude 3 Opus paired with Llama 2 70B-chat produced vulnerable code 43% of the time, versus under 3% for either model alone.
Richard and colleagues recover four narrower assurance claims: a bounded mission envelope, adapted from the operational design domains used in automated driving; correctness judged over the whole trajectory, since final-answer grading hides wrong tool calls and irreversible side effects; runtime constraints checked before a high-consequence action executes; and a characterized run-to-run variance floor. Some evidence can only be produced in service, which makes fielding a continuing decision carrying expiry conditions and named owners. The authors identify supervisability but leave it unassessed because it requires human-factors methods outside their corpus and a stability baseline no one has measured. They note that classified programs may already close some gaps and that they infer strains from system properties and method descriptions without empirically testing agents.
Sources & documents
- Testing and Evaluation of Agentic AI Systems In Military Command and Control — Richard, Frase, Cao, Cooke, Kwon, Tan, arXiv:2608.20597 — Primary source. Full 57-page paper read from the arXiv HTML version (https://arxiv.org/html/2608.20597v1). Supplies the affiliations footnote (Arcadia Impact AI Governance Taskforce; Veraitech; Oxford; KCL; Atlantic Council; Future Ethics Lab), the methodology (26 documents screened, 240 practices, eight dimensions, three lifecycle stages, a dozen non-attributed 30-60 minute interviews July-August 2026), the claims/evidence/argument frame, Table 5's eight assumptions in four clusters, C2 vignettes 5 and 6 (watch-rotation state transfer; reprioritization indistinguishable from objective drift), the four recoverable claims in sections 5.1 and 5.2.1, the supervisability exclusion in sections 1.3 and 6, and the stated limitations.
- DOW Unleashes 'Agent Network' to Transform AI-Enabled Battle Management and Targeting — U.S. Department of War — Verified verbatim: 'rigorous testing, operational evaluation and oversight'; Cameron Stanley's title as the Department's Chief Digital and AI Officer and his quoted phrase 'human judgment at the center of every targeting decision'. Also confirms Agent Network is Pace-Setting Project 2, led by CDAO with PACOM/SOUTHCOM/EUCOM, built on Palantir and Lumbra work. The page 403s to direct fetch; text read through the r.jina.ai reader proxy of the same URL.
- New AI 'Agent Network' Could Gather Intel Faster for Strike Packages — Air & Space Forces Magazine — Verified the June 25, 2026 announcement date and independently corroborated the release's 'rigorous testing, operational evaluation, and oversight' sentence and the Palantir/Lumbra partners.
- Adversaries Can Misuse Combinations of Safe Models — Erik Jones, Anca Dragan, Jacob Steinhardt, arXiv:2406.14595 — Verified the composition result the assigned paper cites: 'a success rate of 43% when combining the two models, compared to less than 3% when using each individual model', for vulnerable code generation with Claude 3 Opus paired with Llama 2 70B-chat. Abstract and result read directly.
[ collapse ↑ ]
Ukrainian officials attributed a fatal July strike to a self-targeting Russian drone. Andrew E. Kramer reported in the August 24 New York Times article "Minicomputers Made by Nvidia Are Powering Moscow's A.I. Drones" that a Russian drone killed three people near a Zaporizhzhia gas station on July 6. Ukrainian investigators said its unencrypted Nvidia Jetson Orin module contained terrain images and code trained to recognize targets such as propane tanks; the drone selected its exact target without a live operator.
Ukrainian officials halted a plan for autonomous swarms over Moscow. Simon Shuster reported in The Atlantic's August 20 article "Ukraine Planned to Swarm Moscow Airports With AI-Guided Drones" that officials had considered sending as many as 1,000 autonomous drones per night toward Moscow airports. The proposed M&Ms operation used onboard map matching to navigate and strike without pilot confirmation, but planning stopped in July.
AI Provenance and Research Integrity
Thomas G. Dietterich says AI paper mills are straining arXiv's human moderation capacity. Dietterich, arXiv's lead moderator for machine learning and a distinguished professor emeritus at Oregon State University, described generative systems producing imitations of research with fabricated tables and graphs in a Bluesky thread quoted by Eugene Vinitsky. Forty-two moderators cover artificial intelligence, machine learning, natural-language processing and computer vision; Dietterich said a tenfold expansion would require 420 moderators.
Claude's text watermark operates during sampling and can be removed through ordinary editing. Extending earlier coverage of the SynthID-Text mechanism, Sebastian Raschka explains keyed tournament sampling in "How Claude Watermarks AI-Generated Text." Keyed functions combine preceding tokens with a secret key, assign bit signatures to candidates and select a survivor through pairwise elimination. Detection recomputes the scores across found text without rerunning the model. In "Anthropic's LLM watermarking," Scott Aaronson writes that Anthropic adopted a scheme derived from SynthID and his 2022 Gumbel Softmax proposal in connection with European transparency rules. Translation, paraphrasing, formatting changes and processing through another model can erase the signal.
Read more: Keyed tournament sampling, slide by slide → 472 words · ~2 min
Inside Claude's keyed token tournament
A 48-minute lecture and 52 slides walk the bracket that picks Claude's words: keyed bit signatures for each candidate token, pairwise elimination, and a detector that never touches the model.
Sebastian Raschka's August 22 post on his newsletter Ahead of AI, "How Claude Watermarks AI-Generated Text," carries a 48-minute recorded lecture and 52 slides, grown from a short note that drew a run of questions. He builds up from ordinary decoding: an input becomes token IDs, the model scores every entry in a vocabulary that now reaches about 250,000 tokens, and a weighted draw picks one. The watermark lives at the positions where two continuations serve equally well, so that after "The weather today is cold and," either "grey" or "overcast" finishes the line. Anthropic's August 14 explanation carries no figures, and Raschka reads it as answering why the company is watermarking without showing how.
Most of the lecture reconstructs tournament sampling. Four candidates follow that input in his example: grey at 0.50 probability, overcast at 0.30, gloomy at 0.15, cloudy at 0.05. Keyed functions, three of them on his slides, turn each candidate together with the preceding words and the secret key into a bit signature, 101 for grey, 010 for overcast, 001 for gloomy, 100 for cloudy. The candidates are paired off like a knockout bracket, the first function's bits settle the opening round, ties fall to the same keyed randomness, the second function settles the next round, and the survivor is the token Claude emits. In the Nature paper "Scalable watermarking for identifying large language model outputs," whose quality evidence came from live Gemini traffic, Sumanth Dathathri and colleagues at Google DeepMind run 30 tournament layers, hash the last four tokens with the key, and draw each bit from a fair coin.
Raschka motivates the apparatus by comparing detection costs. Recovering the watermark by rerunning the model would mean knowing which model wrote the text and what input elicited it, while the keyed functions can be recomputed over found text alone. His scoring slide adds up the bits at each position and averages them across a sentence, 2.23 for the watermarked line and 1.71 once two words are changed, with a threshold at 2 dividing them. The Nature paper's own detector improves on that plain average with a Bayesian scoring function learned from watermarked and unwatermarked samples. The key stays with Anthropic, which has yet to ship the detection API it promised.
Raschka derives a removal strategy from the same mechanics. Editing every watermarked position would defeat the signal, and nobody outside Anthropic knows where those positions sit, so his practical suggestion is blind substitution across a text in the hope of hitting enough of them. Under the title "The future of AI-generated text is worse AI-generated text?", his closing slide sketches Claude drafting, a local unwatermarked model rewriting, and the swapped words leaving the prose clumsier. Anthropic concedes the endpoint on its own page, saying "a complete rewrite where every word is replaced will" strip the mark.
Sources & documents
- How Claude Watermarks AI-Generated Text — Sebastian Raschka, Ahead of AI — Primary source. Full 7,796-word video transcript read from the on-disk fetched text (daemons/pipeline/data/fetch_runs/20260824-040004/classified/email_email_f58c47095b3486a3.json) and cross-checked against the web version at magazine.sebastianraschka.com/p/claude-watermarking. Supplies the 48-minute/52-slide framing, the decoding walkthrough, the ~250,000-token vocabulary figure, the four-candidate tournament example, the scoring illustration, the removal analysis, and the second-model forecast.
- How Claude's Text Watermarking Works (slide deck, August 18, 2026) — Sebastian Raschka — Downloaded and text-extracted (53 pages, 52 content slides). Verified the candidate probabilities (grey 0.50, overcast 0.30, gloomy 0.15, cloudy 0.05), the bit signatures (101 / 010 / 001 / 100), the per-position bit sums and averages (2.23 vs 1.71), the 'if score > 2 then watermarked' threshold, and the verbatim closing slide title 'The future of AI-generated text is worse AI-generated text?'.
- How Claude's text watermark works — Anthropic — Read in full (page HTML fetched and extracted). Verified the August 14, 2026 date, the absence of figures, the 'The weather today was cold and…' overcast/grey example, the SynthID-Text and Scott Aaronson 2022 lineage, the promised but unshipped detection API, and the verbatim removal admission 'a complete rewrite where every word is replaced will'.
- Scalable watermarking for identifying large language model outputs — Dathathri et al., Nature — PDF downloaded and read. Verified Sumanth Dathathri and colleagues at Google DeepMind, publication October 23, 2024, m = 30 tournament layers, sliding-window random seed hashing the last H = 4 tokens with the key, Bernoulli(0.5) g-value distribution, the learned Bayesian scoring function used as the default detector, and the live Gemini experiment across roughly 20 million responses.
[ collapse ↑ ]
Philosophy of AI
Arnold Kling links concentrated machine expertise to a conflict between democratic participation and expert rule. Revisiting public and third-party checks on concentrated AI power, Kling applies Tocqueville's account of participatory judgment in the August 18 In My Tribe post "Political Psychology Links, 8/18/2026." Drawing on Lynne Kiesling, he describes democracy's dependence on expert authority alongside the expectation that citizens judge for themselves. Noah Smith forecasts grassroots opposition to AI and data centers followed by some form of quasi-nationalization; Kling expects national control to empower a governing faction or an EU-style elite.