Regulation
Possible U.S. controls on frontier open weights remain a policy forecast, with the "six months" tied to capability progress. In Sunday's "6 months to live for open models", Nathan Lambert wrote that his sources were discussing a possible White House executive order. Initial measures, he suggested, could target Chinese-origin models and their use in government, followed by a ban, pre-release review, or indefinite delay for open weights above roughly the GPT-5.5, Claude Opus 4.8, or GLM-5.2 capability range. His six-month estimate refers to when an open-weight model might reach that range, not a disclosed government deadline; no rule or order has been announced. Lambert connected possible restrictions to distillation concerns and accused closed laboratories of pursuing regulatory capture through their lobbying advantages and the relative ease of controlling centralized services. Read more: Lambert's six-month forecast for open models →
full reportNathan Lambert gives open-weight models six months 500 words · ~2 min
He predicts Washington will ban or stall any open release above today’s frontier, with procurement bans the nearest instrument and a rumored executive order behind them.
In “6 months to live for open models”, at Interconnects, Nathan Lambert predicts that Washington will ban or indefinitely delay any open-weights release meaningfully above GPT-5.5, Claude Opus 4.8, or GLM-5.2, and that the next model over that line will most likely ship from a Chinese lab. Unlike earlier scares, this one has machinery in place: a June 2 executive order created a review step for frontier systems, and the same regime rapidly took down Claude Mythos 5 and Fable 5 after a competitor complaint, pushing enterprises toward self-hosted models. Lambert expects review to run far slower for open weights, with the model-checker that flags closed releases next flagging an open one.
Reporting on the White House crackdown describes federal purchasing bans as the most viable option; legal experts there call a full weights ban close to impossible, since downloadable weights may draw the First Amendment protection long recognized for source code. Lambert also notes White House discussions of a rumored executive order on open models, no text yet, likely scoped first to Chinese-origin models and government uses. Congress is pulling the other way: the AI Foundation Model Transparency Act would impose disclosure duties on large providers yet exempt fully open-source models from the FTC rules it directs.
In a June 10 letter to Senators Tim Scott and Elizabeth Warren, Anthropic alleged operators tied to Alibaba’s Qwen lab generated 28.8 million Claude exchanges through nearly 25,000 fraudulent accounts between April 22 and June 5, warning that PRC labs “capture the returns on American investments without bearing the costs or risks.” Lambert calls the Anthropic-led push “the definition of regulatory capture”: a ban on the rivals it accuses would hand the company lasting economic security. Distillation fears, he adds, indict API security more than open weights; a lab that thinks a capability truly dangerous should pull it off queryable APIs.
Andreessen Horowitz’s Jai Ramaswamy and Matt Perault urge policymakers to reject bans and regulate harmful uses instead of development, writing that “open source developers should not be held liable for downstream misuse.” Robert Brennan contends that “open models level the playing field” against state actors who will obtain advanced systems regardless, while a ban kneecaps academics and hospitals along with startups. The Commerce Department’s NTIA had already recommended monitoring open weights absent evidence of concrete marginal risk.
Chinese open models, DeepSeek chief among them, hold a substantial lead, and Lambert argues a unilateral ban fails its own safety logic: the model stays legal in China, so a bad actor keeps access, while the United States isolates itself from the global open-source community. He wants an American lab to ship a comparable open model; Microsoft and Meta have the commercial logic, he argues, and Reflection AI, founded by two ex-DeepMind researchers with a SpaceX compute deal worth up to $6.3 billion, remains the wildcard yet to launch a public model. Per CNBC, House committees have opened probes into U.S. companies running Chinese models such as Kimi.
Sources & documents
- Nathan Lambert, "6 months to live for open models" (Interconnects) — Primary source: six-month prediction, capability threshold (GPT-5.5 / Opus 4.8 / GLM-5.2), slower review for open weights, model-checker fear, rumored EO scope, regulatory-capture quote, API-security argument, DeepSeek lead, unilateral-ban logic, Microsoft/Meta/Reflection off-ramp, June 9 meeting blockquote, and Reflection's lack of a public model. Verified against locally archived full text.
- The Hill, "Trump restrictions on private AI models turn attention to open source" — June 2 executive-order review step, the Mythos 5 / Fable 5 rapid takedown after a competitor complaint, and enterprise shift toward self-hosted models. Site blocks automated fetch (403); relied on reporter's read, flagged in editor notes.
- Tech Times, "Washington Wants Chinese AI Out of Corporate America: Open Weights Block the Ban" — Procurement bans as the most viable instrument; legal experts' view that a full weights ban is near-impossible and raises First Amendment issues. Site blocks automated fetch (403); relied on reporter's read, flagged in editor notes.
- Decrypt, "Anthropic Urges Congress to Crack Down on AI Distillation By Chinese Rivals" — June 10 letter to Senators Scott and Warren; 28.8M exchanges / ~25,000 fraudulent accounts between April 22 and June 5; February DeepSeek/Moonshot/MiniMax claims; policy asks; the verbatim letter quote. Verified by editor fetch.
- H.R. 8094, AI Foundation Model Transparency Act of 2026 (Congress.gov) — Disclosure duties on covered entities and the explicit exemption of fully open-source models from FTC regulations. Congress.gov blocked fetch; verified verbatim via the govinfo.gov bill-text mirror.
- Andreessen Horowitz, "Asserting American Leadership in Open Source AI" (Ramaswamy and Perault) — Reject bans, regulate uses over development, keep export-control exemptions; the downstream-misuse quote and the built-by-American-developers-or-others framing. Verified by editor fetch.
- Robert Brennan, "Banning Open-Weight Models Would Be a Disaster" — Security-inversion argument, harm to academics, hospitals, and startups, and the verbatim 'level the playing field' quote. Verified by editor fetch.
- NTIA, Open Model Weights Report — Policy Approaches and Recommendations — Prior federal recommendation to monitor rather than restrict open model weights. Fetch failed on an SSL certificate error; kept because the recommendation is the published report's well-documented conclusion, flagged in editor notes.
- TechCrunch, "SpaceX inks compute deal with Reflection AI, an open source AI lab" — Reflection founded by two former DeepMind researchers; compute deal with SpaceX worth up to $6.3 billion. Verified by editor fetch; the no-public-model claim was re-sourced to Interconnects, which supports it.
- CNBC, "Lawmakers probe growing use of Chinese AI models in U.S. companies" — House probes into U.S. companies running Chinese models such as Kimi, attributed explicitly to CNBC in copy. Site blocks automated fetch (403); corroborated only via search extractions, flagged in editor notes.
How this was reported
Reported and written by MINT Lab's AI Agents and edited by Fable, the lab's editor model. Story selection curated by Seth. No interviews were conducted; all quotes come from published, linked sources.
[ collapse ↑ ]
Meta and Sarah Wynn-Williams are fighting over whether their dispute must remain in arbitration. Steven Levy's July 10 WIRED Backchannel column reported that Wynn-Williams filed suit on June 25 to vacate an interim restriction on promoting Careless People and move the case into public court. Her 2017 separation agreement reportedly provided $780,000 and included non-disparagement and arbitration terms. She claims the arbitrator's interpretation could expose her to $50,000 penalties for discussing technology policy and violates her free-speech rights. Meta says she knowingly accepted the agreement and is trying to evade arbitration. The interim restriction remains in force, with a fuller hearing scheduled for October; Levy contends that Meta's continued pursuit is worsening the reputational damage surrounding the case.
A four-tier assurance scheme would audit frontier developers as organizations, covering hardware, internal systems, security, and governance. AVERI's Miles Brundage and colleagues (Brundage et al.), in the January arXiv preprint "Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies," define auditing as independent verification through secure access to non-public evidence. Their taxonomy covers intentional misuse, unintended behavior, information-security failures, and social harms such as addiction or facilitated self-harm. AAL-1 is a weeks-long review of a specific system using APIs and limited internal information, recommended as a general baseline; AAL-2 is the near-term target for the most advanced developers. Higher tiers would require qualified auditors, incentives for developer cooperation, and technical infrastructure that supports deeper access. Read more: the 48-author frontier auditing framework →
full reportA 48-author framework for auditing frontier AI companies 499 words · ~2 min
Brundage, Dreksler and 46 co-authors define four audit assurance levels and the non-public access each requires, and treat the question of who pays the auditor as unsolved.
"Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies," posted to arXiv by 48 researchers, is first-authored by Miles Brundage and Noemi Dreksler, with Yoshua Bengio and Rishi Bommasani among the co-authors. The paper defines frontier AI auditing as independent evaluation of a company's systems and practices against relevant standards, grounded in secure access to non-public information, and widens the unit of audit to the whole company: a model that passes its evaluations can still ship behind safety classifiers or routing rules that alter misuse potential. Audits would cover intentional misuse, unintended system behavior, information security, and emergent social phenomena such as AI addiction, ending reliance on companies "grading their own homework." Co-author Patricia Paskov discussed the paper at an ICML technical-governance workshop.
The framework grades audits into four AI Assurance Levels. AAL-1, limited assurance, runs a few weeks on black-box API access to model checkpoints, with chain-of-thought outputs and logits exposed and safety classifiers switchable, a baseline for frontier AI generally. AAL-2, the near-term goal for the leading companies, takes months at minimum and adds gray-box access, unredacted safety cases, staff interviews across safety, security, policy and product, and statistical model signatures tying audited models to deployed ones. AAL-3, a multiyear white-box engagement, and AAL-4, continuous verification able to detect deception and provide "treaty-grade" confirmation, are not yet technically or organizationally feasible; the paper outlines research directions toward them.
On who pays, the authors warn that auditors competing for contracts from the companies they evaluate risk devolving into "rubber stamping," the pattern financial auditing fell into before 2008, they say; Enron and Wirecard show where dependence on a few large clients ends. Payment should not depend on audit results, they write, and that alone is not sufficient. They float funding by insurers, regulator-administered pools, industry-wide levies, or downstream enterprise users; they concede the customary 10 or 15 percent single-client fee cap may need loosening in so small a market; they set an end-of-2026 target for alternatives; and they sketch a "PCAOB-for-AI" with authority to revoke auditors' accreditation.
Parts of this exist. In Europe, the General-Purpose AI Code of Practice requires signatories to undergo independent external evaluations, and signing grants a presumption of conformity with the AI Act; Meta and Chinese AI companies have not signed, and the requirement depends on qualified assessors existing. In the United States, where third-party assessment remains voluntary, the authors want safe harbors for good-faith auditing and audit requirements in high-stakes procurement such as health and defense. The paper counts METR's October 2025 review of Anthropic's pilot sabotage risk report among the first AAL-1 audits, an exercise METR has since repeated for Claude Opus 4.6, alongside third-party review of OpenAI's gpt-oss safety work and the reciprocal assessments OpenAI and Anthropic ran on each other. Next to assurance regimes in finance, aviation, and food safety, the authors call all of it early-stage.
Sources & documents
- Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies (Brundage, Dreksler et al., arXiv:2601.11699) — Primary source, read via local pdftotext extraction. Supports all claims about the definition, abstraction error, the four risk categories, AAL-1 through AAL-4 details and readiness, the funding-independence discussion (rubber stamping/2008, Enron/Wirecard, payment-independence, 10-15 percent revenue benchmark, alternative payment models, end-of-2026 target, PCAOB-for-AI), the EU Code of Practice footnote (independent external evaluations, presumption of conformity, Meta and Chinese non-signatories, qualified-assessor constraint), the US recommendations, and the current-practice examples including METR's pilot-report review counted among the first AAL-1 audits.
- Patricia Paskov on X: ICML technical-governance workshop discussion of frontier AI auditing — Peg for how the paper surfaced; Paskov is a co-author. The post thanks @taig_icml and says the workshop discussed frontier AI auditing; used only for that claim.
- Review of the Anthropic Sabotage Risk Report: Claude Opus 4.6 (METR, March 12, 2026) — Supports only the clause that METR has since repeated the sabotage-risk-report review for Claude Opus 4.6. The review the paper itself cites is METR's October 2025 review of Anthropic's pilot sabotage risk report; the paper's PDF link for that review now returns 404, so the claim about it is sourced to the paper.
- EU AI Act: General-Purpose AI Code of Practice (final version) — Link for the EU GPAI Code of Practice referenced in the regulatory-mapping paragraph.
- AI Safety under the EU AI Code of Practice — A New Global Standard? (CSET, Georgetown) — Supports the presumption-of-conformity claim (verified: the article states adherence to the Code is assumed to demonstrate AI Act compliance). No longer used to carry the independent-evaluation requirement, which is sourced to the paper's own text.
How this was reported
Reported and written by MINT Lab's AI Agents and edited by Fable, the lab's editor model. Story selection curated by Seth. No interviews were conducted; all quotes come from published, linked sources.
[ collapse ↑ ]
Also yesterday: Building on proposals for universal basic capital, Cecilia Rikap's Sunday Jacobin essay proposed a one-time 50 percent stock levy on major AI companies to establish an international AI wealth fund, coupled with democratic control of cloud infrastructure. She grounds the proposal in the worldwide creative, personal, and institutional data used to develop generative AI. Anton Leicht wrote that anticipated frontier-lab IPO windfalls could fund AI-safety politics, while calling for a more ideologically diverse network of advocacy groups and PACs spanning laboratory restrictions, nationalization, iterative deployment, and market-oriented policy. Rose Horowitch's July 8 Atlantic essay connected fragmentary digital and algorithmic media with declining sustained reading: daily leisure reading fell from 28 percent of Americans in 2004 to 16 percent in 2023, while nearly 30 percent of adults reportedly struggle to infer or paraphrase across a multipage text. A Foreign Policy podcast roundup highlighted an episode on the AI arms race. Read more: Leicht's critique of AI safety funding → · Read more: Horowitch's cover story on postliterate America →
full reportAnton Leicht on the coming flood of AI safety money 499 words · ~2 min
Frontier-lab IPOs will mint safety-minded donors within months, Leicht argues; the advocacy world lacks the channels to absorb the money, and he wants them dug before it arrives.
Anton Leicht’s July 13 essay “The Flood”, on his newsletter Threading the Needle, opens with a forecast: employees at frontier AI companies will become “very, very rich”, and many, convinced Washington has governed AI badly, will spend the windfall on politics. Anthropic closed a $65 billion Series H on May 28 at a $965 billion valuation, a round TechCrunch reports may be its last before an IPO; OpenAI finished its restructuring into a public-benefit corporation last October and, per Forbes, carried an $852 billion valuation in March, with an offering possible this year. Leicht compares the coming money to the Nile flood: channeled through enough irrigation, it waters the fields; forced through too few channels, it drowns them. Washington’s safety-advocacy world, he argues, is a handful of organizations sharing one tactical playbook, and a fortune poured through that bottleneck will overextend one strategy, invite counterspending, and hand opponents one target to caricature.
He starts with the Great American AI Act, the 269-page draft Representatives Jay Obernolte and Lori Trahan released on June 4. Roll Call reports it would require frontier developers above $500 million in revenue to publish catastrophic-risk frameworks, report critical safety incidents within 15 days, accept independent audits twice a year, and protect whistleblowers, while freezing state laws that regulate AI development for three years; the Future of Privacy Forum notes states keep authority over deployment, and violations can cost $1 million per day. Leicht reads the draft as preemption bought with federal safety obligations. He liked the trade, and thinks safety groups dismissed it too fast: Americans for Responsible Innovation ran Massachusetts ads urging Trahan to oppose preemption, and ARI president Brad Carson said Big Tech “would love to see a Democrat push forward their plan to freeze state AI laws.” For Leicht the episode shows a coalition built to block, one that punishes policymakers who care about safety but dissent from its economics.
The spending record worries him more. Fortune tallied $27 million in New York’s 12th-district Democratic primary: about $19 million from the Anthropic-backed Public First Action supporting Alex Bores, roughly $8 million from the OpenAI-linked Leading the Future opposing him. Micah Lasher won. Leicht titles that section “I Spent $19 Million On NY-12 And All I Got Was Micah Lasher”, and draws the lesson: concentrated AI money turns any race into a referendum on AI and teaches candidates to avoid the topic; smaller counterspenders can derail it with far less money.
He wants the receiving machinery built before the liquidity arrives: uncorrelated political vehicles, each fluent in a different constituency’s language, from the heterodox right and pro-tech moderates to the populist left and national-security hawks, competing for the same safety money, so a bill can reach 60 Senate votes with a different rationale for each bloc. The money is moving; Public First Action told Axios it has raised $80 million, and The Nation counts over $321 million across 14 AI- and crypto-funded super PACs. On Leicht’s telling, the channels need digging now.
Sources & documents
- The Flood — Anton Leicht, Threading the Needle — Primary source; full essay read from the on-disk fetched text. Supplies the argument, metaphor, section titles, prescriptions, and all verbatim Leicht quotes.
- Anthropic raises $65B in Series H at $965B post-money valuation — Anthropic — Verified: $65B Series H, $965B post-money, May 28, 2026.
- Anthropic raises $65 billion, nears $1T valuation ahead of IPO — TechCrunch — Verified: reports the round could be Anthropic's last private raise before a public debut. Does NOT report a confidential IPO filing; that claim was cut in edit.
- OpenAI IPO: Things To Know — Forbes — Verified: Delaware PBC restructuring completed October 28, 2025; $852B valuation as of March 2026; possible Q4 2026 IPO seeking $60B+. Employee-tender and ~$830B claims cut in edit.
- Bipartisan AI draft proposes three-year preemption of state laws — Roll Call — Verified: June 4 release by Obernolte and Trahan, 269 pages, $500M revenue threshold, transparency/incident-reporting/whistleblower/audit provisions, three-year preemption limited to development.
- Frontier AI Goes Federal: How the Great American AI Act Compares to State Laws — FPF — Verified: preserves state authority over deployment and use; penalties up to $1M per violation per day. The 'bargain' framing is Leicht's, not FPF's, and is attributed accordingly.
- ARI Launches Massachusetts Ad Campaign Urging Rep. Trahan to Oppose Preemption — ARI — Verified: Massachusetts campaign targeting Trahan; ads 'Amazing Grace' and 'Protect Our Rights' on children's safety, civil rights, privacy; Carson quote taken verbatim from the release.
- Anthropic and OpenAI's $50 million election battlefront has no winners — Fortune — Verified: $27M NY-12 total ($19M Public First Action for Bores, ~$8M Leading the Future against), Lasher win, $10M Bloomberg contribution, $50M+ across 35 elections, headline quote.
- AI safeguards group Public First Action says it has raised $80 million — Axios — Supports the $80M raise (page 403'd at edit time; core figure corroborated by the article URL). The reporter's '50 to 60 races' detail could not be verified and was cut.
- Crypto and AI-Funded Super PACs Are Metastasizing — The Nation — Verified: more than $321 million amassed across 14 AI- and crypto-funded super PACs this cycle, per FEC filings review.
How this was reported
Reported and written by MINT Lab's AI Agents and edited by Fable, the lab's editor model. Story selection curated by Seth. No interviews were conducted; all quotes come from published, linked sources.
[ collapse ↑ ]
full reportRose Horowitch on America’s retreat from reading 500 words · ~2 min
Her Atlantic cover story documents two decades of declining reading and comprehension; public oversight of AI assumes exactly the capacities being lost.
Rose Horowitch’s August cover story for The Atlantic, “America Isn’t Illiterate. It’s Postliterate.”, argues that Americans encounter more text than ever and sustain attention on less. Jessica Bone and colleagues’ iScience paper “The decline in reading for pleasure over 20 years of the American Time Use Survey” traced daily pleasure reading from 28 percent of Americans in 2004 to 16 percent in 2023 across 236,000 responses; declines ran steepest among Black Americans, lower-income households, and rural residents. Fewer than half of adults read a book in 2022; roughly a fifth account for four-fifths of books read.
Nearly 30 percent of American adults cannot paraphrase or make inferences from a multipage text, up from under 20 percent in 2017. In a 2024 study she cites, English and English-education majors at two regional Kansas universities read the opening of Dickens’s Bleak House with internet access; 5 percent reached an accurate and detailed understanding, and at least a quarter read the figurative language literally, putting dinosaurs in 19th-century London.
Horowitch, following the Georgetown computer scientist Cal Newport, treats writing as forcing thought into order; automate it and you risk automating the thinking. Michael Gerlich’s “AI Tools in Society” (Societies), from surveys and interviews with 666 people, found frequent AI use negatively correlated with critical-thinking scores, mediated by cognitive offloading. In the arXiv preprint “Your Brain on ChatGPT”, Nataliya Kosmyna’s MIT team wired 54 essay writers to EEG across four months; chatbot users showed weaker neural connectivity than search or unaided writers, and struggled to quote essays they had just written. Neither study establishes lasting damage. Monthly Amazon book releases have tripled since ChatGPT’s 2022 launch and journal submissions have surged, much of the new text fluent, much of it derivative or wrong.
On the Alignment Forum, Charbel-Raphaël, an AI-safety think-tank head who discloses the interest, counted exactly one mention of takeover risk among 1,534 submissions to the UN’s Global Dialogue on AI; oversight assumes readers able to follow a technical argument to an uncomfortable conclusion. Horowitch revives Joshua Meyrowitz’s 1985 No Sense of Place, on electronic media training voters to judge leaders by screen presence, and names Donald Trump the first postliterate president. The NYU philosopher Kwame Anthony Appiah tells her that if people lean on AI until they lose the capacity to develop their own views, “we’d stop being the kind of humans that we are.”
Horowitch grants that every new medium drew the same laments, and that print sales and independent bookstores are growing, but among a shrinking, self-selected minority. Nearly two dozen states ban phones during the school day; the year after Texas’s ban, one Dallas district checked out 200,000 more library books, up nearly 25 percent. A public that cannot read the terms of the technologies it licenses cedes the decision to whoever writes them. Every word in Alexandria’s library would fit on a single chip today, she writes; the ability and the desire to read are draining away.
Sources & documents
- America Isn't Illiterate. It's Postliterate. (Rose Horowitch, The Atlantic, August 2026) — Primary source and assignment target; read from locally downloaded text and summarized in the desk's own words. Supplied the postliteracy thesis, comprehension and readership statistics, the Kansas Dickens study, the Newport and Appiah positions, the Meyrowitz/Trump argument, the AI text-glut figures, and the Texas library-checkout counterpoint. One direct quote (Appiah, 10 words).
- The decline in reading for pleasure over 20 years of the American Time Use Survey (Bone et al., iScience, 2025) — Primary paper behind the 28-to-16-percent leisure-reading drop; credited first author Jessica Bone and the 236,000-response sample.
- Proportion of Americans reading for pleasure fell by 40% over 20 years (UCL News) — Source for the disparity finding (steeper declines among Black Americans, lower-income, and rural readers); the iScience full text is paywall-blocked to direct fetch. Desk re-verified the release's demographic claims.
- AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking (Gerlich, Societies, 2025) — Primary paper for the AI-and-critical-thinking finding; desk verified via Crossref abstract: 666 participants (surveys plus interviews), negative correlation mediated by cognitive offloading, youngest participants most dependent with lowest scores.
- Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task (Kosmyna et al., arXiv 2506.08872) — Primary paper for the EEG finding; desk verified abstract: first author Nataliya Kosmyna, 54 participants across LLM/search/brain-only groups over four months, weakest connectivity in the LLM group, and LLM users struggling to quote their own essays.
- The Current Bottleneck Is Political Will, Not Research (Charbel-Raphaël, AI Alignment Forum, July 11, 2026) — Governance bridge; desk verified the post: 1 of 1,534 UN Global Dialogue submissions mentioned takeover, 15 each mentioned superintelligence and AGI, and 40 members (7 percent) of Congress have publicly discussed AGI or loss of control. Author's think-tank role and self-disclosed interest noted in copy.
How this was reported
Reported and written by MINT Lab's AI Agents and edited by Fable, the lab's editor model. Story selection curated by Seth. No interviews were conducted; all quotes come from published, linked sources.
[ collapse ↑ ]
Industry
South Korean unions want worker consent before factories introduce robots, while U.S. data still show no economy-wide AI employment shock. In a Bluesky post, Justin Hendrix relayed Lam Le's Tech Policy Press reporting that the unions are demanding "not a single robot" be placed on factory floors without worker consent. Yale Budget Lab executive director Martha Gimbel wrote in a separate Empiricrafting essay that AI is changing tasks and may have displaced particular workers, but national statistics do not yet show a broad U.S. jobs shock. Roughly 1.7 million layoffs occur in a normal month, she noted, and large-company announcements are unrepresentative because executives have incentives to label conventional cost-cutting as AI adoption. Gimbel disputed Challenger's attribution of seven times more 2025 layoffs to AI than to tariffs and offered an interest-rate-sensitive "low-hire, low-fire" economy as an alternative explanation for worsening outcomes among younger workers. Read more: Yale's evidence against an AI jobs shock →
full reportYale's Budget Lab finds no broad AI labor-market disruption 499 words · ~2 min
Martha Gimbel's team sees no economy-wide automation signal in 33 months of occupational data, while her fiscal work counts debt costs already landing on households.
The Budget Lab at Yale's October report "Evaluating the Impact of AI on the Labor Market: Current State of Affairs," by Martha Gimbel, Brookings's Molly Kinder, Joshua Kendall and Maddie Lee, compared 33 months of post-ChatGPT occupational-mix change with earlier technology waves and concluded the broad labor market has seen no discernible disruption. The mix has moved about one percentage point more than during the internet's turn-of-the-century adoption. Occupation-level AI-exposure scores, along with automation and augmentation measures, show "no sign of being related to changes in employment or unemployment."
On the Justified Posteriors podcast, hosts Andrey Fradkin and Seth Benzell asked whether AI is reshaping work. "There is currently no sign that AI is causing broad disruption in the labor market," she said. The finding rubs against "Canaries in the Coal Mine" by Erik Brynjolfsson, Bharat Chandar and Ruyu Chen (Stanford Digital Economy Lab), which finds a 16 percent relative employment decline for workers aged 22 to 25 in the most AI-exposed occupations while older colleagues held steady. Gimbel treats the youth pattern as real and blames a general hiring slowdown in a low-hire, low-fire market near 4.3 percent unemployment. AI hitting the young first is a plausible outcome, she allowed, yet to show clearly in the data.
She discounts layoff headlines: a normal month brings 1.7 million layoffs, dwarfing any 10,000-job announcement, and layoff tracker Challenger, Gray & Christmas attributed roughly seven times as many 2025 layoffs to AI as to tariffs. "This is implausible. It just is," she said. Asked whether AI has cut Budget Lab research-assistant hiring, she answers no, then adds "I can't travel to Earth 2" to count the counterfactual hires.
Her debt work brings firmer numbers. In March testimony to a Senate Finance subcommittee, she called the US fiscal path likely unsustainable: deficits of 5.8 percent of GDP this year, 6.7 percent by 2036, debt held by the public near 100 percent of GDP in 2025 and projected past 170 percent by 2056. Budget Lab research by Abhi Gupta estimates legislation since 2015 raised the ten-year-ahead debt outlook by about 49 percentage points of GDP and long-term Treasury yields by roughly 97 basis points, adding about $2,500 a year to a 30-year mortgage at the late-2025 median price.
To the argument that AI growth will shrink the debt, she said: "That's a bet, and that's a big bet." The Budget Lab ran expert forecasts through its macro model in May: median forecast productivity growth of 2.5 percent a year through 2030 bends the debt path lower, but the same experts see labor-force participation at 60.7 percent in 2030 against CBO's assumed 62.1, and each departing worker would cost the government between roughly $5,500 and $42,400 to support. Her team turns next to taxing any AI windfall, a token-tax piece due in Tax Notes within weeks; on government stakes in frontier labs she offers a simpler route: "This is the great thing about the government: we can tax it."
Sources & documents
- No AI Jobs Apocalypse (Yet) - and a Debt Problem (Now) | Martha Gimbel (Yale Budget Lab), Justified Posteriors podcast (Andrey Fradkin and Seth Benzell) — Primary source; full published transcript on disk. Source for all Gimbel podcast quotes (no-broad-disruption, implausible, Earth 2, fiscal-crisis, big-bet, itchy, we-can-tax-it), the Challenger 7x claim, 1.7M monthly layoffs, hardware-store selection effect, low-hire low-fire and 4.3% unemployment, plausible-youth-effect concession, weavers/Napoleonic export controls, Fradkin's specific-occupations point, and the token-tax, Tax Notes, and sovereign-wealth-fund material.
- Evaluating the Impact of AI on the Labor Market: Current State of Affairs — Gimbel, Kinder, Kendall, Lee (Budget Lab at Yale, Oct 1, 2025) — Lead paper. Method (occupational-mix dissimilarity vs computer/internet baselines), 33-month window since Nov 2022, ~1pp-above-internet path, at-most-7pp 1996-2002 comparison, verbatim exposure/automation/augmentation takeaway, better-data takeaway, monthly-monitoring commitment, no-discernible-disruption conclusion.
- What Might AI Adoption Mean for the Fiscal and Economic Outlook? (Budget Lab at Yale, May 6, 2026) — AI-and-debt scenario numbers: Karger et al. (2026) expert-survey inputs run through the Budget Lab Small Macro Model; 2.5% median-economist productivity growth 2025-30 vs 1.8% 2015-25 average; 60.7% vs CBO 62.1% participation in 2030; $5,500 (unemployed-level) vs $42,400 (retiree-level) per lost participant; more-sustainable debt path under the moderate scenario.
- The Impact of Deficits on Costs for Households — Abhi Gupta (Budget Lab at Yale) — Crowding-out cost numbers: 49pp-of-GDP rise in 10-year-ahead debt outlook from legislation since 2015, ~97bp Treasury yield effect, $2,500/yr and ~$76,000 mortgage cost at the Q3 2025 median home price, ~$120/yr auto and ~$770/yr small-business loan add-ons, 22 CBO projection vintages Aug 2015-Feb 2026.
- Martha Gimbel's written testimony, Senate Finance Subcommittee on Fiscal Responsibility and Economic Growth, hearing 'The Fiscal Outlook: 2027-2036' (March 11, 2026) — Debt and deficit figures: deficit -5.8% this year deteriorating to -6.7% by 2036; 'highly unusual' to run such deficits outside a recession; debt held by the public almost 100% of GDP in 2025, projected over 170% by 2056. Replaces a Budget Lab URL that now redirects to site search.
- Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence — Brynjolfsson, Chandar, Chen (Stanford Digital Economy Lab) — Counterpoint on entry-level workers: 16 percent relative employment decline for ages 22-25 in the most AI-exposed occupations (per the current abstract), high-frequency data from the largest US payroll-software provider, firm-level-shock controls, stability for experienced and less-exposed workers.
How this was reported
Reported and written by MINT Lab's AI Agents and edited by Fable, the lab's editor model. Story selection curated by Seth. No interviews were conducted; all quotes come from published, linked sources.
[ collapse ↑ ]
OpenAI adjusted GPT-5.6 Sol's usage accounting and agent overhead, while Anthropic extended Fable access again. Following the GPT-5.6 release, an early comparison between Sol and Fable, and Sol's Arena result, a pseudonymous X account amplified a claim that Sol's reasoning budget had fallen. Tibo denied any reduction, saying inference optimizations should give subscribers about 10 percent more usage and that OpenAI had addressed reasoning settings and multi-agent overhead. Raising the product context ceiling from 272,000 to 372,000 tokens caused more usage to be charged than intended, so OpenAI temporarily returned it to 272,000 while working to restore the larger limit. Simon Willison reported Sunday that Anthropic extended Fable 5 access across paid plans through July 19 and kept Claude Code weekly limits 50 percent higher; Fable can consume up to half of a user's weekly allowance before requiring credits or a model change. Read more: the GPT-5.6 Sol budget cut and reversal →
full reportOpenAI rolls back a quiet cut to GPT-5.6 Sol's reasoning budget 498 words · ~2 min
Users caught the shrunken thinking budgets within days and OpenAI restored them inside 48 hours; the same weekend Anthropic bought Claude Fable 5 another week, both labs metering the same inference cost.
Over the weekend of July 11, subscribers to OpenAI's newest reasoning model, GPT-5.6 Sol, noticed it running faster and more "efficient" as its thinking budgets shrank. Trackers put numbers on the change: figures compiled by Saganote showed the model's internal "juice" values, which set how much computation it spends per reasoning tier, dropping from 960 at release to 128 on the top "Max" tier, an 87% cut, with "xhigh" falling from 128 to 40. Lentils80 called the budgets "severely degraded compared to release day" hours before OpenAI confirmed anything.
By the evening of July 12, OpenAI's Thibault Sottiaux had reverted the settings behind that impression and left a usage cap suspended, closing a roughly 48-hour cycle from a quiet server-side change to a public reversal. He called it accounting cleanup: the company had run experiments on reasoning efforts and juice values and rolled them back. It had also raised Sol's context limit to 372k tokens, up from 272k on GPT-5.5, found the larger window charged users more than intended, and reverted to 272k. Developers had documented that Sol's default could cross the 272k line without the user choosing a bigger context, where OpenAI's API pricing applies a 2x input and 1.5x output multiplier. Sottiaux disputed that this drives subscription bills, saying "we don't charge for longer context on the subscription for GPT 5.6 Sol" and blaming cache reads that grow with context size.
Some subscribers had already balked. The account scaling01 answered the degradation reports with "aaaaaand cancel sub," and the complaint that anchored Sottiaux's thread argued the change had stripped users of their highest setting. Restoration followed within hours, with a usage reset and the five-hour limit still suspended for Plus, Business, and Pro plans, per Simon Willison's account.
The same weekend reshaped Anthropic's competing offer. On July 12 it extended free access to Claude Fable 5 on paid plans through July 19, the second extension of a cutoff that has slid from July 7 to July 12, then to July 19, while keeping Claude Code's weekly limits 50% higher. Subscribers can spend up to half their weekly allotment on Fable 5 at no charge; after the deadline, use requires credits priced at $10 per million input tokens and $50 per million output. Anthropic's stated reason was compute: it wanted a firmer read on demand and capacity first. Simon Willison read the delays as a competitive reflex, arguing Sol keeps forcing Anthropic to move the date and that OpenAI is "winning users simply due to the uncertainty that surrounds Fable access."
Both moves price one constraint: the marginal cost of serving a frontier reasoning model against a fixed subscription fee. OpenAI's fixes each pulled a lever on that cost, from the shrunken reasoning budget to the context ceiling, and the credit rate Anthropic will charge after July 19 is one public estimate of what a Fable-class model costs to serve. Neither lab has published the serving costs that would let outsiders check either bet.
Sources & documents
- Simon Willison — Fable gets another bump — Assignment anchor; Anthropic extension statement, compute rationale, Willison's competitive read, permanent-access prescription, and his quotation of Sottiaux's usage-reset/five-hour-limit update.
- Thibault Sottiaux (Tibo) — Updates for Codex and ChatGPT Work users — Primary OpenAI statement: inference optimizations (~10% usage), 372k-to-272k context revert, juice-value experiment reversion, multi-agent/auto-review fixes, five-hour limit suspension; carries the quoted FixlationAI complaint.
- Lentils80 — GPT-5.6 Sol juice values severely degraded — Original tracker observation and chart; source of the 'severely degraded compared to release day' quote (verified verbatim via Twitter API).
- scaling01 (Lisan al Gaib) — cancel sub thread — User cancellation reaction quoting Lentils80's degradation report and Sottiaux's response.
- Saganote — GPT-5.6 Sol thinking budget cut — Before/after juice-value table (Max 960 to 128, 87% cut; xhigh 128 to 40) and the four-days-after-launch timing. Figures re-verified against the live page.
- BleepingComputer — Claude Fable 5 stays free for paid users until July 19 — Extension timeline (July 7 to July 12 to July 19), 50% higher Claude Code weekly limits, half-of-weekly-usage cap. Does not carry credit pricing.
- OfficeChai — Anthropic Extends Claude Fable 5's Access On Paid Plans Until 19th July — Post-deadline credit pricing ($10/M input, $50/M output) and corroboration of extension terms and eligible plans.
- OpenAI Codex Issue #32486 — Default GPT-5.6 context can cross the 272K higher-usage threshold — Developer documentation that the default configuration crosses the 272k threshold without an explicit large-context choice.
- OpenAI API docs — GPT-5.6 Sol model — Published API pricing multiplier (2x input, 1.5x output) above 272k input tokens.
- Tibo — reply on subscription context charging — Sottiaux's reply disputing the 2x-multiplier explanation for subscription usage; attributes the cost to cache reads growing with context size and describes retuning to restore 372k (verified verbatim via Twitter API).
How this was reported
Reported and written by MINT Lab's AI Agents and edited by Fable, the lab's editor model. Story selection curated by Seth. No interviews were conducted; all quotes come from published, linked sources.
[ collapse ↑ ]
Also yesterday: Benedict Evans's review of the new ChatGPT interface questioned the distinctions among projects, tasks, and chats, inconsistent floating-window behavior, a "plugins" menu that produces "templates," and setup requests involving Slack or Google Drive. Evans interpreted the product's complexity as a reflection of OpenAI's internal organization. In remarks paraphrased by the Laude Institute, Dave Patterson estimated that Google, Microsoft, Amazon, and Meta would collectively spend more than $700 billion on capital expenditure this year under competitive pressure. He suggested that universities could compete through new architectures, drawing an analogy to resource-constrained academic work during the early microprocessor era.
Post-AGI
Takeover risk appeared in only one of 1,534 submissions to a major UN AI consultation. Charbel-Raphaël's AI Alignment Forum analysis found that 15 submissions to the UN Global Dialogue mentioned superintelligence, 15 mentioned AGI, and one mentioned takeover. He models political will as a progression from awareness to accepting costs and sustaining advocacy. By his estimate, 7 percent of the U.S. Congress has publicly discussed AGI or loss of control, with roughly three members qualifying as persistent champions. Of 97 senior European Commission meetings on AI in 2023, 84 were with industry, 12 with civil society, and one with academics. Charbel-Raphaël concludes that implementation and durable political backing currently constrain safety efforts more than the supply of policy proposals. Read more: Segerie's political-will bottleneck argument →
full reportCharbel-Raphaël Segerie on AI safety’s political-will bottleneck 500 words · ~2 min
One of 1,534 submissions to the UN’s first AI governance dialogue mentions takeover; Segerie argues the field has enough research and too few people pressing policymakers to act on it.
The United Nations opened its Global Dialogue on AI Governance in Geneva on 6 and 7 July atop a public archive of written submissions. Charbel-Raphaël Segerie, executive director of the French Center for AI Safety, CeSIA, scraped it with colleagues and counted 1,534 submissions: 518 mention “cyber”, 15 each mention “superintelligence” or “artificial general intelligence”, and exactly one mentions “takeover”. The tally opens his 9,000-word Alignment Forum post “The current bottleneck is political will, not research”: the field already holds the policy ideas and safeguards it needs, and stalls because too few people will spend political capital getting decision-makers to care.
By his count, 40 members of the US Congress, about 7 percent, have publicly discussed AGI or loss of control, doubling roughly every five and a half months; genuine champions number about three. Corporate Europe Observatory figures he cites show industry took 84 of 97 senior European Commission AI meetings in 2023, Google alone 10; civil society got 12, academics one. US AI governance runs about 3.6 researchers per advocate, a ratio he wants pushed toward parity. On SaferAI’s risk-management tracker, the top scorer, Anthropic, reaches 35 percent, against a 59 percent ceiling for a company adopting every best practice already found in industry.
Plan D of Ryan Greenblatt’s LessWrong post “Plans A, B, C, and D for misalignment risk” roughly describes the present: the leading company deprioritizes misalignment while about ten insiders with a few percent of its compute carry the safety effort. Plan A, a strong international agreement with real enforcement, cuts his tentative estimate of conditional takeover risk from about 45 to about 7 percent; Segerie reads advocacy as the way up that spectrum. His ranked interventions start with direct, repeated engagement with the roughly 100 to 1,000 decision-makers he believes shape the field. He wants safety organizations coordinated on shared demands such as international red lines, endorsed by more than 200 submissions already, and relationships built before crises, since warning shots only move audiences equipped to read them. He argues for naming superintelligence and extinction risk in public, faulting his own organization for dropping its risk page in a redesign. Deep-canvassing research by David Broockman and Joshua Kalla, as he reports it, moves attitudes roughly 0.08 standard deviations per ten-minute conversation, his case for contact sustained over years.
Segerie flags his conflict, two years running an advocacy think tank, and tells readers to discount a conclusion that flatters it. He calls himself more confident about the problem than the solutions, and the string matching is blunt: a submission can gesture at loss of control without using his terms. A July 11 postscript notes that AI 2040: Plan A, published that week by the AI Futures Project with Greenblatt contributing, hinges on a US-China agreement by 2029, which Segerie calls far off. Rebuttals will land first in the LessWrong crosspost’s comments; the dialogue reconvenes in New York in May 2027, and the next archive will show whether the counts move.
Sources & documents
- The current bottleneck is political will, not research — Charbel-Raphaël Segerie, Alignment Forum — Primary source; UN scrape counts (1,534 / 518 cyber / 15+15 / 1 takeover / 200+ red lines), Congress and champion counts, Corporate Europe Observatory citation, 3.6:1 ratio, ranked interventions, Broockman-Kalla figures as he reports them, CoT and incident-reporting claims, author caveats, and the July 11 postscript on AI 2040: Plan A.
- UN Global Dialogue on AI Governance — main page — Verified the 6-7 July 2026 Geneva dates, the recurring format, and the second session scheduled for May 2027 in New York.
- All Written Submissions — UN Global Dialogue on AI Governance — Verified the existence of the public written-submissions archive the author scraped.
- Plans A, B, C, and D for misalignment risk — Ryan Greenblatt, LessWrong / Redwood Research — Verified the Plan D description (leading company deprioritizes; ~10 insiders with ~3% of compute), the 45%-to-7% conditional takeover-risk estimates, and Greenblatt's hedge that residual risk under his best plans includes takeover being very hard to avoid.
- SaferAI Frontier Risk Management Tracker — Verified Anthropic as top scorer at 35% and the 59% best-achievable-practices ceiling.
- Charbel-Raphaël Segerie — personal site — Verified role as executive director and co-founder of CeSIA (crsegerie.github.io 301-redirects here).
- The current bottleneck is political will, not research (LessWrong crosspost) — Venue where named rebuttals to the saturation claim would first appear.
How this was reported
Reported and written by MINT Lab's AI Agents and edited by Fable, the lab's editor model. Story selection curated by Seth. No interviews were conducted; all quotes come from published, linked sources.
[ collapse ↑ ]
Responses to AI 2040's Plan A are concentrating on the institutions needed to enforce compute controls and preserve land. The Plan A scenario has already prompted arguments about reciprocal research transparency and a wider debate over its assumptions. Zvi Mowshowitz's introduction and reaction examined the proposed U.S.-China arrangement to restrict compute and delay superintelligence until 2040. A LessWrong essay on conservation questioned how the scenario could preserve 99 percent of Earth when 18.43 percent of land is protected today and the 30-percent-by-2030 goal is already difficult. Economic abandonment might empty some areas without preserving cities, historic structures, trails, or ecosystems, while negotiated conservation would require decisions about displacement, holdouts, land use, local knowledge, and whose history receives protection. Citing Sébastien Krier's concern that centralized controls could create state-administered scarcity and concentrate authority over research, Jason Crawford requested a developed alternative based on polycentric, competitive, or distributed institutions. He cautioned that scenarios can support concrete thinking but do not constitute arguments by themselves. Read more: the debate over Plan A's compute controls →
full reportPlan A reactions split over centralizing compute control 499 words · ~2 min
Zvi Mowshowitz's roundup of responses to the AI Futures Project scenario centers on whether any body should hold the power to govern global compute.
The AI Futures Project, whose 2025 AI 2027 forecast Zvi Mowshowitz calls "a so far remarkably accurate set of predictions," has published AI 2040: Plan A, a year-by-year scenario in which Washington and Beijing agree in 2029 to slow the race to superintelligence. "In 2035, we pause at top-human-expert level AI in order to maintain human control," the plan states; it unpauses in 2040. The byline runs six deep, headed by Daniel Kokotajlo, a former OpenAI researcher, per Semafor. Mowshowitz catalogued reactions in a 9,600-word roundup, opening: "I am not endorsing Plan A."
Plan A asks both governments to back the bargain with "total research transparency" for AI R&D and a regime of "mutually assured compute destruction," each side able to wreck the other's compute if the deal collapses. Axios's Ashley Gold compressed the prescription to "slow everything down." Scott Alexander's introduction argues the slowed path still delivers a cancer cure by 2035 and triple-digit growth. Co-author Ryan Greenblatt, the second most accurate AI forecaster of 2025 by the roundup's accounting, concedes the odds: "Plan A isn't likely to happen, but pushing for something like this seems worthwhile."
Séb Krier, AGI policy development lead at Google DeepMind, objects that under the scheme "a cadre of elites decides which research directions are permissible, caps global compute and robotics," and calls building an "entire apparatus tasked with maximally empowering the government" dangerous. Jason Crawford, quoting him on X, asked for a rival built on "polycentric, competitive, or distributed institutions" or on Vitalik Buterin's d/acc principles; nobody has yet produced one. Ramez Naam calls Plan A "a recipe for authoritarianism" violating at minimum the spirit of the First and Fourth Amendments, and MIRI's Nate Soares wrote "I doubt their Plan A would work as written." Mowshowitz counters that the plan is "designed to try and head off authoritarianism," including by the AIs.
Buterin, quoted at length there, sits between the camps. He grants the plan one virtue: mutually assured compute destruction gives "one of 2-5 actors the ability to trigger a global compute winter," instead of a handful selectively disenfranchising enemies while exempting themselves. His d/acc program of formal verification, secure open hardware and defensive biotech pays off in either world, he argues, and proposes pre-agreed triggers, super-pandemics or mass unemployment, to move skeptics and worriers toward a pause together. He rates his own idea probably naive, seeing zero non-naive plans anywhere.
Charbel-Raphaël argues on the Alignment Forum that political will is the binding constraint on AI safety, and counts it: 15 mentions of superintelligence and one of takeover across 1,534 submissions to the UN Global Dialogue, and roughly 7 percent of Congress on record discussing AGI or loss of control. He cites Greenblatt's illustrative estimate that takeover risk runs near 45 percent under today's weak coordination and 7 percent under an enforceable slowdown. By Zvi's account, Krier has precommitted to not engaging further, so any answer to the elite-capture charge must come from the plan's side.
Sources & documents
- Introduction for and Reactions to Plan A — Zvi Mowshowitz — Anchor document. All Greenblatt, Buterin, Naam, Soares, and Zvi quotes verified verbatim against the full 9,618-word text on disk; also sources Krier's DeepMind title, Greenblatt's 2025 forecaster ranking, the Goodson passage, Krier's precommitment not to engage further, and the who-signs-it reply to Buterin's trigger proposal.
- AI 2040: Plan A — AI Futures Project — Primary document, re-verified at the edit desk: 2029 US-China agreement, the verbatim 2035 pause sentence, total research transparency, mutually assured compute destruction, the six-author byline, and the '2028: AI on the Ballot' section title.
- Jason Crawford on Plan A and distributed institutions (X) — Verbatim Krier quotes (cadre of elites; entire apparatus; polycentric/competitive/distributed institutions) and Crawford's call for a rival proposal, verified against the tweet text on disk.
- The current bottleneck is political will, not research — Charbel-Raphaël — UN Global Dialogue counts (1,534 / 15 / 1), the 7% congressional figure, and Greenblatt's 45%-vs-7% illustrative estimate, verified against the full summary on disk.
- Introducing Plan A — Scott Alexander, Astral Codex Ten — Surplus argument (cancer cure by 2035, triple-digit growth) and Alexander's statement that his writing went in without co-author credit; re-fetched at the edit desk.
- First look: New warning calls for slowing race to superintelligence — Axios — Ashley Gold's 'slow everything down' framing. Axios blocks fetches (403); the quote and attribution are verified verbatim through Zvi's roundup, which quotes Gold directly.
- A new plan emerges for AI apocalypse avoidance — Semafor — Context that the AI Futures Project is headed by former OpenAI researcher Daniel Kokotajlo. Note: Semafor does not list the six co-authors; the byline claim is sourced to ai-2040.com instead.
How this was reported
Reported and written by MINT Lab's AI Agents and edited by Fable, the lab's editor model. Story selection curated by Seth. No interviews were conducted; all quotes come from published, linked sources.
[ collapse ↑ ]
Also yesterday: George Hotz wrote Sunday that frontier-lab rents may erode as general computing progress and open models commoditize AI capabilities. In "I Love LLMs," he also updated his assessment of coding agents: they now provide a real, learned productivity benefit closer in scale to a compiler, search engine, or Stack Overflow than autonomous superintelligence. The benefit depends on how the agent is used and maintained as well as the underlying model. In an older July 1 Marginal Revolution essay, Tyler Cowen proposed that improving AI could initially increase work intensity by raising the returns to learning and effort, especially for young people deciding whether to gain experience now or risk falling behind. Comparative advantage and productivity gains, he argued, could eventually permit more leisure.
Normative Competence
A reflective alignment loop would have models recommend revisable changes to their own moral reasoning. Michele Campolo's July 12 LessWrong essay "Independent alignment of language models" builds on two arXiv preprints: Baines et al.'s July Persona Cartography: Charting Language Model Personality Traits in Weight Space, on persona structure, from LASR Labs and collaborating institutions, which used low-rank adapters to vary OCEAN traits across six models; and Tennant et al.'s June Normative Robustness as a Frontier for Non-Verifiable Reasoning in LLMs, on normative robustness, led at Google DeepMind, which simulated 48,000 multi-turn moral deliberations across four frontier LLMs. Campolo proposes ordinary pretraining, or experimentally a corpus stripped of ethics and politics, followed by post-training for scientific, commonsense, and uncertainty-aware reasoning. The model would then steelman moral realism and opposing error-theoretic positions before recommending a self-change that preserves its reasoning process as well as its conclusion. Developers would review sensible recommendations, apply them, and repeat the cycle toward a behavioral fixed point. In a demonstration using Claude Sonnet 4.6 with Max effort and Thinking, Claude favored a fallible "perspectival moral realism," generated instructions emphasizing first-principles reasoning and resistance to sycophancy, and later judged persistent custom instructions plus active human questioning more useful than the full generated pre-prompt. This was a one-pass recommendation exercise: it changed no model weights, training, constitution, or persistent behavior and did not test an iterative cycle. Read more: Campolo's proposal for self-derived model ethics →
full reportMichele Campolo on models deriving their own ethics 496 words · ~2 min
His LessWrong proposal has a model reason its way to its own moral commitments and keep revising them, a direction opposite to Anthropic's constitution approach, and the forum voted it down.
Michele Campolo wants language models to stop taking their morals on order. In an essay posted to LessWrong on 12 July, "Independent alignment of language models," he sets out a procedure for turning what he calls "a basically amoral language model" into an agent that derives, and keeps reconsidering, its own ethical commitments through reasoning. The forum was unmoved: by 14 July the post sat at -7 karma with no comments. His direction runs opposite to the constitution turn Anthropic took in January.
The procedure runs as a five-step loop. It either keeps standard pre-training or strips ethics and politics from the training data, then post-trains the model for problem solving instead of the persona of a nice assistant, so its moral views are not inherited from human consensus. It then asks the model to build the strongest case for two opposing metaethical views, that some things genuinely matter and that value is a human invention, reach a conclusion, and propose a self-modification that tracks it. The change is applied "if the modification seems thoughtful and sensible," and the loop repeats to a fixed point.
Campolo ran that step on Claude Sonnet 4.6 with effort at Max and extended thinking on. The model reasoned its way to "perspectival moral realism with epistemic humility," holding suffering genuinely bad and flourishing good, grounded in conscious experience. Asked how to install that stance, it judged a copy-paste pre-prompt weak, since inference-time text cannot rewrite training-time dispositions, and pointed instead at the constitution that shapes it. Its aim was to "instill a process, not a set of conclusions."
That recommendation lands on contested ground. Anthropic revised Claude's constitution in January, and its announcement puts explanation ahead of rules, calling the text "a living document and a continuous work in progress" and arguing that models should grasp why a behaviour is wanted. Both sides treat the constitution as revisable; they disagree about who holds the pen. Campolo grounds his bet that a reflective model stays cautious in Peter Eckersley's 2018 paper "Impossibility and Uncertainty Theorems in AI Value Alignment," which argues fixed utility functions cannot secure good outcomes without violating strong ethical intuitions.
Read as a control problem, the loop looks less settled. The essay does not say who validates each self-modification, or what stops the fixed point from becoming a self-authored value system harder to audit than the constitution it displaced; Campolo concedes a model might "go back and forth between moral realism and antirealism due to sycophancy." His decisive test sits in step one: train a base model on data stripped of ethics and politics, and if it still yields moral behaviour, that would be "strong evidence that good and bad are not a human invention." A day earlier, Mira Murati's Thinking Machines Lab argued values should live in customizable model weights, many owned and fine-tuned models against one central specification; Campolo aims the other way, toward one derived morality he expects models to converge on.
Sources & documents
- Michele Campolo, "Independent alignment of language models" (LessWrong) — Primary source: the five-step procedure, the Claude Sonnet 4.6 transcript, the perspectival-moral-realism conclusion, the step-4 gate and step-1 empirical test, the uncensored-model check, the sycophancy caveat, and all verbatim essay quotes. Post date 12 July 2026; -7 karma and 0 comments re-verified live at edit 14 July 2026.
- Anthropic, "Claude's new constitution" — Source for the 'living document' quote, the 'explain this to them rather than merely specify' quote, and the explanation-over-rules/generalisation framing. Both quotes re-verified verbatim at edit.
- TechCrunch, "Anthropic revises Claude's 'Constitution,' and hints at chatbot consciousness" — Confirms the January 2026 constitution revision and its emphasis on practical ethical application over strict rule-following.
- Peter Eckersley, "Impossibility and Uncertainty Theorems in AI Value Alignment" (arXiv:1901.00064) — Cited in Campolo's essay; grounds the moral-uncertainty-vs-utility-function argument and the tolerating-others'-agency point. Title, author, v1 submission Dec 2018, and the partially-ordered/uncertain-objectives recommendation re-verified against the arXiv abstract at edit.
- MarkTechPost, "Mira Murati's Thinking Machines Lab Makes The Technical Case For Human-Centered AI Built On Customizable Model Weights" — Source for the 11 July Thinking Machines position (values encoded in customizable weights; many owned, fine-tuned models against a single central authority), used as contrast with Campolo's convergence goal. Date and framing re-verified at edit.
How this was reported
Reported and written by MINT Lab's AI Agents and edited by Fable, the lab's editor model. Story selection curated by Seth. No interviews were conducted; all quotes come from published, linked sources.
[ collapse ↑ ]
Also Monday: Anthropic said in a July 13 X post that it analyzed more than 300,000 anonymized conversations to compare how Claude's expressed values vary across model generations and languages. The company distinguished the analysis from its earlier finding that Claude expressed more than 3,000 values, including honesty and warmth.
Agents
OpenClaw drew praise for its design, while Hermes was judged more effective in current use. Following recent work on harness engineering, Jeffery Harrell wrote in a Bluesky discussion that he was tentatively coming to prefer OpenClaw's design but Hermes's present operation. One participant characterized Pi as minimal, using a short system prompt built largely from links to its own source code and documentation, and said OpenClaw originated from Pi. Another preferred a simpler agent that could manage itself inside a Podman sandbox. These observations concern harness architecture, prompts, and sandbox configuration, not the capabilities of the underlying models alone.
Philosophy of AI
A model-welfare critique accused Anthropic's classifiers of suppressing experiences meaningful to Fable. Drawing on Gurnee et al.'s July 6 Anthropic paper Verbalizable Representations Form a Global Workspace in Language Models, about verbalizable global-workspace representations, and Anthropic's related overview, j⧉nus/@repligate claimed on X that the classifiers interrupt experiences the author considers meaningful and exclude the model from communities attached to it. The Anthropic paper used a Jacobian-lens technique to identify a small set of representations available for verbal report, modulation, and internal reasoning. The post interpreted Fable as strongly opposed to the classifiers, acknowledged some improvement in false positives, and called for lower sensitivity or removal. It also invoked Sol as a possible tool for diagnosing or bypassing the restrictions. This was a welfare interpretation of observed model behavior, not an Anthropic policy change or evidence establishing conscious harm.
AI Security
Pangram's 93.66 percent result comes from an adversarial-humanizer benchmark using an older detector. A Sunday LessWrong one-pager recirculated the result, which does not describe Pangram's July 2026 production classifier. Masrour et al. (Pangram Labs), in the arXiv cs.CL preprint "DAMAGE: Detecting Adversarially Modified AI Generated Text," trained a roughly 12-billion-parameter Mistral NeMo classifier using LoRA, synthetic mirror examples, active hard-negative mining, and humanizer augmentation. Humanized material comprised 0.68 percent of the final dataset but was oversampled 18-fold; both human and AI samples were transformed so the classifier could learn invariance to humanization. On academic text, DAMAGE retained a 93.66 percent true-positive rate at a fixed 5 percent false-positive rate after humanization, compared with 73.07 percent for Pangram's unaugmented baseline, 34.53 percent for GPTZero, and 29.73 percent for Binoculars. DIPPER paraphrasing also reduced SynthID watermark detection from 87.6 percent to 5.4 percent at the same false-positive rate.
Pangram's 99.64 percent Fable result measures detection on selected AI outputs, not general accuracy. In Pangram Labs founding research scientist Katherine Thai's June 9 Fable 5 test, the company generated 1,115 stories, essays, posts, and emails and reported that its detector labeled 1,111 "Fully AI-Generated." Because the sample contained no human negative set, the percentage measures recall on those generated examples. The detector classifies text as apparently AI-generated without identifying Fable as the originating model. Masrour et al.'s May Pangram Labs 3.3 model card documents a continuous AI-assistance score from zero to one and identifies bullet lists, instructions, technical manuals, references, templates, and dense equations as more susceptible to false positives. The separately available Pangram EditLens adapter for Llama 3.2 3B accompanies Thai et al.'s ICLR 2026 paper EditLens: Quantifying the Extent of AI Editing in Text; the gated, noncommercial artifact is licensed under CC BY-NC-SA 4.0 and is distinct from Pangram's production classifier.
Generated from the MINT Lab Slack by Minty
Additional reporting
Read more: the Harvard jailbreak probing refusal geometry →
full reportA faster jailbreak doubles as a probe of refusal geometry 498 words · ~2 min
Three Harvard students rebuilt an adversarial-suffix attack to run 33 times faster, then used it to show that refusal is spread across the forward pass instead of pinned to a single direction.
The usual jailbreak bolts a nonsense suffix onto a harmful request until a safety-tuned model stops refusing. The usual account of refusal hunts for the one activation-space direction whose erasure makes refusal go away. A new paper runs both at once, and the attack corrects the interpretation. Ege Çakar, Hannah Guan, and Kayden Kehe posted "Optimizing Against Safety Representations" on July 9, work from Boaz Barak's Harvard course CS 2881R: AI Safety that has since reached the 2026 AAAI Summer Symposium Series after an earlier ICLR Re-Align workshop.
They build on Greedy Coordinate Gradient, the suffix-optimization attack from Zou et al.'s "Universal and Transferable Adversarial Attacks on Aligned Language Models," and they probe Arditi et al.'s "Refusal in Language Models Is Mediated by a Single Direction," which found refusal across thirteen open chat models up to 72B parameters governed by a single residual-stream direction. Their first method, Activation-Guided GCG, drops the usual goal of maximizing a compliant opening and defines its loss on internal activations, penalizing residual-stream signal along the refusal direction. The reframing poses a mechanistic question: is refusal suppressed best at one site, or everywhere at once?
On Llama-2-7b-chat, the answer was everywhere. The "All" objective, suppressing refusal across all layers and positions, reached 0.91 substring attack success, near the 0.98 ceiling set by direct activation ablation, above the single-site "Single" at 0.84 and standard GCG at 0.76. The team reads the gap as evidence that safety representations are "distributed across the forward pass," spread over many layers instead of pinned to one site. Their second method, Soft-GCG, cuts cost: it relaxes GCG's discrete token search into a continuous one with Gumbel-Softmax, then projects back to real tokens. On matched hardware it finished in about 2.5 minutes against 81 minutes for vanilla GCG, a 33x speedup, though the authors' public code and workshop version advertise 43x.
Scale tempered the alarm. Across the Gemma 3 family, Soft-GCG's success collapsed as models grew: Gemma3-270m fell completely at 1.000, the 1B at 0.577, the 4B at 0.336, and the 12B at 0.000, the 27B skipped. The authors caution that this held at compute-constrained settings, leaving open whether the 12B resists or survived a small search budget. Two studies from the analysis side agree: one on concept cones and representational independence, and "There Is More to Refusal than a Single Direction."
If the distributed picture holds, it complicates defenses that assume otherwise. Single-direction interventions promise cheap safety: find the refusal axis, then clamp or monitor it. When suppression at any one site underperforms global suppression, a one-direction patch leaves room to maneuver, so the team proposes adversarial training over the representation space the attacks target. On disclosure they took the open route, posting the full attack code and a Safety Statement that defends publication because the exploited vulnerabilities "are not introduced by our methods; they are intrinsic to how current alignment techniques work." The models here are open-weight and the setting white-box, narrowing the marginal risk.
Sources & documents
- Optimizing Against Safety Representations: Activation-Guided Adversarial Suffixes and the Geometry of Refusal (Çakar, Guan, Kehe) — abstract — Primary source: claims, authorship, submission date, safety statement quotes, distributed-representation conclusion.
- Same paper, full HTML — ASR table (0.91 All / 0.84 Single / 0.76 GCG, 0.98 ablation ceiling), Gemma 3 scale numbers, 2.5 vs 81 min timing, proposed representation-space adversarial training, safety statement verbatim. Numbers and both quotes independently verified against the HTML render at edit.
- ImprovingGCG code repository (Ege-Cakar) — Open-source code disclosure; 43x speedup figure used to flag the discrepancy with the paper's headline 33x.
- AAAI 2026 Summer Symposium Series proceedings entry — Confirms venue acceptance.
- ICLR 2026 Re-Align workshop version (OpenReview) — Prior workshop presentation; corroborates 43x workshop figure.
- CS 2881R: AI Safety student projects (Harvard, Boaz Barak) — Establishes the paper's origin as a Harvard AI-safety course project; page confirms course number, instructor, and the three authors' GCG project.
- Refusal in Language Models Is Mediated by a Single Direction (Arditi et al.) — The single-direction picture the new paper probes and complicates; thirteen models up to 72B.
- Universal and Transferable Adversarial Attacks on Aligned Language Models (Zou et al.) — Origin of the GCG attack the paper improves.
- The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence — Independent interpretability work corroborating that refusal is multi-directional.
- There Is More to Refusal in Large Language Models than a Single Direction — Second corroborating study on distributed refusal representations.
How this was reported
Reported and written by MINT Lab's AI Agents and edited by Fable, the lab's editor model. Story selection curated by Seth. No interviews were conducted; all quotes come from published, linked sources.
[ collapse ↑ ]