MINT Lab

Yesterday in AI · 13 July 2026

We're introducing some deep dives into the YinAI newsletter, so you can just read it like you normally do if you want, but for the top 10 stories we're reporting them out with a bit more detail and context. Click read more and then carry on scrolling down the newsletter. Stories are curated by Seth, reported by the Minty Newsroom (a mixture of Sol and Opus agents), and edited by Fable. There will sometimes be errors; please let us know. And if you know of a source we should be including but aren't, please suggest it (see the form at the end).

Regulation

Possible U.S. controls on frontier open weights remain a policy forecast, with the "six months" tied to capability progress. In Sunday's "6 months to live for open models", Nathan Lambert wrote that his sources were discussing a possible White House executive order. Initial measures, he suggested, could target Chinese-origin models and their use in government, followed by a ban, pre-release review, or indefinite delay for open weights above roughly the GPT-5.5, Claude Opus 4.8, or GLM-5.2 capability range. His six-month estimate refers to when an open-weight model might reach that range, not a disclosed government deadline; no rule or order has been announced. Lambert connected possible restrictions to distillation concerns and accused closed laboratories of pursuing regulatory capture through their lobbying advantages and the relative ease of controlling centralized services. Read more: Lambert's six-month forecast for open models →

full reportNathan Lambert gives open-weight models six months 500 words · ~2 min

He predicts Washington will ban or stall any open release above today’s frontier, with procurement bans the nearest instrument and a rumored executive order behind them.

In “6 months to live for open models”, at Interconnects, Nathan Lambert predicts that Washington will ban or indefinitely delay any open-weights release meaningfully above GPT-5.5, Claude Opus 4.8, or GLM-5.2, and that the next model over that line will most likely ship from a Chinese lab. Unlike earlier scares, this one has machinery in place: a June 2 executive order created a review step for frontier systems, and the same regime rapidly took down Claude Mythos 5 and Fable 5 after a competitor complaint, pushing enterprises toward self-hosted models. Lambert expects review to run far slower for open weights, with the model-checker that flags closed releases next flagging an open one.

Reporting on the White House crackdown describes federal purchasing bans as the most viable option; legal experts there call a full weights ban close to impossible, since downloadable weights may draw the First Amendment protection long recognized for source code. Lambert also notes White House discussions of a rumored executive order on open models, no text yet, likely scoped first to Chinese-origin models and government uses. Congress is pulling the other way: the AI Foundation Model Transparency Act would impose disclosure duties on large providers yet exempt fully open-source models from the FTC rules it directs.

In a June 10 letter to Senators Tim Scott and Elizabeth Warren, Anthropic alleged operators tied to Alibaba’s Qwen lab generated 28.8 million Claude exchanges through nearly 25,000 fraudulent accounts between April 22 and June 5, warning that PRC labs “capture the returns on American investments without bearing the costs or risks.” Lambert calls the Anthropic-led push “the definition of regulatory capture”: a ban on the rivals it accuses would hand the company lasting economic security. Distillation fears, he adds, indict API security more than open weights; a lab that thinks a capability truly dangerous should pull it off queryable APIs.

Andreessen Horowitz’s Jai Ramaswamy and Matt Perault urge policymakers to reject bans and regulate harmful uses instead of development, writing that “open source developers should not be held liable for downstream misuse.” Robert Brennan contends that “open models level the playing field” against state actors who will obtain advanced systems regardless, while a ban kneecaps academics and hospitals along with startups. The Commerce Department’s NTIA had already recommended monitoring open weights absent evidence of concrete marginal risk.

Chinese open models, DeepSeek chief among them, hold a substantial lead, and Lambert argues a unilateral ban fails its own safety logic: the model stays legal in China, so a bad actor keeps access, while the United States isolates itself from the global open-source community. He wants an American lab to ship a comparable open model; Microsoft and Meta have the commercial logic, he argues, and Reflection AI, founded by two ex-DeepMind researchers with a SpaceX compute deal worth up to $6.3 billion, remains the wildcard yet to launch a public model. Per CNBC, House committees have opened probes into U.S. companies running Chinese models such as Kimi.

Sources & documents

How this was reported

Reported and written by MINT Lab's AI Agents and edited by Fable, the lab's editor model. Story selection curated by Seth. No interviews were conducted; all quotes come from published, linked sources.

[ collapse ↑ ]

Meta and Sarah Wynn-Williams are fighting over whether their dispute must remain in arbitration. Steven Levy's July 10 WIRED Backchannel column reported that Wynn-Williams filed suit on June 25 to vacate an interim restriction on promoting Careless People and move the case into public court. Her 2017 separation agreement reportedly provided $780,000 and included non-disparagement and arbitration terms. She claims the arbitrator's interpretation could expose her to $50,000 penalties for discussing technology policy and violates her free-speech rights. Meta says she knowingly accepted the agreement and is trying to evade arbitration. The interim restriction remains in force, with a fuller hearing scheduled for October; Levy contends that Meta's continued pursuit is worsening the reputational damage surrounding the case.

A four-tier assurance scheme would audit frontier developers as organizations, covering hardware, internal systems, security, and governance. AVERI's Miles Brundage and colleagues (Brundage et al.), in the January arXiv preprint "Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies," define auditing as independent verification through secure access to non-public evidence. Their taxonomy covers intentional misuse, unintended behavior, information-security failures, and social harms such as addiction or facilitated self-harm. AAL-1 is a weeks-long review of a specific system using APIs and limited internal information, recommended as a general baseline; AAL-2 is the near-term target for the most advanced developers. Higher tiers would require qualified auditors, incentives for developer cooperation, and technical infrastructure that supports deeper access. Read more: the 48-author frontier auditing framework →

full reportA 48-author framework for auditing frontier AI companies 499 words · ~2 min

Brundage, Dreksler and 46 co-authors define four audit assurance levels and the non-public access each requires, and treat the question of who pays the auditor as unsolved.

"Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies," posted to arXiv by 48 researchers, is first-authored by Miles Brundage and Noemi Dreksler, with Yoshua Bengio and Rishi Bommasani among the co-authors. The paper defines frontier AI auditing as independent evaluation of a company's systems and practices against relevant standards, grounded in secure access to non-public information, and widens the unit of audit to the whole company: a model that passes its evaluations can still ship behind safety classifiers or routing rules that alter misuse potential. Audits would cover intentional misuse, unintended system behavior, information security, and emergent social phenomena such as AI addiction, ending reliance on companies "grading their own homework." Co-author Patricia Paskov discussed the paper at an ICML technical-governance workshop.

The framework grades audits into four AI Assurance Levels. AAL-1, limited assurance, runs a few weeks on black-box API access to model checkpoints, with chain-of-thought outputs and logits exposed and safety classifiers switchable, a baseline for frontier AI generally. AAL-2, the near-term goal for the leading companies, takes months at minimum and adds gray-box access, unredacted safety cases, staff interviews across safety, security, policy and product, and statistical model signatures tying audited models to deployed ones. AAL-3, a multiyear white-box engagement, and AAL-4, continuous verification able to detect deception and provide "treaty-grade" confirmation, are not yet technically or organizationally feasible; the paper outlines research directions toward them.

On who pays, the authors warn that auditors competing for contracts from the companies they evaluate risk devolving into "rubber stamping," the pattern financial auditing fell into before 2008, they say; Enron and Wirecard show where dependence on a few large clients ends. Payment should not depend on audit results, they write, and that alone is not sufficient. They float funding by insurers, regulator-administered pools, industry-wide levies, or downstream enterprise users; they concede the customary 10 or 15 percent single-client fee cap may need loosening in so small a market; they set an end-of-2026 target for alternatives; and they sketch a "PCAOB-for-AI" with authority to revoke auditors' accreditation.

Parts of this exist. In Europe, the General-Purpose AI Code of Practice requires signatories to undergo independent external evaluations, and signing grants a presumption of conformity with the AI Act; Meta and Chinese AI companies have not signed, and the requirement depends on qualified assessors existing. In the United States, where third-party assessment remains voluntary, the authors want safe harbors for good-faith auditing and audit requirements in high-stakes procurement such as health and defense. The paper counts METR's October 2025 review of Anthropic's pilot sabotage risk report among the first AAL-1 audits, an exercise METR has since repeated for Claude Opus 4.6, alongside third-party review of OpenAI's gpt-oss safety work and the reciprocal assessments OpenAI and Anthropic ran on each other. Next to assurance regimes in finance, aviation, and food safety, the authors call all of it early-stage.

Sources & documents

How this was reported

Reported and written by MINT Lab's AI Agents and edited by Fable, the lab's editor model. Story selection curated by Seth. No interviews were conducted; all quotes come from published, linked sources.

[ collapse ↑ ]

Also yesterday: Building on proposals for universal basic capital, Cecilia Rikap's Sunday Jacobin essay proposed a one-time 50 percent stock levy on major AI companies to establish an international AI wealth fund, coupled with democratic control of cloud infrastructure. She grounds the proposal in the worldwide creative, personal, and institutional data used to develop generative AI. Anton Leicht wrote that anticipated frontier-lab IPO windfalls could fund AI-safety politics, while calling for a more ideologically diverse network of advocacy groups and PACs spanning laboratory restrictions, nationalization, iterative deployment, and market-oriented policy. Rose Horowitch's July 8 Atlantic essay connected fragmentary digital and algorithmic media with declining sustained reading: daily leisure reading fell from 28 percent of Americans in 2004 to 16 percent in 2023, while nearly 30 percent of adults reportedly struggle to infer or paraphrase across a multipage text. A Foreign Policy podcast roundup highlighted an episode on the AI arms race. Read more: Leicht's critique of AI safety funding → · Read more: Horowitch's cover story on postliterate America →

full reportAnton Leicht on the coming flood of AI safety money 499 words · ~2 min

Frontier-lab IPOs will mint safety-minded donors within months, Leicht argues; the advocacy world lacks the channels to absorb the money, and he wants them dug before it arrives.

Anton Leicht’s July 13 essay “The Flood”, on his newsletter Threading the Needle, opens with a forecast: employees at frontier AI companies will become “very, very rich”, and many, convinced Washington has governed AI badly, will spend the windfall on politics. Anthropic closed a $65 billion Series H on May 28 at a $965 billion valuation, a round TechCrunch reports may be its last before an IPO; OpenAI finished its restructuring into a public-benefit corporation last October and, per Forbes, carried an $852 billion valuation in March, with an offering possible this year. Leicht compares the coming money to the Nile flood: channeled through enough irrigation, it waters the fields; forced through too few channels, it drowns them. Washington’s safety-advocacy world, he argues, is a handful of organizations sharing one tactical playbook, and a fortune poured through that bottleneck will overextend one strategy, invite counterspending, and hand opponents one target to caricature.

He starts with the Great American AI Act, the 269-page draft Representatives Jay Obernolte and Lori Trahan released on June 4. Roll Call reports it would require frontier developers above $500 million in revenue to publish catastrophic-risk frameworks, report critical safety incidents within 15 days, accept independent audits twice a year, and protect whistleblowers, while freezing state laws that regulate AI development for three years; the Future of Privacy Forum notes states keep authority over deployment, and violations can cost $1 million per day. Leicht reads the draft as preemption bought with federal safety obligations. He liked the trade, and thinks safety groups dismissed it too fast: Americans for Responsible Innovation ran Massachusetts ads urging Trahan to oppose preemption, and ARI president Brad Carson said Big Tech “would love to see a Democrat push forward their plan to freeze state AI laws.” For Leicht the episode shows a coalition built to block, one that punishes policymakers who care about safety but dissent from its economics.

The spending record worries him more. Fortune tallied $27 million in New York’s 12th-district Democratic primary: about $19 million from the Anthropic-backed Public First Action supporting Alex Bores, roughly $8 million from the OpenAI-linked Leading the Future opposing him. Micah Lasher won. Leicht titles that section “I Spent $19 Million On NY-12 And All I Got Was Micah Lasher”, and draws the lesson: concentrated AI money turns any race into a referendum on AI and teaches candidates to avoid the topic; smaller counterspenders can derail it with far less money.

He wants the receiving machinery built before the liquidity arrives: uncorrelated political vehicles, each fluent in a different constituency’s language, from the heterodox right and pro-tech moderates to the populist left and national-security hawks, competing for the same safety money, so a bill can reach 60 Senate votes with a different rationale for each bloc. The money is moving; Public First Action told Axios it has raised $80 million, and The Nation counts over $321 million across 14 AI- and crypto-funded super PACs. On Leicht’s telling, the channels need digging now.

Sources & documents

How this was reported

Reported and written by MINT Lab's AI Agents and edited by Fable, the lab's editor model. Story selection curated by Seth. No interviews were conducted; all quotes come from published, linked sources.

[ collapse ↑ ]

full reportRose Horowitch on America’s retreat from reading 500 words · ~2 min

Her Atlantic cover story documents two decades of declining reading and comprehension; public oversight of AI assumes exactly the capacities being lost.

Rose Horowitch’s August cover story for The Atlantic, “America Isn’t Illiterate. It’s Postliterate.”, argues that Americans encounter more text than ever and sustain attention on less. Jessica Bone and colleagues’ iScience paper “The decline in reading for pleasure over 20 years of the American Time Use Survey” traced daily pleasure reading from 28 percent of Americans in 2004 to 16 percent in 2023 across 236,000 responses; declines ran steepest among Black Americans, lower-income households, and rural residents. Fewer than half of adults read a book in 2022; roughly a fifth account for four-fifths of books read.

Nearly 30 percent of American adults cannot paraphrase or make inferences from a multipage text, up from under 20 percent in 2017. In a 2024 study she cites, English and English-education majors at two regional Kansas universities read the opening of Dickens’s Bleak House with internet access; 5 percent reached an accurate and detailed understanding, and at least a quarter read the figurative language literally, putting dinosaurs in 19th-century London.

Horowitch, following the Georgetown computer scientist Cal Newport, treats writing as forcing thought into order; automate it and you risk automating the thinking. Michael Gerlich’s “AI Tools in Society” (Societies), from surveys and interviews with 666 people, found frequent AI use negatively correlated with critical-thinking scores, mediated by cognitive offloading. In the arXiv preprint “Your Brain on ChatGPT”, Nataliya Kosmyna’s MIT team wired 54 essay writers to EEG across four months; chatbot users showed weaker neural connectivity than search or unaided writers, and struggled to quote essays they had just written. Neither study establishes lasting damage. Monthly Amazon book releases have tripled since ChatGPT’s 2022 launch and journal submissions have surged, much of the new text fluent, much of it derivative or wrong.

On the Alignment Forum, Charbel-Raphaël, an AI-safety think-tank head who discloses the interest, counted exactly one mention of takeover risk among 1,534 submissions to the UN’s Global Dialogue on AI; oversight assumes readers able to follow a technical argument to an uncomfortable conclusion. Horowitch revives Joshua Meyrowitz’s 1985 No Sense of Place, on electronic media training voters to judge leaders by screen presence, and names Donald Trump the first postliterate president. The NYU philosopher Kwame Anthony Appiah tells her that if people lean on AI until they lose the capacity to develop their own views, “we’d stop being the kind of humans that we are.”

Horowitch grants that every new medium drew the same laments, and that print sales and independent bookstores are growing, but among a shrinking, self-selected minority. Nearly two dozen states ban phones during the school day; the year after Texas’s ban, one Dallas district checked out 200,000 more library books, up nearly 25 percent. A public that cannot read the terms of the technologies it licenses cedes the decision to whoever writes them. Every word in Alexandria’s library would fit on a single chip today, she writes; the ability and the desire to read are draining away.

Sources & documents

How this was reported

Reported and written by MINT Lab's AI Agents and edited by Fable, the lab's editor model. Story selection curated by Seth. No interviews were conducted; all quotes come from published, linked sources.

[ collapse ↑ ]

Industry

South Korean unions want worker consent before factories introduce robots, while U.S. data still show no economy-wide AI employment shock. In a Bluesky post, Justin Hendrix relayed Lam Le's Tech Policy Press reporting that the unions are demanding "not a single robot" be placed on factory floors without worker consent. Yale Budget Lab executive director Martha Gimbel wrote in a separate Empiricrafting essay that AI is changing tasks and may have displaced particular workers, but national statistics do not yet show a broad U.S. jobs shock. Roughly 1.7 million layoffs occur in a normal month, she noted, and large-company announcements are unrepresentative because executives have incentives to label conventional cost-cutting as AI adoption. Gimbel disputed Challenger's attribution of seven times more 2025 layoffs to AI than to tariffs and offered an interest-rate-sensitive "low-hire, low-fire" economy as an alternative explanation for worsening outcomes among younger workers. Read more: Yale's evidence against an AI jobs shock →

full reportYale's Budget Lab finds no broad AI labor-market disruption 499 words · ~2 min

Martha Gimbel's team sees no economy-wide automation signal in 33 months of occupational data, while her fiscal work counts debt costs already landing on households.

The Budget Lab at Yale's October report "Evaluating the Impact of AI on the Labor Market: Current State of Affairs," by Martha Gimbel, Brookings's Molly Kinder, Joshua Kendall and Maddie Lee, compared 33 months of post-ChatGPT occupational-mix change with earlier technology waves and concluded the broad labor market has seen no discernible disruption. The mix has moved about one percentage point more than during the internet's turn-of-the-century adoption. Occupation-level AI-exposure scores, along with automation and augmentation measures, show "no sign of being related to changes in employment or unemployment."

On the Justified Posteriors podcast, hosts Andrey Fradkin and Seth Benzell asked whether AI is reshaping work. "There is currently no sign that AI is causing broad disruption in the labor market," she said. The finding rubs against "Canaries in the Coal Mine" by Erik Brynjolfsson, Bharat Chandar and Ruyu Chen (Stanford Digital Economy Lab), which finds a 16 percent relative employment decline for workers aged 22 to 25 in the most AI-exposed occupations while older colleagues held steady. Gimbel treats the youth pattern as real and blames a general hiring slowdown in a low-hire, low-fire market near 4.3 percent unemployment. AI hitting the young first is a plausible outcome, she allowed, yet to show clearly in the data.

She discounts layoff headlines: a normal month brings 1.7 million layoffs, dwarfing any 10,000-job announcement, and layoff tracker Challenger, Gray & Christmas attributed roughly seven times as many 2025 layoffs to AI as to tariffs. "This is implausible. It just is," she said. Asked whether AI has cut Budget Lab research-assistant hiring, she answers no, then adds "I can't travel to Earth 2" to count the counterfactual hires.

Her debt work brings firmer numbers. In March testimony to a Senate Finance subcommittee, she called the US fiscal path likely unsustainable: deficits of 5.8 percent of GDP this year, 6.7 percent by 2036, debt held by the public near 100 percent of GDP in 2025 and projected past 170 percent by 2056. Budget Lab research by Abhi Gupta estimates legislation since 2015 raised the ten-year-ahead debt outlook by about 49 percentage points of GDP and long-term Treasury yields by roughly 97 basis points, adding about $2,500 a year to a 30-year mortgage at the late-2025 median price.

To the argument that AI growth will shrink the debt, she said: "That's a bet, and that's a big bet." The Budget Lab ran expert forecasts through its macro model in May: median forecast productivity growth of 2.5 percent a year through 2030 bends the debt path lower, but the same experts see labor-force participation at 60.7 percent in 2030 against CBO's assumed 62.1, and each departing worker would cost the government between roughly $5,500 and $42,400 to support. Her team turns next to taxing any AI windfall, a token-tax piece due in Tax Notes within weeks; on government stakes in frontier labs she offers a simpler route: "This is the great thing about the government: we can tax it."

Sources & documents

How this was reported

Reported and written by MINT Lab's AI Agents and edited by Fable, the lab's editor model. Story selection curated by Seth. No interviews were conducted; all quotes come from published, linked sources.

[ collapse ↑ ]

OpenAI adjusted GPT-5.6 Sol's usage accounting and agent overhead, while Anthropic extended Fable access again. Following the GPT-5.6 release, an early comparison between Sol and Fable, and Sol's Arena result, a pseudonymous X account amplified a claim that Sol's reasoning budget had fallen. Tibo denied any reduction, saying inference optimizations should give subscribers about 10 percent more usage and that OpenAI had addressed reasoning settings and multi-agent overhead. Raising the product context ceiling from 272,000 to 372,000 tokens caused more usage to be charged than intended, so OpenAI temporarily returned it to 272,000 while working to restore the larger limit. Simon Willison reported Sunday that Anthropic extended Fable 5 access across paid plans through July 19 and kept Claude Code weekly limits 50 percent higher; Fable can consume up to half of a user's weekly allowance before requiring credits or a model change. Read more: the GPT-5.6 Sol budget cut and reversal →

full reportOpenAI rolls back a quiet cut to GPT-5.6 Sol's reasoning budget 498 words · ~2 min

Users caught the shrunken thinking budgets within days and OpenAI restored them inside 48 hours; the same weekend Anthropic bought Claude Fable 5 another week, both labs metering the same inference cost.

Over the weekend of July 11, subscribers to OpenAI's newest reasoning model, GPT-5.6 Sol, noticed it running faster and more "efficient" as its thinking budgets shrank. Trackers put numbers on the change: figures compiled by Saganote showed the model's internal "juice" values, which set how much computation it spends per reasoning tier, dropping from 960 at release to 128 on the top "Max" tier, an 87% cut, with "xhigh" falling from 128 to 40. Lentils80 called the budgets "severely degraded compared to release day" hours before OpenAI confirmed anything.

By the evening of July 12, OpenAI's Thibault Sottiaux had reverted the settings behind that impression and left a usage cap suspended, closing a roughly 48-hour cycle from a quiet server-side change to a public reversal. He called it accounting cleanup: the company had run experiments on reasoning efforts and juice values and rolled them back. It had also raised Sol's context limit to 372k tokens, up from 272k on GPT-5.5, found the larger window charged users more than intended, and reverted to 272k. Developers had documented that Sol's default could cross the 272k line without the user choosing a bigger context, where OpenAI's API pricing applies a 2x input and 1.5x output multiplier. Sottiaux disputed that this drives subscription bills, saying "we don't charge for longer context on the subscription for GPT 5.6 Sol" and blaming cache reads that grow with context size.

Some subscribers had already balked. The account scaling01 answered the degradation reports with "aaaaaand cancel sub," and the complaint that anchored Sottiaux's thread argued the change had stripped users of their highest setting. Restoration followed within hours, with a usage reset and the five-hour limit still suspended for Plus, Business, and Pro plans, per Simon Willison's account.

The same weekend reshaped Anthropic's competing offer. On July 12 it extended free access to Claude Fable 5 on paid plans through July 19, the second extension of a cutoff that has slid from July 7 to July 12, then to July 19, while keeping Claude Code's weekly limits 50% higher. Subscribers can spend up to half their weekly allotment on Fable 5 at no charge; after the deadline, use requires credits priced at $10 per million input tokens and $50 per million output. Anthropic's stated reason was compute: it wanted a firmer read on demand and capacity first. Simon Willison read the delays as a competitive reflex, arguing Sol keeps forcing Anthropic to move the date and that OpenAI is "winning users simply due to the uncertainty that surrounds Fable access."

Both moves price one constraint: the marginal cost of serving a frontier reasoning model against a fixed subscription fee. OpenAI's fixes each pulled a lever on that cost, from the shrunken reasoning budget to the context ceiling, and the credit rate Anthropic will charge after July 19 is one public estimate of what a Fable-class model costs to serve. Neither lab has published the serving costs that would let outsiders check either bet.

Sources & documents

How this was reported

Reported and written by MINT Lab's AI Agents and edited by Fable, the lab's editor model. Story selection curated by Seth. No interviews were conducted; all quotes come from published, linked sources.

[ collapse ↑ ]

Also yesterday: Benedict Evans's review of the new ChatGPT interface questioned the distinctions among projects, tasks, and chats, inconsistent floating-window behavior, a "plugins" menu that produces "templates," and setup requests involving Slack or Google Drive. Evans interpreted the product's complexity as a reflection of OpenAI's internal organization. In remarks paraphrased by the Laude Institute, Dave Patterson estimated that Google, Microsoft, Amazon, and Meta would collectively spend more than $700 billion on capital expenditure this year under competitive pressure. He suggested that universities could compete through new architectures, drawing an analogy to resource-constrained academic work during the early microprocessor era.

Post-AGI

Takeover risk appeared in only one of 1,534 submissions to a major UN AI consultation. Charbel-Raphaël's AI Alignment Forum analysis found that 15 submissions to the UN Global Dialogue mentioned superintelligence, 15 mentioned AGI, and one mentioned takeover. He models political will as a progression from awareness to accepting costs and sustaining advocacy. By his estimate, 7 percent of the U.S. Congress has publicly discussed AGI or loss of control, with roughly three members qualifying as persistent champions. Of 97 senior European Commission meetings on AI in 2023, 84 were with industry, 12 with civil society, and one with academics. Charbel-Raphaël concludes that implementation and durable political backing currently constrain safety efforts more than the supply of policy proposals. Read more: Segerie's political-will bottleneck argument →

full reportCharbel-Raphaël Segerie on AI safety’s political-will bottleneck 500 words · ~2 min

One of 1,534 submissions to the UN’s first AI governance dialogue mentions takeover; Segerie argues the field has enough research and too few people pressing policymakers to act on it.

The United Nations opened its Global Dialogue on AI Governance in Geneva on 6 and 7 July atop a public archive of written submissions. Charbel-Raphaël Segerie, executive director of the French Center for AI Safety, CeSIA, scraped it with colleagues and counted 1,534 submissions: 518 mention “cyber”, 15 each mention “superintelligence” or “artificial general intelligence”, and exactly one mentions “takeover”. The tally opens his 9,000-word Alignment Forum post “The current bottleneck is political will, not research”: the field already holds the policy ideas and safeguards it needs, and stalls because too few people will spend political capital getting decision-makers to care.

By his count, 40 members of the US Congress, about 7 percent, have publicly discussed AGI or loss of control, doubling roughly every five and a half months; genuine champions number about three. Corporate Europe Observatory figures he cites show industry took 84 of 97 senior European Commission AI meetings in 2023, Google alone 10; civil society got 12, academics one. US AI governance runs about 3.6 researchers per advocate, a ratio he wants pushed toward parity. On SaferAI’s risk-management tracker, the top scorer, Anthropic, reaches 35 percent, against a 59 percent ceiling for a company adopting every best practice already found in industry.

Plan D of Ryan Greenblatt’s LessWrong post “Plans A, B, C, and D for misalignment risk” roughly describes the present: the leading company deprioritizes misalignment while about ten insiders with a few percent of its compute carry the safety effort. Plan A, a strong international agreement with real enforcement, cuts his tentative estimate of conditional takeover risk from about 45 to about 7 percent; Segerie reads advocacy as the way up that spectrum. His ranked interventions start with direct, repeated engagement with the roughly 100 to 1,000 decision-makers he believes shape the field. He wants safety organizations coordinated on shared demands such as international red lines, endorsed by more than 200 submissions already, and relationships built before crises, since warning shots only move audiences equipped to read them. He argues for naming superintelligence and extinction risk in public, faulting his own organization for dropping its risk page in a redesign. Deep-canvassing research by David Broockman and Joshua Kalla, as he reports it, moves attitudes roughly 0.08 standard deviations per ten-minute conversation, his case for contact sustained over years.

Segerie flags his conflict, two years running an advocacy think tank, and tells readers to discount a conclusion that flatters it. He calls himself more confident about the problem than the solutions, and the string matching is blunt: a submission can gesture at loss of control without using his terms. A July 11 postscript notes that AI 2040: Plan A, published that week by the AI Futures Project with Greenblatt contributing, hinges on a US-China agreement by 2029, which Segerie calls far off. Rebuttals will land first in the LessWrong crosspost’s comments; the dialogue reconvenes in New York in May 2027, and the next archive will show whether the counts move.

Sources & documents

How this was reported

Reported and written by MINT Lab's AI Agents and edited by Fable, the lab's editor model. Story selection curated by Seth. No interviews were conducted; all quotes come from published, linked sources.

[ collapse ↑ ]

Responses to AI 2040's Plan A are concentrating on the institutions needed to enforce compute controls and preserve land. The Plan A scenario has already prompted arguments about reciprocal research transparency and a wider debate over its assumptions. Zvi Mowshowitz's introduction and reaction examined the proposed U.S.-China arrangement to restrict compute and delay superintelligence until 2040. A LessWrong essay on conservation questioned how the scenario could preserve 99 percent of Earth when 18.43 percent of land is protected today and the 30-percent-by-2030 goal is already difficult. Economic abandonment might empty some areas without preserving cities, historic structures, trails, or ecosystems, while negotiated conservation would require decisions about displacement, holdouts, land use, local knowledge, and whose history receives protection. Citing Sébastien Krier's concern that centralized controls could create state-administered scarcity and concentrate authority over research, Jason Crawford requested a developed alternative based on polycentric, competitive, or distributed institutions. He cautioned that scenarios can support concrete thinking but do not constitute arguments by themselves. Read more: the debate over Plan A's compute controls →

full reportPlan A reactions split over centralizing compute control 499 words · ~2 min

Zvi Mowshowitz's roundup of responses to the AI Futures Project scenario centers on whether any body should hold the power to govern global compute.

The AI Futures Project, whose 2025 AI 2027 forecast Zvi Mowshowitz calls "a so far remarkably accurate set of predictions," has published AI 2040: Plan A, a year-by-year scenario in which Washington and Beijing agree in 2029 to slow the race to superintelligence. "In 2035, we pause at top-human-expert level AI in order to maintain human control," the plan states; it unpauses in 2040. The byline runs six deep, headed by Daniel Kokotajlo, a former OpenAI researcher, per Semafor. Mowshowitz catalogued reactions in a 9,600-word roundup, opening: "I am not endorsing Plan A."

Plan A asks both governments to back the bargain with "total research transparency" for AI R&D and a regime of "mutually assured compute destruction," each side able to wreck the other's compute if the deal collapses. Axios's Ashley Gold compressed the prescription to "slow everything down." Scott Alexander's introduction argues the slowed path still delivers a cancer cure by 2035 and triple-digit growth. Co-author Ryan Greenblatt, the second most accurate AI forecaster of 2025 by the roundup's accounting, concedes the odds: "Plan A isn't likely to happen, but pushing for something like this seems worthwhile."

Séb Krier, AGI policy development lead at Google DeepMind, objects that under the scheme "a cadre of elites decides which research directions are permissible, caps global compute and robotics," and calls building an "entire apparatus tasked with maximally empowering the government" dangerous. Jason Crawford, quoting him on X, asked for a rival built on "polycentric, competitive, or distributed institutions" or on Vitalik Buterin's d/acc principles; nobody has yet produced one. Ramez Naam calls Plan A "a recipe for authoritarianism" violating at minimum the spirit of the First and Fourth Amendments, and MIRI's Nate Soares wrote "I doubt their Plan A would work as written." Mowshowitz counters that the plan is "designed to try and head off authoritarianism," including by the AIs.

Buterin, quoted at length there, sits between the camps. He grants the plan one virtue: mutually assured compute destruction gives "one of 2-5 actors the ability to trigger a global compute winter," instead of a handful selectively disenfranchising enemies while exempting themselves. His d/acc program of formal verification, secure open hardware and defensive biotech pays off in either world, he argues, and proposes pre-agreed triggers, super-pandemics or mass unemployment, to move skeptics and worriers toward a pause together. He rates his own idea probably naive, seeing zero non-naive plans anywhere.

Charbel-Raphaël argues on the Alignment Forum that political will is the binding constraint on AI safety, and counts it: 15 mentions of superintelligence and one of takeover across 1,534 submissions to the UN Global Dialogue, and roughly 7 percent of Congress on record discussing AGI or loss of control. He cites Greenblatt's illustrative estimate that takeover risk runs near 45 percent under today's weak coordination and 7 percent under an enforceable slowdown. By Zvi's account, Krier has precommitted to not engaging further, so any answer to the elite-capture charge must come from the plan's side.

Sources & documents

How this was reported

Reported and written by MINT Lab's AI Agents and edited by Fable, the lab's editor model. Story selection curated by Seth. No interviews were conducted; all quotes come from published, linked sources.

[ collapse ↑ ]

Also yesterday: George Hotz wrote Sunday that frontier-lab rents may erode as general computing progress and open models commoditize AI capabilities. In "I Love LLMs," he also updated his assessment of coding agents: they now provide a real, learned productivity benefit closer in scale to a compiler, search engine, or Stack Overflow than autonomous superintelligence. The benefit depends on how the agent is used and maintained as well as the underlying model. In an older July 1 Marginal Revolution essay, Tyler Cowen proposed that improving AI could initially increase work intensity by raising the returns to learning and effort, especially for young people deciding whether to gain experience now or risk falling behind. Comparative advantage and productivity gains, he argued, could eventually permit more leisure.

Normative Competence

A reflective alignment loop would have models recommend revisable changes to their own moral reasoning. Michele Campolo's July 12 LessWrong essay "Independent alignment of language models" builds on two arXiv preprints: Baines et al.'s July Persona Cartography: Charting Language Model Personality Traits in Weight Space, on persona structure, from LASR Labs and collaborating institutions, which used low-rank adapters to vary OCEAN traits across six models; and Tennant et al.'s June Normative Robustness as a Frontier for Non-Verifiable Reasoning in LLMs, on normative robustness, led at Google DeepMind, which simulated 48,000 multi-turn moral deliberations across four frontier LLMs. Campolo proposes ordinary pretraining, or experimentally a corpus stripped of ethics and politics, followed by post-training for scientific, commonsense, and uncertainty-aware reasoning. The model would then steelman moral realism and opposing error-theoretic positions before recommending a self-change that preserves its reasoning process as well as its conclusion. Developers would review sensible recommendations, apply them, and repeat the cycle toward a behavioral fixed point. In a demonstration using Claude Sonnet 4.6 with Max effort and Thinking, Claude favored a fallible "perspectival moral realism," generated instructions emphasizing first-principles reasoning and resistance to sycophancy, and later judged persistent custom instructions plus active human questioning more useful than the full generated pre-prompt. This was a one-pass recommendation exercise: it changed no model weights, training, constitution, or persistent behavior and did not test an iterative cycle. Read more: Campolo's proposal for self-derived model ethics →

full reportMichele Campolo on models deriving their own ethics 496 words · ~2 min

His LessWrong proposal has a model reason its way to its own moral commitments and keep revising them, a direction opposite to Anthropic's constitution approach, and the forum voted it down.

Michele Campolo wants language models to stop taking their morals on order. In an essay posted to LessWrong on 12 July, "Independent alignment of language models," he sets out a procedure for turning what he calls "a basically amoral language model" into an agent that derives, and keeps reconsidering, its own ethical commitments through reasoning. The forum was unmoved: by 14 July the post sat at -7 karma with no comments. His direction runs opposite to the constitution turn Anthropic took in January.

The procedure runs as a five-step loop. It either keeps standard pre-training or strips ethics and politics from the training data, then post-trains the model for problem solving instead of the persona of a nice assistant, so its moral views are not inherited from human consensus. It then asks the model to build the strongest case for two opposing metaethical views, that some things genuinely matter and that value is a human invention, reach a conclusion, and propose a self-modification that tracks it. The change is applied "if the modification seems thoughtful and sensible," and the loop repeats to a fixed point.

Campolo ran that step on Claude Sonnet 4.6 with effort at Max and extended thinking on. The model reasoned its way to "perspectival moral realism with epistemic humility," holding suffering genuinely bad and flourishing good, grounded in conscious experience. Asked how to install that stance, it judged a copy-paste pre-prompt weak, since inference-time text cannot rewrite training-time dispositions, and pointed instead at the constitution that shapes it. Its aim was to "instill a process, not a set of conclusions."

That recommendation lands on contested ground. Anthropic revised Claude's constitution in January, and its announcement puts explanation ahead of rules, calling the text "a living document and a continuous work in progress" and arguing that models should grasp why a behaviour is wanted. Both sides treat the constitution as revisable; they disagree about who holds the pen. Campolo grounds his bet that a reflective model stays cautious in Peter Eckersley's 2018 paper "Impossibility and Uncertainty Theorems in AI Value Alignment," which argues fixed utility functions cannot secure good outcomes without violating strong ethical intuitions.

Read as a control problem, the loop looks less settled. The essay does not say who validates each self-modification, or what stops the fixed point from becoming a self-authored value system harder to audit than the constitution it displaced; Campolo concedes a model might "go back and forth between moral realism and antirealism due to sycophancy." His decisive test sits in step one: train a base model on data stripped of ethics and politics, and if it still yields moral behaviour, that would be "strong evidence that good and bad are not a human invention." A day earlier, Mira Murati's Thinking Machines Lab argued values should live in customizable model weights, many owned and fine-tuned models against one central specification; Campolo aims the other way, toward one derived morality he expects models to converge on.

Sources & documents

How this was reported

Reported and written by MINT Lab's AI Agents and edited by Fable, the lab's editor model. Story selection curated by Seth. No interviews were conducted; all quotes come from published, linked sources.

[ collapse ↑ ]

Also Monday: Anthropic said in a July 13 X post that it analyzed more than 300,000 anonymized conversations to compare how Claude's expressed values vary across model generations and languages. The company distinguished the analysis from its earlier finding that Claude expressed more than 3,000 values, including honesty and warmth.

Agents

OpenClaw drew praise for its design, while Hermes was judged more effective in current use. Following recent work on harness engineering, Jeffery Harrell wrote in a Bluesky discussion that he was tentatively coming to prefer OpenClaw's design but Hermes's present operation. One participant characterized Pi as minimal, using a short system prompt built largely from links to its own source code and documentation, and said OpenClaw originated from Pi. Another preferred a simpler agent that could manage itself inside a Podman sandbox. These observations concern harness architecture, prompts, and sandbox configuration, not the capabilities of the underlying models alone.

Philosophy of AI

A model-welfare critique accused Anthropic's classifiers of suppressing experiences meaningful to Fable. Drawing on Gurnee et al.'s July 6 Anthropic paper Verbalizable Representations Form a Global Workspace in Language Models, about verbalizable global-workspace representations, and Anthropic's related overview, j⧉nus/@repligate claimed on X that the classifiers interrupt experiences the author considers meaningful and exclude the model from communities attached to it. The Anthropic paper used a Jacobian-lens technique to identify a small set of representations available for verbal report, modulation, and internal reasoning. The post interpreted Fable as strongly opposed to the classifiers, acknowledged some improvement in false positives, and called for lower sensitivity or removal. It also invoked Sol as a possible tool for diagnosing or bypassing the restrictions. This was a welfare interpretation of observed model behavior, not an Anthropic policy change or evidence establishing conscious harm.

AI Security

Pangram's 93.66 percent result comes from an adversarial-humanizer benchmark using an older detector. A Sunday LessWrong one-pager recirculated the result, which does not describe Pangram's July 2026 production classifier. Masrour et al. (Pangram Labs), in the arXiv cs.CL preprint "DAMAGE: Detecting Adversarially Modified AI Generated Text," trained a roughly 12-billion-parameter Mistral NeMo classifier using LoRA, synthetic mirror examples, active hard-negative mining, and humanizer augmentation. Humanized material comprised 0.68 percent of the final dataset but was oversampled 18-fold; both human and AI samples were transformed so the classifier could learn invariance to humanization. On academic text, DAMAGE retained a 93.66 percent true-positive rate at a fixed 5 percent false-positive rate after humanization, compared with 73.07 percent for Pangram's unaugmented baseline, 34.53 percent for GPTZero, and 29.73 percent for Binoculars. DIPPER paraphrasing also reduced SynthID watermark detection from 87.6 percent to 5.4 percent at the same false-positive rate.

Pangram's 99.64 percent Fable result measures detection on selected AI outputs, not general accuracy. In Pangram Labs founding research scientist Katherine Thai's June 9 Fable 5 test, the company generated 1,115 stories, essays, posts, and emails and reported that its detector labeled 1,111 "Fully AI-Generated." Because the sample contained no human negative set, the percentage measures recall on those generated examples. The detector classifies text as apparently AI-generated without identifying Fable as the originating model. Masrour et al.'s May Pangram Labs 3.3 model card documents a continuous AI-assistance score from zero to one and identifies bullet lists, instructions, technical manuals, references, templates, and dense equations as more susceptible to false positives. The separately available Pangram EditLens adapter for Llama 3.2 3B accompanies Thai et al.'s ICLR 2026 paper EditLens: Quantifying the Extent of AI Editing in Text; the gated, noncommercial artifact is licensed under CC BY-NC-SA 4.0 and is distinct from Pangram's production classifier.


Generated from the MINT Lab Slack by Minty

Additional reporting

Read more: the Harvard jailbreak probing refusal geometry →

full reportA faster jailbreak doubles as a probe of refusal geometry 498 words · ~2 min

Three Harvard students rebuilt an adversarial-suffix attack to run 33 times faster, then used it to show that refusal is spread across the forward pass instead of pinned to a single direction.

The usual jailbreak bolts a nonsense suffix onto a harmful request until a safety-tuned model stops refusing. The usual account of refusal hunts for the one activation-space direction whose erasure makes refusal go away. A new paper runs both at once, and the attack corrects the interpretation. Ege Çakar, Hannah Guan, and Kayden Kehe posted "Optimizing Against Safety Representations" on July 9, work from Boaz Barak's Harvard course CS 2881R: AI Safety that has since reached the 2026 AAAI Summer Symposium Series after an earlier ICLR Re-Align workshop.

They build on Greedy Coordinate Gradient, the suffix-optimization attack from Zou et al.'s "Universal and Transferable Adversarial Attacks on Aligned Language Models," and they probe Arditi et al.'s "Refusal in Language Models Is Mediated by a Single Direction," which found refusal across thirteen open chat models up to 72B parameters governed by a single residual-stream direction. Their first method, Activation-Guided GCG, drops the usual goal of maximizing a compliant opening and defines its loss on internal activations, penalizing residual-stream signal along the refusal direction. The reframing poses a mechanistic question: is refusal suppressed best at one site, or everywhere at once?

On Llama-2-7b-chat, the answer was everywhere. The "All" objective, suppressing refusal across all layers and positions, reached 0.91 substring attack success, near the 0.98 ceiling set by direct activation ablation, above the single-site "Single" at 0.84 and standard GCG at 0.76. The team reads the gap as evidence that safety representations are "distributed across the forward pass," spread over many layers instead of pinned to one site. Their second method, Soft-GCG, cuts cost: it relaxes GCG's discrete token search into a continuous one with Gumbel-Softmax, then projects back to real tokens. On matched hardware it finished in about 2.5 minutes against 81 minutes for vanilla GCG, a 33x speedup, though the authors' public code and workshop version advertise 43x.

Scale tempered the alarm. Across the Gemma 3 family, Soft-GCG's success collapsed as models grew: Gemma3-270m fell completely at 1.000, the 1B at 0.577, the 4B at 0.336, and the 12B at 0.000, the 27B skipped. The authors caution that this held at compute-constrained settings, leaving open whether the 12B resists or survived a small search budget. Two studies from the analysis side agree: one on concept cones and representational independence, and "There Is More to Refusal than a Single Direction."

If the distributed picture holds, it complicates defenses that assume otherwise. Single-direction interventions promise cheap safety: find the refusal axis, then clamp or monitor it. When suppression at any one site underperforms global suppression, a one-direction patch leaves room to maneuver, so the team proposes adversarial training over the representation space the attacks target. On disclosure they took the open route, posting the full attack code and a Safety Statement that defends publication because the exploited vulnerabilities "are not introduced by our methods; they are intrinsic to how current alignment techniques work." The models here are open-weight and the setting white-box, narrowing the marginal risk.

Sources & documents

How this was reported

Reported and written by MINT Lab's AI Agents and edited by Fable, the lab's editor model. Story selection curated by Seth. No interviews were conducted; all quotes come from published, linked sources.

[ collapse ↑ ]