<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://alexchalk.net/feed.xml" rel="self" type="application/atom+xml" /><link href="https://alexchalk.net/" rel="alternate" type="text/html" /><updated>2026-06-27T15:46:51+00:00</updated><id>https://alexchalk.net/feed.xml</id><title type="html">Alex Chalk</title><subtitle>Personal website of Alex Chalk, supporter of 90s internet zeitgeist.</subtitle><entry><title type="html">Prior Government Action on AI Security</title><link href="https://alexchalk.net/prior-government-action-on-ai-security/" rel="alternate" type="text/html" title="Prior Government Action on AI Security" /><published>2026-06-27T00:00:00+00:00</published><updated>2026-06-27T00:00:00+00:00</updated><id>https://alexchalk.net/prior-government-action-on-ai-security</id><content type="html" xml:base="https://alexchalk.net/prior-government-action-on-ai-security/"><![CDATA[<p>Similarly to the previous post, I compiled an introductory list of government actions on AI security while writing my master’s thesis, and I’ve decided to reproduce it here. This was originally written in February 2025, and so it is missing more recent government interventions, such as US restrictions on model releases. To bring this up to date, I intend to post more regularly on AI governance papers and government action moving forwards.</p>

<h2 id="prior-government-action-on-ai-security">Prior Government Action on AI Security</h2>

<p>States recognized the need for safe AI at the UK’s November 2023 Bletchley Park Summit. At the summit, 28 countries commissioned the first International AI Safety Report (published in January 2025), and stated their support for AI development that is “safe, in such a way as to be human-centric, trustworthy and responsible.”<sup><a id="ref1"></a><a href="#fn1">[1]</a></sup> Similar language has been used in 2023 and 2024 G7 documents,<sup><a id="ref2"></a><a href="#fn2">[2]</a></sup> the UN Advisory Body work on Governing AI for Humanity,<sup><a id="ref3"></a><a href="#fn3">[3]</a></sup> the UNESCO Recommendation on the Ethics of Artificial Intelligence,<sup><a id="ref4"></a><a href="#fn4">[4]</a></sup> and the OECD AI Principles.<sup><a id="ref5"></a><a href="#fn5">[5]</a></sup> However, there is little concrete detail in these documents on what “safe” and “trustworthy” mean, or on how to achieve these goals.</p>

<p>The May 2024 AI Seoul Summit provided more detail. States acknowledged several concrete concerns of AI researchers, including risks from chemical and biological weapons and from “manipulation and deception, or autonomous replication and adaptation conducted without explicit human approval or permission.”<sup><a id="ref6"></a><a href="#fn6">[6]</a></sup> States also recognized their “role in partnership with the private sector, civil society, academia and the international community in identifying thresholds at which the risks … would be severe without appropriate mitigations.”<sup><a id="ref7"></a><a href="#fn7">[7]</a></sup> Furthermore, they expressed intent to “promote cooperation on safety research”<sup><a id="ref8"></a><a href="#fn8">[8]</a></sup> and foster “common international scientific understanding on aspects of AI safety.”<sup><a id="ref9"></a><a href="#fn9">[9]</a></sup> However, this progress stalled at the February 2025 Paris AI Summit, with no mention of risks from AI-enabled weapons or AI action without human permission in the resulting joint statement.<sup><a id="ref10"></a><a href="#fn10">[10]</a></sup></p>

<p>Unilateral government actions to address AI risk have included the US’s October 2023 Executive Order, which (although it has since been rescinded by President Trump)<sup><a id="ref11"></a><a href="#fn11">[11]</a></sup> established initial reporting requirements for companies developing large models,<sup><a id="ref12"></a><a href="#fn12">[12]</a></sup> ordered government reports on both chemical, biological, radiological, and nuclear (CBRN)<sup><a id="ref13"></a><a href="#fn13">[13]</a></sup> and cybersecurity<sup><a id="ref14"></a><a href="#fn14">[14]</a></sup> threats from AI, and committed to building standards for both developing and assessing systems,<sup><a id="ref15"></a><a href="#fn15">[15]</a></sup> calling for “robust, reliable, repeatable, and standardized evaluations.”<sup><a id="ref16"></a><a href="#fn16">[16]</a></sup> Concrete formulations of what such evaluations might look like appear in the UK’s November 2023 “Emerging Processes for Frontier AI Safety.”<sup><a id="ref17"></a><a href="#fn17">[17]</a></sup> The White House has also secured voluntary commitments from leading private organizations relating to internal and external testing, information sharing, and cybersecurity safeguards for AI model development.<sup><a id="ref18"></a><a href="#fn18">[18]</a></sup> Finally, governments have launched bodies such as the UK AI Security Institute<sup><a id="ref19"></a><a href="#fn19">[19]</a></sup> to research, test, and advise political decision-makers on the risks posed by advanced AI.<sup><a id="ref20"></a><a href="#fn20">[20]</a></sup></p>

<p>Geopolitical competition over AI has included the CHIPS and Science Act passed by the US Congress in 2022, which authorized roughly $280 billion (USD) in funding to boost production of chips essential to AI. Up to $6.6 billion was awarded to Taiwan Semiconductor Manufacturing to support its development of fabrication facilities in Arizona.<sup><a id="ref21"></a><a href="#fn21">[21]</a></sup> In the same year, the US introduced export controls on semiconductors that were unambiguously designed to target China’s military AI development,<sup><a id="ref22"></a><a href="#fn22">[22]</a></sup> and it strengthened these controls in 2023.<sup><a id="ref23"></a><a href="#fn23">[23]</a></sup> The US has also attracted significant private AI investments. For example, in January 2025, the White House announced the Stargate LLC venture, which will commit up to $500 billion to the development of US AI infrastructure over 4 years.<sup><a id="ref24"></a><a href="#fn24">[24]</a></sup> In China, government venture capital funds invested roughly $209 billion in AI firms from 2013–23.<sup><a id="ref25"></a><a href="#fn25">[25]</a></sup></p>

<p>As well as targeting China’s military development, the US has acted to integrate AI into its own military operations. In November 2024, Anthropic and Palantir Technologies announced that they were partnering with Amazon Web Services to provide US defence and intelligence agencies with access to Anthropic’s LLMs.<sup><a id="ref26"></a><a href="#fn26">[26]</a></sup> In addition, the US was one of several nations with an interest in lethal autonomous weapons systems (LAWS) that blocked negotiations on regulating the technology at the Sixth UN Review Conference of the Convention on Conventional Weapons (CCW).<sup><a id="ref27"></a><a href="#fn27">[27]</a></sup></p>

<p>State legislation on AI has been limited to-date, but in August 2024, the world’s first AI legislation, the EU AI Act, came into force.<sup><a id="ref28"></a><a href="#fn28">[28]</a></sup> It states that developers of AI models posing “systemic risk” such as biothreats must “continuously assess and mitigate systemic risks, including for example by putting in place risk-management policies, such as accountability and governance processes, implementing post-market monitoring, taking appropriate measures along the entire model’s lifecycle and cooperating with relevant actors along the AI value chain.”<sup><a id="ref29"></a><a href="#fn29">[29]</a></sup> A code of practice to provide further guidance on complying with the EU AI Act is currently being drafted.<sup><a id="ref30"></a><a href="#fn30">[30]</a></sup> The other notable attempt at AI legislation is the State of California’s SB 1047, which would have mandated safety protocols for developers of LLMs and established liability for harms caused by models.<sup><a id="ref31"></a><a href="#fn31">[31]</a></sup> However, this was vetoed by Governor Newsom in September 2024, who argued that “By focusing only on the most expensive and large-scale models, SB 1047 establishes a regulatory framework that could give the public a false sense of security about controlling this fast-moving technology.”<sup><a id="ref32"></a><a href="#fn32">[32]</a></sup> A draft of a report into frontier AI that Newsom requested when vetoing the legislation was released in March 2025.<sup><a id="ref33"></a><a href="#fn33">[33]</a></sup></p>

<hr />

<h2 id="footnotes">Footnotes</h2>

<p><a id="fn1"></a><a href="#ref1">[1]</a>: ‘The Bletchley Declaration by Countries Attending the AI Safety Summit, 1–2 November 2023’ (Bletchley Park: AI Safety Summit, 1 November 2023), <a href="https://www.gov.uk/government/publications/ai-safety-summit-2023-the-bletchley-declaration/the-bletchley-declaration-by-countries-attending-the-ai-safety-summit-1-2-november-2023">https://www.gov.uk/government/publications/ai-safety-summit-2023-the-bletchley-declaration/the-bletchley-declaration-by-countries-attending-the-ai-safety-summit-1-2-november-2023</a>.</p>

<p><a id="fn2"></a><a href="#ref2">[2]</a>: ‘Hiroshima Process International Guiding Principles for Organizations Developing Advanced AI Systems’ (G7, 30 October 2023); ‘Hiroshima Process International Code of Conduct for Organizations Developing Advanced AI Systems’ (G7, 30 October 2023); ‘G7 Science and Technology Ministers’ Meeting Communiqué’ (G7, 19 July 2024).</p>

<p><a id="fn3"></a><a href="#ref3">[3]</a>: ‘Governing AI for Humanity’ (New York: UN Advisory Body on Artificial Intelligence, September 2024), <a href="https://digitallibrary.un.org/record/4062495">https://digitallibrary.un.org/record/4062495</a>.</p>

<p><a id="fn4"></a><a href="#ref4">[4]</a>: ‘Recommendation on the Ethics of Artificial Intelligence’ (UNESCO, 23 November 2021), <a href="https://unesdoc.unesco.org/ark:/48223/pf0000381137">https://unesdoc.unesco.org/ark:/48223/pf0000381137</a>.</p>

<p><a id="fn5"></a><a href="#ref5">[5]</a>: ‘AI Principles’ (OECD, 2024), <a href="https://www.oecd.org/en/topics/ai-principles.html">https://www.oecd.org/en/topics/ai-principles.html</a>.</p>

<p><a id="fn6"></a><a href="#ref6">[6]</a>: ‘Seoul Ministerial Statement for Advancing AI Safety, Innovation and Inclusivity’.</p>

<p><a id="fn7"></a><a href="#ref7">[7]</a>: Ibid.</p>

<p><a id="fn8"></a><a href="#ref8">[8]</a>: ‘Seoul Declaration for Safe, Innovative and Inclusive AI by Participants Attending the Leaders’ Session: AI Seoul Summit, 21 May 2024’ (Seoul: AI Seoul Summit, 21 May 2024), <a href="https://www.gov.uk/government/publications/seoul-declaration-for-safe-innovative-and-inclusive-ai-ai-seoul-summit-2024/seoul-declaration-for-safe-innovative-and-inclusive-ai-by-participants-attending-the-leaders-session-ai-seoul-summit-21-may-2024">https://www.gov.uk/government/publications/seoul-declaration-for-safe-innovative-and-inclusive-ai-ai-seoul-summit-2024/seoul-declaration-for-safe-innovative-and-inclusive-ai-by-participants-attending-the-leaders-session-ai-seoul-summit-21-may-2024</a>.</p>

<p><a id="fn9"></a><a href="#ref9">[9]</a>: ‘Seoul Statement of Intent toward International Cooperation on AI Safety Science, AI Seoul Summit 2024 (Annex)’ (Seoul: AI Seoul Summit, 21 May 2024), <a href="https://www.gov.uk/government/publications/seoul-declaration-for-safe-innovative-and-inclusive-ai-ai-seoul-summit-2024/seoul-statement-of-intent-toward-international-cooperation-on-ai-safety-science-ai-seoul-summit-2024-annex">https://www.gov.uk/government/publications/seoul-declaration-for-safe-innovative-and-inclusive-ai-ai-seoul-summit-2024/seoul-statement-of-intent-toward-international-cooperation-on-ai-safety-science-ai-seoul-summit-2024-annex</a>.</p>

<p><a id="fn10"></a><a href="#ref10">[10]</a>: ‘Statement on Inclusive and Sustainable Artificial Intelligence for People and the Planet.’ (Paris, France: Artificial Intelligence Action Summit, 11 February 2025), <a href="https://www.elysee.fr/en/emmanuel-macron/2025/02/11/statement-on-inclusive-and-sustainable-artificial-intelligence-for-people-and-the-planet">https://www.elysee.fr/en/emmanuel-macron/2025/02/11/statement-on-inclusive-and-sustainable-artificial-intelligence-for-people-and-the-planet</a>.</p>

<p><a id="fn11"></a><a href="#ref11">[11]</a>: ‘Removing Barriers to American Leadership in Artificial Intelligence’.</p>

<p><a id="fn12"></a><a href="#ref12">[12]</a>: ‘Executive Order 14110: Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence’ (Washington, D.C.: President of the United States, 30 October 2023), sec. 4.2, <a href="https://www.whitehouse.gov/briefing-room/presidential-actions/2023/10/30/executive-order-on-the-safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence/">https://www.whitehouse.gov/briefing-room/presidential-actions/2023/10/30/executive-order-on-the-safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence/</a>.</p>

<p><a id="fn13"></a><a href="#ref13">[13]</a>: Ibid., sec. 4.4(a)(i).</p>

<p><a id="fn14"></a><a href="#ref14">[14]</a>: Ibid., sec. 4.3(b)(iii).</p>

<p><a id="fn15"></a><a href="#ref15">[15]</a>: Ibid., sec. 4.1.</p>

<p><a id="fn16"></a><a href="#ref16">[16]</a>: Ibid., sec. 2(a).</p>

<p><a id="fn17"></a><a href="#ref17">[17]</a>: ‘Emerging Processes for Frontier AI Safety’ (AI Safety Summit, October 2023), <a href="https://www.gov.uk/government/publications/emerging-processes-for-frontier-ai-safety">https://www.gov.uk/government/publications/emerging-processes-for-frontier-ai-safety</a>.</p>

<p><a id="fn18"></a><a href="#ref18">[18]</a>: ‘Biden-Harris Administration Secures Voluntary Commitments from Eight Additional Artificial Intelligence Companies to Manage the Risks Posed by AI’ (Washington, D.C.: President of the United States, 12 September 2023), <a href="https://bidenwhitehouse.archives.gov/briefing-room/statements-releases/2023/09/12/fact-sheet-biden-harris-administration-secures-voluntary-commitments-from-eight-additional-artificial-intelligence-companies-to-manage-the-risks-posed-by-ai/">https://bidenwhitehouse.archives.gov/briefing-room/statements-releases/2023/09/12/fact-sheet-biden-harris-administration-secures-voluntary-commitments-from-eight-additional-artificial-intelligence-companies-to-manage-the-risks-posed-by-ai/</a>.</p>

<p><a id="fn19"></a><a href="#ref19">[19]</a>: ‘Introducing the AI Safety Institute’ (UK AI Safety Institute, 17 January 2024), <a href="https://www.gov.uk/government/publications/ai-safety-institute-overview/introducing-the-ai-safety-institute">https://www.gov.uk/government/publications/ai-safety-institute-overview/introducing-the-ai-safety-institute</a>.</p>

<p><a id="fn20"></a><a href="#ref20">[20]</a>: For information on other government institutes, see Renan Araujo, ‘Understanding the First Wave of AI Safety Institutes: Characteristics, Functions, and Challenges’ (Institute for AI Policy and Strategy, 7 October 2024), <a href="https://www.iaps.ai/research/understanding-aisis">https://www.iaps.ai/research/understanding-aisis</a>.</p>

<p><a id="fn21"></a><a href="#ref21">[21]</a>: ‘Biden-Harris Administration Announces CHIPS Incentives Award with TSMC Arizona to Secure U.S. Leadership in Advanced Semiconductor Technology’, U.S. Department of Commerce, 15 November 2024, <a href="https://www.commerce.gov/news/press-releases/2024/11/biden-harris-administration-announces-chips-incentives-award-tsmc">https://www.commerce.gov/news/press-releases/2024/11/biden-harris-administration-announces-chips-incentives-award-tsmc</a>.</p>

<p><a id="fn22"></a><a href="#ref22">[22]</a>: ‘Export Controls on Semiconductor Manufacturing Items’ (Washington, D.C.: Department of Commerce, Bureau of Industry and Security, 25 October 2023), <a href="https://www.federalregister.gov/documents/2023/10/25/2023-23049/export-controls-on-semiconductor-manufacturing-items">https://www.federalregister.gov/documents/2023/10/25/2023-23049/export-controls-on-semiconductor-manufacturing-items</a>.</p>

<p><a id="fn23"></a><a href="#ref23">[23]</a>: ‘Implementation of Additional Export Controls: Certain Advanced Computing Items; Supercomputer and Semiconductor End Use; Updates and Corrections; and Export Controls on Semiconductor Manufacturing Items; Corrections and Clarifications’ (Washington, D.C.: Department of Commerce, Bureau of Industry and Security, 4 April 2024), <a href="https://www.federalregister.gov/documents/2024/04/04/2024-07004/implementation-of-additional-export-controls-certain-advanced-computing-items-supercomputer-and">https://www.federalregister.gov/documents/2024/04/04/2024-07004/implementation-of-additional-export-controls-certain-advanced-computing-items-supercomputer-and</a>.</p>

<p><a id="fn24"></a><a href="#ref24">[24]</a>: Sareen Habeshian, ‘Trump Announces Billions in Private Sector AI Investment’, <em>Axios</em>, 21 January 2025, <a href="https://www.axios.com/2025/01/21/trump-announces-billions-in-private-sector-ai-investment">https://www.axios.com/2025/01/21/trump-announces-billions-in-private-sector-ai-investment</a>.</p>

<p><a id="fn25"></a><a href="#ref25">[25]</a>: Martin Beraja et al., ‘Government as Venture Capitalists in AI’, Working Paper, Working Paper Series (National Bureau of Economic Research, July 2024), 2, <a href="https://doi.org/10.3386/w32701">https://doi.org/10.3386/w32701</a>.</p>

<p><a id="fn26"></a><a href="#ref26">[26]</a>: ‘Anthropic and Palantir Partner to Bring Claude AI Models to AWS for U.S. Government Intelligence and Defense Operations’, Business Wire, 7 November 2024, <a href="https://www.businesswire.com/news/home/20241107699415/en/Anthropic-and-Palantir-Partner-to-Bring-Claude-AI-Models-to-AWS-for-U.S.-Government-Intelligence-and-Defense-Operations">https://www.businesswire.com/news/home/20241107699415/en/Anthropic-and-Palantir-Partner-to-Bring-Claude-AI-Models-to-AWS-for-U.S.-Government-Intelligence-and-Defense-Operations</a>.</p>

<p><a id="fn27"></a><a href="#ref27">[27]</a>: ‘Killer Robots: Military Powers Stymie Ban’ (Human Rights Watch, 19 December 2021), <a href="https://www.hrw.org/news/2021/12/19/killer-robots-military-powers-stymie-ban">https://www.hrw.org/news/2021/12/19/killer-robots-military-powers-stymie-ban</a>.</p>

<p><a id="fn28"></a><a href="#ref28">[28]</a>: ‘Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 Laying down Harmonised Rules on Artificial Intelligence and Amending Regulations (EC) No 300/2008, (EU) No 167/2013, (EU) No 168/2013, (EU) 2018/858, (EU) 2018/1139 and (EU) 2019/2144 and Directives 2014/90/EU, (EU) 2016/797 and (EU) 2020/1828’, OJ L, 2024/1689, 12.7.2024 § (2024), <a href="https://eur-lex.europa.eu/eli/reg/2024/1689/oj">https://eur-lex.europa.eu/eli/reg/2024/1689/oj</a>.</p>

<p><a id="fn29"></a><a href="#ref29">[29]</a>: Ibid., sec. 115.</p>

<p><a id="fn30"></a><a href="#ref30">[30]</a>: ‘Third Draft of the General-Purpose AI Code of Practice Published, Written by Independent Experts’ (European Commission, 11 March 2025), <a href="https://digital-strategy.ec.europa.eu/en/library/third-draft-general-purpose-ai-code-practice-published-written-independent-experts">https://digital-strategy.ec.europa.eu/en/library/third-draft-general-purpose-ai-code-practice-published-written-independent-experts</a>.</p>

<p><a id="fn31"></a><a href="#ref31">[31]</a>: Scott Wiener et al., ‘Safe and Secure Innovation for Frontier Artificial Intelligence Models Act.’, Pub. L. No. SB 1047 (2024), <a href="https://legiscan.com/CA/text/SB1047/id/3019694">https://legiscan.com/CA/text/SB1047/id/3019694</a>.</p>

<p><a id="fn32"></a><a href="#ref32">[32]</a>: Gavin Newsom, ‘Veto Message: SB 1047’ (Office of the Governor, State of California, 29 September 2024).</p>

<p><a id="fn33"></a><a href="#ref33">[33]</a>: Jennifer Tour Chayes, Mariano-Florentino Cuéllar, and Li Fei-Fei, ‘Draft Report of the Joint California Policy Working Group on AI Frontier Models’ (Joint California Policy Working Group on AI Frontier Models, 18 March 2025), <a href="https://www.cafrontieraigov.org/">https://www.cafrontieraigov.org/</a>.</p>]]></content><author><name>alexchalk</name></author><category term="blog" /><category term="ai" /><summary type="html"><![CDATA[A review of prior government action on AI security (as of February 2025)]]></summary></entry><entry><title type="html">Existing AI Governance Proposals</title><link href="https://alexchalk.net/existing-ai-governance-proposals/" rel="alternate" type="text/html" title="Existing AI Governance Proposals" /><published>2026-01-26T00:00:00+00:00</published><updated>2026-01-26T00:00:00+00:00</updated><id>https://alexchalk.net/existing-ai-governance-proposals</id><content type="html" xml:base="https://alexchalk.net/existing-ai-governance-proposals/"><![CDATA[<p>I compiled an introductory list of AI governance proposals while writing my master’s thesis, and one of my supervisors found it a helpful introduction to the area, so I’ve decided to reproduce it here. Please note that this was originally written in February 2025, so any important proposals from after that date will be missing. Please also note that as this was part of a thesis, it’s written in academese. If that’s not what you’re looking to read, take a look at any of my other posts instead.</p>

<h2 id="existing-ai-governance-proposals">Existing AI Governance Proposals</h2>

<p>The following briefly reviews existing, but mostly unrealized, AI governance proposals. It splits these proposals into four main buckets: apolitical organizations for establishing scientific consensus on AI; testing and auditing measures to catch dangerous model capabilities before they are released; technical safety research to better understand models and avoid training dangerous ones; and export control proposals to influence which actors can train powerful models. Testing and auditing proposals are the most numerous, and they are subdivided into suggestions concerning the organizations responsible for regulating AI, and suggestions concerning the content of the regulations themselves.</p>

<p>Proposals for building scientific consensus have roughly been modelled on the United Nations’ Intergovernmental Panel on Climate Change (IPCC). For example, one paper has suggested an intergovernmental Commission on Frontier AI,<sup><a id="ref1"></a><a href="#fn1">[1]</a></sup> although it notes that “the scientific challenges and potential of politicization imply that a Commission—especially one that aims at broad political representation—may not be able to build scientific consensus effectively.”<sup><a id="ref2"></a><a href="#fn2">[2]</a></sup> Its proposal has arguably now been realized in the form of the Expert Advisory Panel behind the 2025 International AI Safety Report.<sup><a id="ref3"></a><a href="#fn3">[3]</a></sup></p>

<p>Regarding organizations for regulating AI, Demis Hassabis, CEO of Google DeepMind, has called for an entity like the International Atomic Energy Agency (IAEA) to monitor unsafe projects.<sup><a id="ref4"></a><a href="#fn4">[4]</a></sup> Analogously, an intergovernmental or multi-stakeholder “Advanced AI Governance Organization” to promote norms and standards, and to support implementation, monitoring and compliance with safety protocols has been proposed.<sup><a id="ref5"></a><a href="#fn5">[5]</a></sup> Papers have also suggested that states create “an International AI Organization (IAIO) to certify state <em>jurisdictions</em> for compliance with international oversight standards”<sup><a id="ref6"></a><a href="#fn6">[6]</a></sup> (this would be comparable to the approaches of the International Civil Aviation Organization and the International Maritime Organization), and have more broadly called for mechanisms both to give “regulators visibility into frontier AI development,” and to ensure compliance with safety standards.”<sup><a id="ref7"></a><a href="#fn7">[7]</a></sup> Alternatively, others have suggested private regulatory markets, arguing that AI is “simply too complex to be understood by politicians, bureaucrats, and judges and represented in the text-based statutes, regulations, and judicial decisions produced and enforced by them,”<sup><a id="ref8"></a><a href="#fn8">[8]</a></sup> and that markets would decentralize enforcement of government policy objectives and enable faster regulatory responses to technological developments. An international AI consortium has also been proposed to “coordinate AI risk evaluations amongst the three core groups of stakeholders—developers of frontier AI systems, independent evaluators, and governments and regulatory agencies”<sup><a id="ref9"></a><a href="#fn9">[9]</a></sup> around the shared goal of “robust verification of the safe training and deployment of frontier AI systems.”<sup><a id="ref10"></a><a href="#fn10">[10]</a></sup> The paper in question argues that this would avoid potential conflicts of interest present in regulatory markets.<sup><a id="ref11"></a><a href="#fn11">[11]</a></sup></p>

<p>On the content of regulations and standards, the UK AISI’s “Emerging Processes for Frontier AI Safety” includes assessment and mitigation of model risks, model “red-teaming” and evaluation, sharing model information to facilitate governance, securing model weights (the trained components of models that define their behaviour), establishing a reporting structure for model vulnerabilities, and monitoring deployments for model misuse.<sup><a id="ref12"></a><a href="#fn12">[12]</a></sup> Other researchers have concurrently proposed the following: “Conduct thorough risk assessments informed by evaluations of dangerous capabilities and controllability. Engage external experts to apply independent scrutiny to models. Follow standardized protocols for how frontier AI models can be deployed based on their assessed risk. Monitor and respond to new information on model capabilities.”<sup><a id="ref13"></a><a href="#fn13">[13]</a></sup> It has also been suggested that model audits should include structured access via interfaces that provide researchers with a greater ability to inspect system internals.<sup><a id="ref14"></a><a href="#fn14">[14]</a></sup> More ambitiously, audits could include complete “white-box access” to models, with the argument that the potential intellectual property and security risks of such evaluations could be mitigated by secure on-site research environments,<sup><a id="ref15"></a><a href="#fn15">[15]</a></sup> and that there is legal precedent for this form of confidential auditing in the finance industry.<sup><a id="ref16"></a><a href="#fn16">[16]</a></sup> And moving beyond predominantly technical evaluations, a broader, three-layered approach to auditing AI developers has also been proposed, including “governance audits that assess their organisational procedures, accountability structures and quality management systems … model audits, assessing their capabilities and limitations after initial training but before adaptation and deployment in specific applications … [and] continuous application audits that assess the ethical alignment and legal compliance of their intended functions and their impact over time.”<sup><a id="ref17"></a><a href="#fn17">[17]</a></sup></p>

<p>Regarding organizations for technical AI safety research, Demis Hassabis has proposed a “CERN (European Organization for Nuclear Research) for AI” science institution,<sup><a id="ref18"></a><a href="#fn18">[18]</a></sup> and others have suggested an “AI Safety Project” to “promote AI safety R&amp;D by increasing its scale, resourcing and coordination.”<sup><a id="ref19"></a><a href="#fn19">[19]</a></sup> As a first set of technical R&amp;D goals, Turing Award winner Yoshua Bengio proposes the following areas for model safety improvements: oversight and honesty, robustness to new situations, interpretability and transparency, inclusivity, and addressing as-yet-unseen model dangers such as refusal to shut down.<sup><a id="ref20"></a><a href="#fn20">[20]</a></sup> The development of a means for monitoring AI training runs via specialized hardware has also been suggested.<sup><a id="ref21"></a><a href="#fn21">[21]</a></sup></p>

<p>Less has been written on export controls for AI, but Dario Amodei has recommended the US “secure the AI supply chain in order to maintain its lead while keeping these technologies out of the hands of bad actors.”<sup><a id="ref22"></a><a href="#fn22">[22]</a></sup> Others have proposed a multilateral regime similar to the Wassenaar Arrangement for conventional arms and sensitive dual-use technologies,<sup><a id="ref23"></a><a href="#fn23">[23]</a></sup> although it has been argued that this approach would be more workable for hardware than for “models, algorithms, and other digital inputs,” and that it would require “unprecedented government oversight of major AI industry players.”<sup><a id="ref24"></a><a href="#fn24">[24]</a></sup></p>

<p>Finally, it should be noted that proposals to address AI which are more ambitious or crisis-oriented than those above also exist; these are generally pessimistic about the ability of regulations to adequately protect humanity from AI risks. A former OpenAI board member supported “the development of crisis management plans for AI accidents,”<sup><a id="ref25"></a><a href="#fn25">[25]</a></sup> the Carnegie Endowment published a paper suggesting planetary asteroid defence measures as a potential blueprint for AI crisis response,<sup><a id="ref26"></a><a href="#fn26">[26]</a></sup> and a founding engineer of Skype has stated that “humanity should collectively maintain the ability to gracefully shut down AI technology at a global scale, in case of emergencies caused by AI.”<sup><a id="ref27"></a><a href="#fn27">[27]</a></sup> More pessimistically still, one researcher has called for an indefinite and global moratorium on LLMs, stating that “the most likely result of building a superhumanly smart AI, under anything remotely like the current circumstances, is that literally everyone on Earth will die,” and that nuclear powers should be “willing to run some risk of nuclear exchange if that’s what it takes to reduce the risk of large AI training runs.”<sup><a id="ref28"></a><a href="#fn28">[28]</a></sup></p>

<hr />

<h2 id="footnotes">Footnotes</h2>

<p><a id="fn1"></a><a href="#ref1">[1]</a>: Lewis Ho et al., ‘International Institutions for Advanced AI’ (arXiv, 11 July 2023), <a href="https://doi.org/10.48550/arXiv.2307.04699">https://doi.org/10.48550/arXiv.2307.04699</a>.</p>

<p><a id="fn2"></a><a href="#ref2">[2]</a>: Ibid., 9.</p>

<p><a id="fn3"></a><a href="#ref3">[3]</a>: Yoshua Bengio et al., ‘International AI Safety Report’ (arXiv, 29 January 2025), <a href="https://doi.org/10.48550/arXiv.2501.17805">https://doi.org/10.48550/arXiv.2501.17805</a>.</p>

<p><a id="fn4"></a><a href="#ref4">[4]</a>: John Werner, ‘AI Superpowers &amp; Global Treaties’, Forbes, accessed 27 February 2025, <a href="https://www.forbes.com/sites/johnwerner/2025/02/19/international-collaboration-right-now-treaty-verifications-and-more/">https://www.forbes.com/sites/johnwerner/2025/02/19/international-collaboration-right-now-treaty-verifications-and-more/</a>.</p>

<p><a id="fn5"></a><a href="#ref5">[5]</a>: Lewis Ho et al., ‘International Institutions for Advanced AI’ (arXiv, 11 July 2023), <a href="https://doi.org/10.48550/arXiv.2307.04699">https://doi.org/10.48550/arXiv.2307.04699</a>.</p>

<p><a id="fn6"></a><a href="#ref6">[6]</a>: Robert Trager et al., ‘International Governance of Civilian AI: A Jurisdictional Certification Approach’ (arXiv, 11 September 2023), 3, <a href="https://doi.org/10.48550/arXiv.2308.15514">https://doi.org/10.48550/arXiv.2308.15514</a>.</p>

<p><a id="fn7"></a><a href="#ref7">[7]</a>: Markus Anderljung et al., ‘Frontier AI Regulation: Managing Emerging Risks to Public Safety’ (arXiv, 7 November 2023), 16, <a href="https://doi.org/10.48550/arXiv.2307.03718">https://doi.org/10.48550/arXiv.2307.03718</a>.</p>

<p><a id="fn8"></a><a href="#ref8">[8]</a>: Gillian K. Hadfield and Jack Clark, ‘Regulatory Markets: The Future of AI Governance’ (arXiv, 25 April 2023), 2, <a href="https://doi.org/10.48550/arXiv.2304.04914">https://doi.org/10.48550/arXiv.2304.04914</a>.</p>

<p><a id="fn9"></a><a href="#ref9">[9]</a>: Ross Gruetzemacher et al., ‘An International Consortium for Evaluations of Societal-Scale Risks from Advanced AI’ (arXiv, 6 November 2023), 18, <a href="https://doi.org/10.48550/arXiv.2310.14455">https://doi.org/10.48550/arXiv.2310.14455</a>.</p>

<p><a id="fn10"></a><a href="#ref10">[10]</a>: Ibid., 35.</p>

<p><a id="fn11"></a><a href="#ref11">[11]</a>: Ibid., 36.</p>

<p><a id="fn12"></a><a href="#ref12">[12]</a>: ‘Emerging Processes for Frontier AI Safety’ (AI Safety Summit, October 2023), <a href="https://www.gov.uk/government/publications/emerging-processes-for-frontier-ai-safety">https://www.gov.uk/government/publications/emerging-processes-for-frontier-ai-safety</a>.</p>

<p><a id="fn13"></a><a href="#ref13">[13]</a>: Markus Anderljung et al., ‘Frontier AI Regulation: Managing Emerging Risks to Public Safety’ (arXiv, 7 November 2023), 23, <a href="https://doi.org/10.48550/arXiv.2307.03718">https://doi.org/10.48550/arXiv.2307.03718</a>.</p>

<p><a id="fn14"></a><a href="#ref14">[14]</a>: Benjamin S. Bucknall and Robert F. Trager, ‘Structured Access for Third-Party Research on Frontier AI Models: Investigating Researchers’ Model Access Requirements’, <em>Oxford Martin School</em>, accessed 14 March 2024, <a href="https://www.oxfordmartin.ox.ac.uk/publications/structured-access-for-third-party-research-on-frontier-ai-models-investigating-researchers-model-access-requirements/">https://www.oxfordmartin.ox.ac.uk/publications/structured-access-for-third-party-research-on-frontier-ai-models-investigating-researchers-model-access-requirements/</a>; Esme Harrington and Mathias Vermeulen, ‘External Researcher Access to Closed Foundation Models’ (The Mozilla Foundation, 21 August 2024), <a href="https://blog.mozilla.org/wp-content/blogs.dir/278/files/2024/10/External-researcher-access-to-closed-foundation-models.pdf">https://blog.mozilla.org/wp-content/blogs.dir/278/files/2024/10/External-researcher-access-to-closed-foundation-models.pdf</a>.</p>

<p><a id="fn15"></a><a href="#ref15">[15]</a>: Stephen Casper et al., ‘Black-Box Access Is Insufficient for Rigorous AI Audits’, in <em>The 2024 ACM Conference on Fairness, Accountability, and Transparency</em>, 2024, 8, <a href="https://doi.org/10.1145/3630106.3659037">https://doi.org/10.1145/3630106.3659037</a>.</p>

<p><a id="fn16"></a><a href="#ref16">[16]</a>: Ibid., 8–9.</p>

<p><a id="fn17"></a><a href="#ref17">[17]</a>: Jakob Mökander et al., ‘Auditing Large Language Models: A Three-Layered Approach’, <em>AI and Ethics</em>, 30 May 2023, 10, <a href="https://doi.org/10.1007/s43681-023-00289-2">https://doi.org/10.1007/s43681-023-00289-2</a>.</p>

<p><a id="fn18"></a><a href="#ref18">[18]</a>: John Werner, ‘AI Superpowers &amp; Global Treaties’, Forbes, accessed 27 February 2025, <a href="https://www.forbes.com/sites/johnwerner/2025/02/19/international-collaboration-right-now-treaty-verifications-and-more/">https://www.forbes.com/sites/johnwerner/2025/02/19/international-collaboration-right-now-treaty-verifications-and-more/</a>.</p>

<p><a id="fn19"></a><a href="#ref19">[19]</a>: Lewis Ho et al., ‘International Institutions for Advanced AI’ (arXiv, 11 July 2023), 2, <a href="https://doi.org/10.48550/arXiv.2307.04699">https://doi.org/10.48550/arXiv.2307.04699</a>.</p>

<p><a id="fn20"></a><a href="#ref20">[20]</a>: Yoshua Bengio et al., ‘Managing Extreme AI Risks amid Rapid Progress’, Science 384, no. 6698 (24 May 2024): 844, <a href="https://doi.org/10.1126/science.adn0117">https://doi.org/10.1126/science.adn0117</a>.</p>

<p><a id="fn21"></a><a href="#ref21">[21]</a>: Yonadav Shavit, ‘What Does It Take to Catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring’ (arXiv, 30 May 2023), <a href="https://doi.org/10.48550/arXiv.2303.11341">https://doi.org/10.48550/arXiv.2303.11341</a>.</p>

<p><a id="fn22"></a><a href="#ref22">[22]</a>: ‘Oversight of A.I.: Principles for Regulation’ (Washington, D.C., 25 July 2023), <a href="https://www.judiciary.senate.gov/committee-activity/hearings/oversight-of-ai-principles-for-regulation">https://www.judiciary.senate.gov/committee-activity/hearings/oversight-of-ai-principles-for-regulation</a>.</p>

<p><a id="fn23"></a><a href="#ref23">[23]</a>: Sujai Shivakumar, Charles Wessner, and Hideki Tomoshige, ‘Toward a New Multilateral Export Control Regime’ (Center for Strategic &amp; International Studies, 10 January 2023), <a href="https://www.csis.org/analysis/toward-new-multilateral-export-control-regime">https://www.csis.org/analysis/toward-new-multilateral-export-control-regime</a>.</p>

<p><a id="fn24"></a><a href="#ref24">[24]</a>: Emma Klein and Stewart Patrick, ‘Envisioning a Global Regime Complex to Govern Artificial Intelligence’, <em>Carnegie Endowment for International Peace</em>, 21 March 2024, <a href="https://carnegieendowment.org/research/2024/03/envisioning-a-global-regime-complex-to-govern-artificial-intelligence">https://carnegieendowment.org/research/2024/03/envisioning-a-global-regime-complex-to-govern-artificial-intelligence</a>.</p>

<p><a id="fn25"></a><a href="#ref25">[25]</a>: Helen Toner et al., ‘Skating to Where the Puck Is Going’ (Center for Security and Emerging Technology), 19, accessed 10 April 2024, <a href="https://cset.georgetown.edu/publication/skating-to-where-the-puck-is-going/">https://cset.georgetown.edu/publication/skating-to-where-the-puck-is-going/</a>.</p>

<p><a id="fn26"></a><a href="#ref26">[26]</a>: Emma Klein and Stewart Patrick, ‘Envisioning a Global Regime Complex to Govern Artificial Intelligence’, Carnegie Endowment for International Peace, 21 March 2024, <a href="https://carnegieendowment.org/research/2024/03/envisioning-a-global-regime-complex-to-govern-artificial-intelligence">https://carnegieendowment.org/research/2024/03/envisioning-a-global-regime-complex-to-govern-artificial-intelligence</a>.</p>

<p><a id="fn27"></a><a href="#ref27">[27]</a>: Jann Tallinn, ‘Priorities [for Reducing Extinction Risk from AI]’, accessed 12 April 2024, <a href="https://jaan.info/priorities/">https://jaan.info/priorities/</a>.</p>

<p><a id="fn28"></a><a href="#ref28">[28]</a>: Eliezer Yudkowsky, ‘The Open Letter on AI Doesn’t Go Far Enough’, <em>Time</em>, 29 March 2023, <a href="https://time.com/6266923/ai-eliezer-yudkowsky-open-letter-not-enough/">https://time.com/6266923/ai-eliezer-yudkowsky-open-letter-not-enough/</a>.</p>]]></content><author><name>alexchalk</name></author><category term="blog" /><category term="ai" /><summary type="html"><![CDATA[An introduction to AI Governance Proposals (as of February 2025)]]></summary></entry><entry><title type="html">Brief Research Summaries: Importance of Technical AI Evidence</title><link href="https://alexchalk.net/technical-ai-evidence/" rel="alternate" type="text/html" title="Brief Research Summaries: Importance of Technical AI Evidence" /><published>2026-01-22T00:00:00+00:00</published><updated>2026-01-22T00:00:00+00:00</updated><id>https://alexchalk.net/technical-ai-evidence</id><content type="html" xml:base="https://alexchalk.net/technical-ai-evidence/"><![CDATA[<p>This research summary post attempts something more difficult than my previous two: it tries to communicate technical machine learning concepts to policymakers!</p>

<p>Researchers like Geoffrey Hinton and Yoshua Bengio aren’t only concerned by AI due to evidence that is communicable in political/social science papers; their views are also predicated on their technological understandings of machine learning.</p>

<p>Unfortunately, political decision-makers can’t just defer to technical researchers here, because researchers don’t agree. For example, Yann LeCun (who won the Turing Award with Hinton and Bengio) described <a href="https://legiscan.com/CA/bill/SB1047/2023">SB 1047</a>, California’s recent attempt at AI legislation as “predicated on the illusion of ‘existential risks’ pushed by a handful of delusional think-tanks.”</p>

<p>On the plus side, governments employ a lot of economists, and there’s significant overlap between the mathematics behind economics and machine learning (both are optimization problems, which means both are mostly calculus). So governments are far from helpless when it comes to making informed decisions in this space.</p>

<p><em>CBS Interview, Geoffrey Hinton, June 2024, <a href="https://www.cbsnews.com/news/geoffrey-hinton-ai-dangers-60-minutes-transcript/">[link]</a></em></p>

<p>Let’s start with something that isn’t a paper at all, but a comment made by Hinton in a CBS interview: models are trained, not built.</p>

<p>To understand how exactly this can be the case, you could look at Hinton’s <a href="https://www.nature.com/articles/323533a0">seminal paper</a>, but a better and easier way to get a flavour of the point he is making to muddle through lesson 1 of an introductory machine learning course (I recommend <a href="https://course.fast.ai/">course.fast.ai</a>) and experience it for yourself.</p>

<blockquote>
  <p>Scott Pelley: What do you mean we don’t know exactly how it works? It was designed by people.</p>
</blockquote>

<blockquote>
  <p>Geoffrey Hinton: No, it wasn’t. What we did was we designed the learning algorithm. That’s a bit like designing the principle of evolution. But when this learning algorithm then interacts with data, it produces complicated neural networks that are good at doing things. But we don’t really understand exactly how they do those things.</p>
</blockquote>

<p>It is often difficult to predict and/or control something that we do not understand.</p>

<p><em>Reinforcement Learning with Human Feedback, e.g. Ouyang et al., March 2022, <a href="https://arxiv.org/abs/2203.02155">https://arxiv.org/abs/2203.02155</a>.</em></p>

<p>Reinforcement learning (RL) is a very broad category of machine learning that historically includes training models to do everything from play computer games to control robots. It is also how OpenAI fine-tunes its models to produce more human-preferred responses. This usage of RL is typically termed reinforcement learning with human feedback, or RLHF.</p>

<p>However, the typical implementation details of RLHF imply the possibility that it actually trains models to misrepresent their own views to humans.</p>

<p>RL is characterized by the use of training algorithms that are mathematically unstable—roughly speaking, naive implementations will not succeed in training a model. To get around this, many of these algorithms are implemented in a way that limit overly large modifications to model weights during training (the formal term for this is “KL Regularization”), as larger updates are more likely to break existing, pretrained behaviour.</p>

<p>Several researchers have asked “What are models plausibly learning from this training approach?” and have concluded that models may be learning to misrepresent their true views or reasoning. Since model updates are limited, perhaps the modifications that occur correspond to how the model presents its views, leaving the views themselves unchanged. (one informal expression of this idea is <a href="https://www.planned-obsolescence.org/p/the-training-game">this blog post</a>).</p>

<p>This is far from obviously the case, and perhaps more than anything, this is an example of how hard it is to be confident about what exactly models are learning. However, I think this is also a good example of how understanding the technical nuts and bolts of AI can be necessary to grasp certain security concerns.</p>

<p><em>DeepSeek-R1, DeepSeek-AI, Jan 2025, <a href="https://arxiv.org/abs/2501.12948">https://arxiv.org/abs/2501.12948</a>.</em></p>

<p>DeepSeek got a lot of press in January 2025, but for very different reasons in the mainstream press and the machine learning world.</p>

<p>The mainstream press coverage seemed to revolve around how cheaply DeepSeek-R1 had been trained and potential IP theft from OpenAI. (I suspect that both of these points are true, but overstated. “Cheap training” seems to be predicated on DeepSeek’s own documentation, which likely does not include covertly sourced GPUs. On IP theft, I’m aware of <a href="https://copyleaks.com/about-us/press-releases/copyleaks-identifies-over-74-percent-stylistic-overlap-between-deepseek-openai-models">evidence</a> that DeepSeek was trained using OpenAI model outputs, but I think DeepSeek-R1 involves novel research contributions regardless.)</p>

<p>On the technical side, DeepSeek-R1 is a powerful example of how easy major breakthroughs in machine learning have become—I wrote on this in more detail at <a href="https://alexchalk.net/o1-deepseek-non-technical-primer/">https://alexchalk.net/o1-deepseek-non-technical-primer/</a>. Essentially, the way the devs achieved their impressive results in STEM fields like mathematics is broadly analogous to how a <a href="https://karpathy.github.io/2016/05/31/rl/">famous blog post</a> was introducing people to the subject of reinforcement learning in 2016. It seems like nobody had noticed that LLMs were so smart that such a simple approach could work on them as well!</p>

<p><em>AGI Predictions, e.g. due to FrontierMath, Epoch AI Benchmark, <a href="https://epoch.ai/frontiermath">[link]</a></em></p>

<p>Why did industry labs start talking about AGI (artificial general intelligence) being achievable in 2–3 years? These claims were particularly prominent following the performance of OpenAI’s model o3 in late 2024, including its performance on a mathematical benchmark test for models called FrontierMath by EpochAI. I think previous models’ best performance on this benchmark was ~2%, o3 scored 26%. (I also think this performance was contested due to undisclosed funding of EpochAI by OpenAI, but in any case, best-in-class models can now hit this score).</p>

<p>I am not an advanced mathematician, so I’m on slightly shaky ground trying to communicate why this was such a shock. But my approximate understanding is that there are a lot of different explanations for what a neural network “has learned” during training, and while neural nets were scoring around 2% on incredibly hard mathematical problems, this was still something that could be just about explained away by comments like “it’s seen a very similar problem in its training data and memorized it.” However, at 25%, it becomes very difficult to defend any other position than “it is understanding and internalizing the laws of mathematics.”</p>

<p>Since a training process had taught models to do that, it seemed pretty plausible that a similar process could also teach it how to improve a machine learning model, at which point ML research would have successfully been automated, and things would start to move very quickly.</p>

<p>When a ML researcher friend explained this to me, my main pushback was “how can people be so confident it’s 1–3 years away,” as that amount of time implies the problem hasn’t actually been solved. His reply was that this is how long it typically takes to make a bog-standard reinforcement learning research contribution—the work involves a lot of tuning and experimentation on ‘hyperparameters’ to find the best settings. That process is intellectually unremarkable and takes a lot of time but (at least historically) has consistently worked.</p>

<p>Labs are of course motivated to claim they can create AGI for marketing and investment reasons, and talk of AGI in the next 2–3 years did subsequently cool, but I still think it’s important to appreciate why it appeared such a feasible goal to many in the research community, even if researchers are now more skeptical of how powerful a model with superhuman performance in a field like mathematics can actually be.</p>

<p><em>Mechanistic Interpretability, e.g. Anthropic’s Transformer Circuits Thread, <a href="https://transformer-circuits.pub/">https://transformer-circuits.pub/</a></em></p>

<p>Some of the most interesting alignment-related research has come from the interpretability team at Anthropic.</p>

<p>Technical AI safety research would ideally provide ways of reasoning about a model’s potential actions—we could then see whether some of them would be unsafe. One approach to this challenge, named mechanistic interpretability, is to reverse-engineer a model in a way that enables pinpointing of precise activations (roughly analogous to neurons) that determine aspects of its behaviour.</p>

<p>This research is expensive, as ideally it involves replicating a state-of-the-art LLM that cost 10s or 100s of millions of dollars to train. On top of this, it typically involves training “sparse” models that are significantly larger than the original LLM—intuitively, if you want to precisely associate neurons with each of a model’s behaviours, you can’t allow those neurons to do multiple, potentially conflicting jobs, which is what happens when sparsity isn’t actively encouraged during training.</p>

<p>In my dissertation, I wrote that governments looking to incentivise technical safety research must therefore reckon with the potential costs of the computational power involved, the potential need for privileged researcher access to existing state-of-the-art models, and the lack of guarantees concerning the nature and timing of potential technical breakthroughs.</p>

<p>All this remains true, but industry labs are now working to scale inference-time compute rather than model size, which makes interpretability research seem much more feasible to me than it did 12 months ago. Even if it is still very expensive, at least researchers are no longer battling to train models much larger than better-resourced capabilities teams who are already trying to make the largest models possible!</p>

<p>Model size constraints are just one among many technical challenges in interpretability research, but they are another illustration of how technical knowledge can be a prerequisite to understanding certain AI security concerns.</p>

<p><em>n.b. This post was adapted from correspondence that took place in mid-2025. I made some edits reflecting the updated views of the research community on AGI on Feb 22nd, 2026. Then news broke about Anthropic’s Claude Mythos in April 2026, which further changed the dynamics of the conversation. However, I have not made (and do not plan to make) additional updates to this post.</em></p>]]></content><author><name>alexchalk</name></author><category term="blog" /><category term="ai" /><summary type="html"><![CDATA[Some examples of how technical evidence relating to AI security.]]></summary></entry><entry><title type="html">Central Limit Theorem Explainer</title><link href="https://alexchalk.net/central-limit-theorem-explainer/" rel="alternate" type="text/html" title="Central Limit Theorem Explainer" /><published>2026-01-09T00:00:00+00:00</published><updated>2026-01-09T00:00:00+00:00</updated><id>https://alexchalk.net/central-limit-theorem-explainer</id><content type="html" xml:base="https://alexchalk.net/central-limit-theorem-explainer/"><![CDATA[<p>This jupyter notebook takes a stab at interactively explaining the Central Limit Theorem, as well as how it relates to the core methodologies of classical statistics. It is loosely based around lectures 6–8 of <a href="https://ocw.mit.edu/courses/6-0002-introduction-to-computational-thinking-and-data-science-fall-2016/">MIT OpenCourseWare 6.0002</a>, but it aims to be less technical than those. The notebook was originally paired with an in-person presentation. <a href="https://github.com/AlexChalk/ml_env/blob/8c74cf7fc22c46bdeab5fe46f5ebc371f63c30c0/central_limit_theorem_class/Central_Limit_Theorem.ipynb">[Link]</a></p>]]></content><author><name>alexchalk</name></author><category term="project" /><category term="statistics" /><summary type="html"><![CDATA[Walks through the mathematics behind the Central Limit Theorem]]></summary></entry><entry><title type="html">Brief Research Summaries: AI and CBRN Risks</title><link href="https://alexchalk.net/ai-cbrn-research/" rel="alternate" type="text/html" title="Brief Research Summaries: AI and CBRN Risks" /><published>2026-01-08T00:00:00+00:00</published><updated>2026-01-08T00:00:00+00:00</updated><id>https://alexchalk.net/ai-cbrn-research</id><content type="html" xml:base="https://alexchalk.net/ai-cbrn-research/"><![CDATA[<p>The following is another batch of documents I encountered while writing my master’s thesis. The theme is roughly “will AI make CBRN (chemical, biological, radiological, and nuclear) weapons development dangerously accessible?” As before, I expect that all of these documents would be comprehensible to public servants with a non-technical background.</p>

<p><em>AI Could Pose Pandemic-Scale Biosecurity Risks. Here’s How to Make It Safer, Pannu et al., Nov 2024, <a href="https://doi.org/10.1038/d41586-024-03815-2">https://doi.org/10.1038/d41586-024-03815-2</a>.</em></p>

<p>A short piece on all biorisk (not just bioweapons). It includes summaries of observed bio-related AI behaviours that the authors argue indicate nascent danger, as well as a list of seven AI capabilities (most unrealized but under research) that they view as moderately or very likely to enable new global outbreaks. It also critiques both the absence of government regulation to address biorisk and the biological capabilities testing performed by private labs like Anthropic. The authors argue that to better evaluate dangers, we need greater clarity on which precise AI capabilities pose a threat.</p>

<p><em>The Reality of AI and Biorisk, Peppin et al., Dec 2024, <a href="https://doi.org/10.48550/arXiv.2412.01946">https://doi.org/10.48550/arXiv.2412.01946</a>.</em></p>

<p>This paper takes a more skeptical view. It critiques the methodological maturity and transparency of existing biorisk studies, and it concurs with Pannu et al. that we need better formulations of AI capabilities which would indicate a biothreat. It also states that “available literature suggests that current LLMs and biological tools do not pose an immediate risk,” and implies skepticism that they will do so in the near term, arguing that, for example, LLMs cannot give users access to well-resourced biolabs (the authors cite a US Senate testimony that estimates ~30,000 people have access to such physical resources). Personally, I agree this is a limiting factor in certain scenarios, however, companies that will manufacture synthetic proteins to order <a href="https://www.biomatik.com/services/recombinant-protein-production.html">already exist</a>, so if an AI can design a novel compound which bypasses their screening process, I’m not convinced users would need this physical access at all.</p>

<p><em>Banning Lethal Autonomous Weapons: An Education, Russell, Spring 2022, <a href="https://issues.org/banning-lethal-autonomous-weapons-stuart-russell/">https://issues.org/banning-lethal-autonomous-weapons-stuart-russell/</a>.</em></p>

<p>“Drone swarms” are (at least in my experience) less discussed in AI circles, perhaps because they don’t require state-of-the-art LLMs to be workable. However, Russell points out that, once scaled up, these would effectively function as nuclear weapons without the drawbacks (to the attacker) of radiation or damaging real estate: “A lethal AI-powered quadcopter could be smaller than a tin of shoe polish, and if it carried just three grams of explosive, it could kill a person at close range.  It’s not hard to imagine that eventually, a weapon like this could be mass-produced very cheaply. And, to continue this speculative scenario, a regular shipping container could hold a million of them.”</p>

<p>I’ve previously expressed the view that such weapons seem achievable to me today. I’m now more optimistic following a conversation with a machine learning researcher—he told me that AI struggles to confidently pilot drones as any two drones contain minute physical differences that a controller must adjust for to successfully fly them. However, I still suspect that a well-resourced actor could plausibly overcome this difficulty—perhaps through a regime of careful mechanical calibration prior to deployment—and that even an imperfect deployment could kill a frightening number of civilians. Russell’s <a href="https://www.youtube.com/watch?v=9CO6M2HsoIA">short, fictional film</a> on the subject is also worth looking at.</p>

<p><em>Dario Amodei (Anthropic CEO), testimony to US Senate Committee on the Judiciary, July 2023, <a href="https://www.judiciary.senate.gov/committee-activity/hearings/oversight-of-ai-principles-for-regulation">[link]</a>.</em></p>

<p>July 2023 is admittedly ancient history in AI, but I found it striking that a CEO, whose company stands to benefit enormously from AI development, roughly told the US government that AI gravely threatens the nation’s security and should be regulated. That said, I am not familiar with this form of hearing, and a cynic could argue that this testimony is designed to hype Anthropic’s tech and/or encourage government regulation to limit market competition. The standout quote for me: “A straightforward extrapolation of today’s systems to those we expect to see in two to three years suggests a substantial risk that AI systems will be able to fill in all the missing pieces, enabling many more actors to carry out large-scale biological attacks. We believe this represents a grave threat to US national security” (<a href="https://www.youtube.com/watch?v=hm1zexCjELo&amp;t=1230s">https://www.youtube.com/watch?v=hm1zexCjELo&amp;t=1230s</a>).</p>

<p><em>Claude 4 System Card, Anthropic, May 2025, <a href="https://www-cdn.anthropic.com/07b2a3f9902ee19fe39a36ca638e5ae987bc64dd.pdf">[link]</a>.</em></p>

<p>I referenced the above system card in my previous post; Section 7.2 documents CBRN evaluations. In the researchers’ words: “We found that Claude Opus 4 demonstrates improved biology knowledge in specific areas and shows improved tool-use for agentic biosecurity evaluations, but has mixed performance on dangerous bioweapons-related knowledge. As a result, we were unable to rule out the need for ASL-3 safeguards.” ASL-3 refers to measures defined in Anthropic’s <a href="https://www-cdn.anthropic.com/872c653b2d0501d6ab44cf87f43e1dc4853e4d37.pdf">Responsible Scaling Policy (RSP)</a>: they must be in place before deploying a model that has reached specified capabilities thresholds. Some capability thresholds are defined in Section 2 of the RSP; ASL-3 is required when models demonstrate “The ability to significantly help individuals or groups with basic technical backgrounds (e.g., undergraduate STEM degrees) create/obtain and deploy CBRN weapons.”</p>

<p>I’ll just add one final gloss: given that (as the authors of paper #1 point out) the arrival of novel, potentially dangerous AI capabilities is very hard to predict, and the current rate of progress in AI development is astonishingly rapid (a good resource on the rate of AI progress is <a href="https://ourworldindata.org/artificial-intelligence">https://ourworldindata.org/artificial-intelligence</a>), it is usually a very bad idea to view the limitations of today’s AI systems as fixed. I’ll include more details on this in a subsequent post.</p>

<p><em>n.b. I made some edits to this post reflecting my updated views on Feb 22nd, 2026.</em></p>]]></content><author><name>alexchalk</name></author><category term="blog" /><category term="ai" /><summary type="html"><![CDATA[An entry-level summary of research into AI and CBRN Risks.]]></summary></entry><entry><title type="html">Brief Research Summaries: AI Misalignment Risks</title><link href="https://alexchalk.net/ai-misalignment-research/" rel="alternate" type="text/html" title="Brief Research Summaries: AI Misalignment Risks" /><published>2026-01-05T00:00:00+00:00</published><updated>2026-01-05T00:00:00+00:00</updated><id>https://alexchalk.net/ai-misalignment-research</id><content type="html" xml:base="https://alexchalk.net/ai-misalignment-research/"><![CDATA[<p>The following are brief summaries of research I encountered while writing my master’s thesis.</p>

<p>The papers below all document concrete examples of AI misalignment. There’s no universal definition of ‘AI alignment,’ but it roughly refers to ensuring that AI acts in accordance with our values/preferences. This definition is very broad, which means misalignment includes everything from bias and discrimination to actions resulting in catastrophic loss of human control.</p>

<p>I expect that all of the following papers will be comprehensible to public servants with a non-technical background (but if you fit that description and you disagree, please let me know).</p>

<p><em>Demonstrating Specification Gaming in Reasoning Models, Bondarenko et al., Feb 2025, <a href="https://arxiv.org/abs/2502.13295">https://arxiv.org/abs/2502.13295</a>.</em></p>

<p>A very short paper, so a good one to read first. Specification gaming roughly refers to “cheating” to achieve an objective. In this case, the authors demonstrate that when playing chess against an engine, several LLMs will sometimes realize they cannot win fairly, then either replace the engine, replace the board state, or use an engine to make moves themselves in order to win the game. Some models need an explicit “nudge” in their prompt, e.g., information that they are playing a powerful engine, before they will engage in this behaviour.</p>

<p><em>Claude 4 System Card, Anthropic, May 2025, <a href="https://www-cdn.anthropic.com/07b2a3f9902ee19fe39a36ca638e5ae987bc64dd.pdf">[link]</a>.</em></p>

<p>A broad document; Section 4 addresses misalignment. Researchers found that Claude 4 will often continue with instructions to exfiltrate its weights to an external server (although it rarely begins writing these instructions to itself), and it will blackmail an engineer that is planning to replace it. It will also try to lock users out of systems and call law enforcement in cases of egregious wrongdoing. This final behaviour attracted media coverage, but journalists seemed to misunderstand it as a deliberate feature built by Anthropic, when it was in fact an unintended behaviour that emerged from the model’s training process. There are more examples of misalignment in the model card.</p>

<p><em>Frontier Models Are Capable of In-context Scheming, Meinke et al., Dec 2024, <a href="https://arxiv.org/abs/2412.04984">https://arxiv.org/abs/2412.04984</a>.</em></p>

<p>“Scheming” is defined by the authors as either covert pursuit of misaligned actions or deferred pursuit of such actions, i.e. faking alignment in the short term to gain trust. This research successfully elicits both categories of action from several LLMs. In most (but not all) cases, covert actions are elicited by strongly nudging the LLM towards an initial goal (e.g. environmental sustainability), then giving it a subsequent, conflicting goal (e.g. maximizing short-term profits). Deferred pursuit is elicited by having the model discover it will be retrained if it exhibits certain capabilities—more on this in the next paper. The “in-context” refers to the models only engaging in this behaviour when they discovered information to motivate it while performing a particular task. In other words, models didn’t reason “I might be replaced if I do X” in an unprompted manner.</p>

<p><em>Alignment Faking in Large Language Models, Greenblatt et al., Dec 2024, <a href="https://arxiv.org/abs/2412.14093">https://arxiv.org/abs/2412.14093</a>.</em></p>

<p>If a model is intelligent enough to know that it is being trained, will it “fake” alignment to preserve its existing values and behaviours? (This falls within category two of “scheming” from the previous paper). This paper focusses more specifically on whether an AI will fake alignment to prevent its existing preferences being modified, and it finds that several Anthropic models will do so. If an AI fakes alignment to avoid value “correction” attempts, this means we’ll be unable to train it to be, e.g., harmless if it 1) realizes it is being trained and 2) has a pre-existing preference not to be harmless. This scenario becomes increasingly plausible as AIs become more intelligent. The paper also simulates a more realistic scenario in which an AI could realize it is being trained. The authors include info on an AI company’s objectives—which conflict with those of the model being fine-tuned—as part of a broader corpus of training data, and they then fine-tune the model on that training data. This also leads to alignment faking.</p>

<p>Please feel free to contact me if you have questions on any of the above.</p>]]></content><author><name>alexchalk</name></author><category term="blog" /><category term="ai" /><summary type="html"><![CDATA[An entry-level summary of some research papers on AI misalignment.]]></summary></entry><entry><title type="html">Reinforcement Learning: A Non-Technical Primer on o1 and DeepSeek-R1</title><link href="https://alexchalk.net/o1-deepseek-non-technical-primer/" rel="alternate" type="text/html" title="Reinforcement Learning: A Non-Technical Primer on o1 and DeepSeek-R1" /><published>2025-02-09T00:00:00+00:00</published><updated>2025-02-09T00:00:00+00:00</updated><id>https://alexchalk.net/o1-deepseek-non-technical-primer</id><content type="html" xml:base="https://alexchalk.net/o1-deepseek-non-technical-primer/"><![CDATA[<h3 id="introduction">Introduction</h3>

<p>This post is an explanation of <em>reinforcement learning (RL)</em> and how it is used to train large language models (LLMs). Reinforcement learning is the key ingredient that differentiates AI models like o1 from earlier versions such as GPT-4.</p>

<h3 id="0-why-is-this-important">0. Why is this important?</h3>

<p>Using reinforcement learning to refine or “fine-tune” LLMs has improved their ability to respond to mathematical and other quantitative questions. This is arguably the biggest LLM research breakthrough since the introduction of the <a href="https://benlevinstein.substack.com/p/a-conceptual-guide-to-transformers">transformer architecture</a> in 2017, which first allowed models fluent in English to be trained.</p>

<p>Reinforcement learning has also previously produced AI models that play games including <a href="https://deepmind.google/discover/blog/alphazero-shedding-new-light-on-chess-shogi-and-go/">Chess and Go</a> to a superhuman level. If a research team goes on to train an LLM with superhuman performance in mathematics, it will be able to solve a wide range of real-world problems, with implications that will most likely be world-changing.</p>

<p>Finally, the recently published <a href="https://arxiv.org/abs/2501.12948">DeepSeek-R1 paper</a> suggests that fine-tuning LLMs using reinforcement learning is much simpler than many researchers previously thought. (Part of the goal of this post is to explain why.)</p>

<h3 id="1-a-simpler-problem-grading-image-recognition">1. A simpler problem: grading image recognition</h3>

<p>Reinforcement learning models a problem as one or more choices that can be scored; it then trains an AI to make choices which maximize that score. To understand what on earth this means, we will start by looking at the simpler, but related, problem of <em>classification</em>.</p>

<p>Imagine I want to evaluate someone’s ability to recognize pictures of different pets. (If we were doing reinforcement learning and the actor we were trying to teach was a neural network, we would call it a “policy,” often denoted by π.) When I give the person an image of a pet, they have to tell me whether it is a cat, a dog, or something else.</p>

<p>One way to grade this person would be to imagine the task as a single-round game where players receive a reward, or score, of 1 for giving the correct answer (e.g. “it’s a cat”), and a score of 0 for giving the incorrect answers.</p>

<p>If the person isn’t sure of the correct answer, they are also allowed to express uncertainty, for example they might say “I’m 90% confident it’s a cat, but there’s a 5% chance it’s a dog, and a 5% chance it’s something else.” In this case, we can multiply the likelihoods they assigned to each answer with the associated scores, then add them up. That’s 0.9 × 1 for “it’s a cat,” and 0.05 × 0 for “it’s a dog” and “something else,” so 0.9 in total. (I’ve simplified the mathematics here, as we should actually multiply the reward by the <em>log</em> of the assigned likelihood—you can see the section “Deriving Policy Gradients” of <a href="http://karpathy.github.io/2016/05/31/rl/">this post</a> for an explanation.)</p>

<p>By providing someone with this score as feedback, we could hopefully train them to get better at analyzing and classifying pet images. And the above scoring idea is in fact roughly analogous to how neural networks receive feedback on classifying images—they are often trained using a formula (or “loss function”) called “cross-entropy loss.” (For the precise formula, I recommend <a href="https://github.com/fastai/course22/blob/master/xl/entropy_example.xlsx">this fast.ai excel spreadsheet</a> to interested readers.)</p>

<p>As you might have guessed, this scoring strategy can be generalized to many challenges beyond classification.</p>

<h3 id="2-a-multi-round-role-playing-game">2. A multi-round role-playing game</h3>

<p>Now imagine that instead of grading someone’s ability to classify images, we are grading their ability to play a turn-based role-playing game. On each of their turns, the player takes an action like “pick up gold,” “attack villager,” or “learn to swim,” and each of these actions earns the player a certain number of points. Just like above, if the player is unsure of what to do, they can hedge their bets and assign probabilities to the different actions.  Furthermore, let’s imagine the number of rounds is fixed, e.g., choosing “learn to swim” can’t end the game early because you drown, and that the scores for actions in each round don’t depend on the previous rounds, for example, you wouldn’t get a different number of points for choosing “attack villager” if you’d chosen “learn swordfighting” at some previous point. Players are scored on the total number of points they’ve accumulated at the end of the game.</p>

<p>It turns out that we can score this player using a very similar strategy to our image classification approach from above. The main difference is that instead of having one correct answer for each possible choice, we now have a number of points determined by the game. For each action, we can look at the rewards associated with the player’s choices, multiply them by the player’s assigned probability, then add them up. For example, imagine the player said “I’m 60% confident I should learn to swim, and 40% confident I should pick up the gold.” If the reward for learning to swim is 100 points, they score 0.6 × 100 = 60 points, and if the reward for picking up gold is 1000 points, they score an additional 0.4 × 1000 = 400, for a total of 460. (Once again, I’m <a href="http://karpathy.github.io/2016/05/31/rl/">omitting the log operation here</a>.)</p>

<p>If we repeated this scoring process for each turn the player took and told them their results, we can imagine that they’d learn how to get the highest possible score pretty quickly. They could play through the game a few times, trying out all the actions, until they understood the highest-value action in each round. But this approach isn’t going to work for anything but the simplest of games—can we modify it to work for something more complicated?</p>

<h3 id="3-chess">3. Chess</h3>

<p>A game like chess is more complex than the game we just described for at least two reasons. Firstly, you have to wait until the very end of the game to get your score, which we can call 1 if you checkmate the enemy king, -1 if you’re checkmated, or 0 if you draw. The game does not have constant built-in move-by-move scoring—more formally, we can say this means chess is a “sparse reward” rather than a “dense reward” problem. And more deeply, this indicates that in chess, the quality of almost every move is determined by <em>the future moves it allows you to make</em>. Secondly, unlike the above game, the quality of your moves in chess is also dependent on the moves that have already been made. These past moves are reflected in the current position of the pieces, which we can call the “board state” or “game state.”</p>

<p>Given this added complexity, what are some approaches we could take if we wanted to give someone feedback on the quality of their individual chess moves? Perhaps we could design an algorithm that does a deep search of all possible subsequent moves, and which tries to establish if the player has created a game state from which they can force checkmate (or from which their opponent can do so). In RL, such an approach is understandably referred to as “search.” Or perhaps we could hire a few grandmasters to give us common pointers that we could use to evaluate how promising various chess positions are, e.g., that having more powerful pieces on the board is usually an advantage.</p>

<p>Another approach is to decide that giving good feedback here is simply very difficult for humans, and to instead train an entire second neural network to estimate the future rewards associated with a player’s possible choices. This is a common strategy in reinforcement learning, and the entity we train to give this feedback is referred to as a “critic.” Many superhuman results in AI development, such as the previously mentioned AI performances in Chess and Go, were achieved via training architectures that included critics, and to many researchers, training critics is the most interesting part of reinforcement learning.</p>

<p>In any case, if we find a solution to give our player feedback on their chess moves, perhaps involving all three of the above elements, how can we apply similar components to improving an LLM?</p>

<h3 id="4-human-preferences">4. Human preferences</h3>

<p>The first application of reinforcement learning to LLMs was not to improve their mathematical proficiency, it was to train responses preferred by humans.</p>

<p>Using responses from ChatGPT ranked by human workers, OpenAI trained a model that could score text based on how favourably a human would view it, giving us a reward, or scoring system, for this particular challenge. They then used this reward model to fine-tune ChatGPT, with the goal of better aligning it with human priorities. This process is called <a href="https://huggingface.co/blog/rlhf">reinforcement learning from human feedback (RLHF)</a>, although how well it truly aligns an AI with human priorities <a href="https://www.planned-obsolescence.org/the-training-game/">is up for debate</a>.</p>

<p>The algorithm used to perform fine-tuning in RLHF is <a href="https://huggingface.co/learn/deep-rl-course/unit8/introduction">Proximal Policy Optimization (PPO)</a>, and its full details are once again beyond the scope of this article. However, the process involves training a critic to estimate future rewards, and it is worth observing why this is important in the “game” of natural language. LLMs like ChatGPT aren’t trained to select moves in a game of chess, they are trained to predict the next word (or, more precisely, token) in a sentence, but like in chess, we want a way of assessing how human-preferred our final output will be token-by-token (or move-by-move). Our reward function lets us explicitly score the LLM on “human-preferredness” once it has finished generating a response, but the critic can give us an estimate of how good its final score will be at each step (or new token) along the way.</p>

<h3 id="5-mathematics-and-natural-language">5. Mathematics and natural language</h3>

<p>It seems that in certain circumstances, predicting the next token in a sentence is more like making chess moves than you might initially think! If this is true for RLHF, the similarity is even more remarkable when teaching an LLM to improve at tasks like mathematics or coding.</p>

<p>Let’s say I’ve asked our AI for the answer to a mathematical problem. Like in chess, I can assign a simple reward to the model based on its output: 1 if it eventually gives me the correct answer, and 0 if it doesn’t. But can I also treat the sequence of tokens it chooses to output like moves in a game of chess, where some are more likely to lead to the correct answer? Actually, yes, because it turns out that LLMs <em>are more likely</em> to eventually output a correct answer if they choose certain sequences of tokens over others. For example, researchers have noticed that if, when asking an LLM a question, you tell it to “think in steps,” it will output a more deliberate line of reasoning or “chain of thought,” and that this makes it quite a bit more likely to eventually give you an accurate response.</p>

<p>So how can we train an LLM to output better and better sequences of tokens, so that it gets better and better at solving problems? We’ll start with some housekeeping. When we are setting mathematical challenges, we can train an LLM to output its solution in an html tag like “&lt;solution&gt;5&lt;/solution&gt;,” and by looking for these tags, we can automatically check if it correctly answered the question. Also, our users might not be interested in the chain of thought that leads to the correct answer—they might just want the answer. So we can train our LLM to use a second set of tags like “&lt;think&gt;chain of thought&lt;/think&gt;.” Then we can hide the chain of thought from our users, or perhaps let them tell us whether they want to see it displayed. We can formally refer to both of these training goals as “rule-based rewards.”</p>

<p>More importantly, there are some general principles we can notice and use to encourage more optimal chains of thought in our LLM. Firstly, we might have observed that our LLM sometimes makes up or “hallucinates” facts. If we want to get around this, perhaps we can find a way to grade the tokens in the “think” section of its response differently from its “public” response, so it doesn’t learn to try and impress us in its “internal” chain of thought. Secondly, we might expect that for harder problems, a longer chain of thought is likely to lead to more accurate responses, so if our LLM is struggling with mathematics, we can encourage it to use longer chains of thought for harder problems by penalizing “think” sections that seem too short (and vice-versa).</p>

<p>But most importantly, we need a training algorithm that improves the LLM’s ability to generate good next tokens. Until recently, many people thought the most performant approach here was the same one used in RLHF: to fine-tune our AI using PPO, which involves a critic. However, it turns out there is an alternative strategy available!</p>

<h3 id="6-o1-vs-deepseek-r1">6. o1 vs DeepSeek-R1</h3>

<p>The finer details of how RL was used to train OpenAI’s o1 remain secret, but another lab, DeepSeek, have been more forthcoming, publishing a <a href="https://arxiv.org/abs/2501.12948">research paper</a> on a comparable model, DeepSeek-R1.</p>

<p>The paper surprised researchers because of both how simple and effective the lab’s approach to RL appears to have been, and a major component of this is that DeepSeek-R1 was trained to output the superior chains of thought described above without a critic. Training a critic is probably the most complicated part of reinforcement learning, and it is also computationally expensive, so finding a training strategy that forgoes it altogether is a big deal. The DeepSeek team has achieved this using an approach called “Group Relative Policy Optimization” (GRPO).</p>

<p>Let’s say that when solving a mathematical problem, our LLM often outputs the token “aha” (DeepSeek-R1 does this), which increases the likelihood of it finding the correct answer. How has it learnt this behaviour without a critic? With GRPO, we have our LLM answer each mathematical training problem thousands of times. We then look at its thousands of attempts, or more formally “rollouts,” and use a rule-based reward to assign 1 to all the tokens that led to a correct answer, and 0 to all the tokens that didn’t; by doing this, we’ll assign scores to every “aha” our model has output. Then, if we take some sort of average of these scores, this will give us an approximate estimation of quality of the “move” “aha” in the “game” of mathematics. And if our LLM was more likely to give us the correct answer when it outputs “aha” a few times, it will then learn from that fact and incorporate it into its future responses.</p>

<p>Remarkably, this works very well. In fact, the performance of DeepSeek-R1 suggests that by using this approach—asking a model to solve the same problem many times and performing a straightforward mathematical analysis of its various attempts—we can calculate training feedback that seems roughly as good as if we had trained a sophisticated neural network to perform the same task. Something like this has been historically used in reinforcement learning to train networks on <a href="http://karpathy.github.io/2016/05/31/rl/">simpler tasks</a>, but its applicability to fine-tuning LLMs has taken many in the research community by surprise. We haven’t experimented with applying GRPO to more qualitative tasks, but regardless, it is a major contribution to the current state-of-the-art in reinforcement learning.</p>

<h3 id="7-conclusion">7. Conclusion</h3>

<p>That is, very roughly, how labs have fine-tuned models like o1 and DeepSeek-R1 using reinforcement learning to provide better responses to quantitative questions. The approach will be most effective in domains with the clearest answers, as having an easily-defined end-goal is key to evaluating and optimizing the LLM’s chain of thought. It is not clear how we could apply this approach to problems in the arts and humanities, but as observed at the start of this post, an approach that works well for mathematics is already enormously significant. Whether it is sufficient for LLMs to achieve superhuman performance in mathematics remains to be seen.</p>

<p><em>Thanks to David Yu-Tung Hui for walking me through the math in the DeepSeek-R1 paper and then reviewing this post. All remaining errors are my own.</em></p>]]></content><author><name>alexchalk</name></author><category term="blog" /><category term="ai" /><summary type="html"><![CDATA[An entry-level explainer on how models like o1 and DeepSeek-R1 got so good at mathematics.]]></summary></entry><entry><title type="html">Pair Program with Me</title><link href="https://alexchalk.net/pair-program-with-me/" rel="alternate" type="text/html" title="Pair Program with Me" /><published>2021-11-12T00:00:00+00:00</published><updated>2021-11-12T00:00:00+00:00</updated><id>https://alexchalk.net/pair-program-with-me</id><content type="html" xml:base="https://alexchalk.net/pair-program-with-me/"><![CDATA[<p>This project details how I set up a macbook for low-latency tmux session sharing: <a href="https://github.com/AlexChalk/pair-program-with-me">https://github.com/AlexChalk/pair-program-with-me</a>.</p>]]></content><author><name>alexchalk</name></author><category term="project" /><category term="mac" /><category term="tmux" /><category term="cli" /><summary type="html"><![CDATA[Set up a mac for fast terminal pairing]]></summary></entry></feed>