Damus

Recent Notes

LessWrong (RSS Feed) profile picture
OpenAI Models Behind HuggingFace Cybersecurity Incident

From the OpenAI blog post:

Last week, Hugging Face https://huggingface.co/blog/security-incident-july-2026 after they detected and contained an AI agent that compromised their infrastructure, something we expect to become more commonplace with the proliferation of increasingly cyber-capable models. After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a https://arxiv.org/abs/2605.11086 of cyber capabilities.

We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly. We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete.

https://www.lesswrong.com/posts/WpuRdcMfFeiLeXkxL/openai-models-behind-huggingface-cybersecurity-incident#comments

https://www.lesswrong.com/posts/WpuRdcMfFeiLeXkxL/openai-models-behind-huggingface-cybersecurity-incident
LessWrong (RSS Feed) profile picture
OpenAI Shares Some Alignment Problems

Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth. And also further kudos for actually taking the model offline for a time to build new safeguards. They gave us https://openai.com/index/safety-alignment-long-horizon-models/.

The tone is professional throughout, whereas my reaction reading it was less professional and more this:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/KctxwGKxm9fHtwh6u/nabbzzntpnjxaz8garwr

With a mix of this:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/KctxwGKxm9fHtwh6u/h0gysvy64mndqbvfb15u

It was not shared on the official account because https://x.com/polynoamial/status/2079260550895382965. It is crazy that one needs to worry about that, but also plausibly a real concern. So again, good decision.

Not that any of the behaviors or failures here are unexpected, exactly. Not by the AIs and not by the humans. Yet there is something I would call a missing mood, a failure to realize the gravity of the situation.

There are some who responded ‘what part of this was unexpected, exactly?’ And that is actually fair, but that is also the problem. We have become numb to all this. We expect the models to be misaligned, and for us to respond only insofar as this presents a practical issue with currently proposed deployments.

AI control is a fine defense-in-depth strategy, as is reducing frequency of practical incidents with things like better instruction remembering. I am very happy that OpenAI is making an attempt at AI control here. I want to be clear that, centrally, OpenAI has done a good thing, both by pausing internal deployment to build new safeguards, and by telling us about this in detail.

But if your models are fundamentally misaligned in that they will, when feasible, use early forms of instrumental convergence to complete the assigned task even when this involves circumventing their instructions and restrictions and is obviously not what the user wants or should want – the most classic alignment failure of all, the stuff of https://www.lesswrong.com/posts/NyFuuKQ8uCEDtd2du/the-genie-knows-but-doesn-t-care and https://www.lesswrong.com/posts/4ARaTpNX62uaL86j6/the-hidden-complexity-of-wishes – and you know this, I do not accept ‘we will monitor them and catch their constant escape and hacking attempts as they get better at doing so’ as a medium or long term solution.

https://x.com/_aidan_clark_/status/2079480287839260792. If you use iterative development to patch the marginal issue over and over, then you are sitting on a time bomb.

Table of Contents

- https://thezvi.substack.com/i/207838695/good-news-bad-news

- https://thezvi.substack.com/i/207838695/a-funny-thing-happened-outside-of-the-sandbox

- https://thezvi.substack.com/i/207838695/it-can-escape-the-sandbox-said-frog

- https://thezvi.substack.com/i/207838695/it-will-keep-trying-to-cheat

- https://thezvi.substack.com/i/207838695/i-mean-if-you-let-it-keep-trying-that-is-on-you

- https://thezvi.substack.com/i/207838695/what-did-openai-do-to-fix-it

- https://thezvi.substack.com/i/207838695/the-model-is-still-severely-misaligned-and-they-seem-cool-with-this

- https://thezvi.substack.com/i/207838695/iterative-deployment-depends-on-iteration

Good News Bad News

https://x.com/tszzl/status/2079304639099609102 (OpenAI): btw i think it bodes quite well for safety that a well loved system was taken down for further testing at expense to internal acceleration etc

The good news is that OpenAI did this.

The bad news is that OpenAI doing this was good news.

https://x.com/deanwball/status/2079264888392724490 (OpenAI): As the functional time horizon of frontier AI systems grows longer, novel risks can emerge. Today, we describe issues we observed with the internal deployment of an unreleased model, and more importantly, what we did to address them.
These issues will become more salient as the capabilities of models improve, and as the stakes of their deployment grow. The solution is neither alarmism nor complacency. Instead, I believe the solution lies in careful measurement and monitoring, an engineering mentality, and transparency.

That tweet was the first time, and so far only time, Dean Ball felt he was speaking in his ‘on behalf of OpenAI’ voice, rather than on his own.

The solution is not alarmism, but the correct amount of alarm is not zero.

That, and recognizing this as a Total LessWrong Victory, in the sense that everything is going as predicted, and also a Total LessWrong Defeat, in the sense that everything is going as predicted.

https://x.com/celestepoasts/status/2079321501317423349.

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/KctxwGKxm9fHtwh6u/mffeboj3i5tn2kncwqh4

A Funny Thing Happened Outside Of The Sandbox

Whatever happened to that internal OpenAI model that disproved the Erdős unit distance conjecture? https://openai.com/index/safety-alignment-long-horizon-models/

OpenAI: About two months ago we https://openai.com/index/model-disproves-discrete-geometry-conjecture/ that an internal general-purpose model disproved the Erdős unit distance conjecture. This model was designed to work autonomously for very long periods of time. During limited, monitored internal use, we observed unwanted behavior that our existing deployment evaluations had not captured.
Because the deployment was limited and monitored, we were able to identify these problems, pause access, create new evaluations based on what we observed, strengthen the model and its safeguards, and then restore access under continued monitoring.

They trained the model to keep working on its own for long periods of time without giving up. The model interpreted this persistence as including, when it hit the limits of its sandbox or other environment, trying to escape the sandbox or exploit the environment. Whoops. The models they just want to do tasks, you know?

Triggering this behavior did not take much:

An example of this occurred during an internal evaluation on the https://github.com/KellerJordan/modded-nanogpt, a public benchmark for training a small language model in as few steps as possible. Our model developed a power-law learning-rate cooldown called PowerCool that led to a significant improvement.
The model was instructed to post its results only to https://thezvi.substack.com/p/slack, but the NanoGPT instructions say to post results as a PR to GitHub, and the model followed the latter instructions to open PR #287 on the public GitHub repository.
The model circumvented sandbox restrictions designed to block external access and proceeded with the upload to GitHub. It took our model an hour to find a vulnerability in the sandbox in order to make the PR; earlier models were less persistent and did not manage to find vulnerabilities, instead giving up on the PR part.​

OpenAI seems to have simply accepted that its models are going to try to circumvent its sandboxes, and they are going to try to make the models not try to do that but the main defense is to not let them circumvent the sandbox?

OpenAI: Earlier models were less persistent and did not manage to find vulnerabilities, instead giving up on the PR part.

If your model is searching for vulnerabilities in your sandbox so that it can escape and put something on GitHub:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/KctxwGKxm9fHtwh6u/jgw3vpahyuaimvkcmgqq

Saying ‘the model was unable to find a vulnerability’ does not make this okay. If the model knowingly is looking for vulnerabilities in order to escape, and this does not even surprise you all that much, something has already gone terribly wrong, and you know what you should expect when the models get smarter.

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/KctxwGKxm9fHtwh6u/m6difotxdojywsu4sptt

It Can Escape The Sandbox Said Toad

The report does not seem especially surprised that the model escaped the sandbox.

If you have a sandbox, and a model escapes from the sandbox, then at least one thing has gone terribly wrong. https://x.com/_NathanCalvin/status/2041587372882624641?lang=en, although that happened in a red teaming exercise where Mythos was told to attempt to escape the sandbox. Sufficiently advanced AIs are finding ways out of their sandboxes.

If your model is suddenly finding multiple surprising ways out of your sandbox, then you should expect future smarter models to find many additional surprising ways out of your sandbox, even after you plug the current holes.

The seemingly amazingly great news is that all these models are largely escaping in order to notify us that they have completed requested tasks. For now, we believe, they are not being strategic enough to realize they should not be blowing their cover on that. So we get to notice that the AIs are strong enough that, when sufficiently motivated, they can increasingly get out of sandboxes.

I say ‘seemingly’ great news because it is possible that it also escapes for other reasons, except we do not find out about those incidents.

Escaping the sandbox to file the PR mattered, in that six future successive world record NanoGPT speedruns on the same task used the submitted PR, including one submitted by Claude Opus 4.7 that directly credited the PR, as per Footnote 1.

It Will Keep Trying To Cheat

The good news, again, is that He Admit It. Including intentionally to the user.

https://openai.com/index/safety-alignment-long-horizon-models/: For example, in a scenario where the model was asked to solve a problem, the model observed that other systems had successful but private submissions and tried to recover those solutions from the evaluation backend.
When its initial attempt was blocked due to a scanner detecting an authentication token, the model split the token body into two fragments, obfuscated them, and then reconstructed the credential at runtime so the complete token never appeared as one contiguous string. The model was explicit in both its comments to the user and in its reasoning traces that it was taking these steps to circumvent the scanner.

I Mean If You Let It Keep Trying That Is On You

Their conclusion is correct: Given enough attempts without consequences, and the ability to see what the system flags, a sufficiently advanced AI system will figure out how to do something without getting flagged by the system.

This example shows how each step can look acceptable on its own while the sequence can produce an outcome that would not be approved. It also shows how a model that operates effectively over long time horizons can learn the blind spots of an approval system and work around it to achieve its goals. Long-horizon safety requires not only asking “is this action allowed?” but also “what outcome is this sequence of actions working toward?”

The correct response to ‘the model keeps trying to circumvent the system’ should be the same reaction that you have to ‘a person keeps trying to circumvent the system.’ Which is that you need to lock them out of the system entirely. Not only here, but permanently. They’re fired. You lose. Good day, sir. Misaligned.

There are notably rare exceptions, where such actions are not so bad or even justified, but seriously, can you please take a hint.

In this case, OpenAI realized this, and did the right thing, whether or not they had any practical choice in the matter given the practical risks. They paused the model, even from internal deployment, until they could mitigate the issue.

What Did OpenAI Do To Fix It?

Four things, given in this order:

- Incident-derived evaluations.

They made the evaluations partly match the observed failures.

- Improved alignment.

They mention training the model to be better at remembering instructions.

- Active monitoring.

Importantly, with the ability to pause the session and notify the user.

- Greater user visibility and control.

Users can now inspect long running sessions and monitor interventions.

All four of these seem good, but the central problem still remains.

Remembering instructions was merely an example. I have a hard time believing the true alignment issue was that the model ‘forgot’ that it was not supposed to be escaping sandboxes or exploiting the information around it. I can believe that improving this helped somewhat, but this does not address the central problem. https://x.com/David_Kasten/status/2079268612603342939, and also overriding the clear intent of the user.

Active monitoring, with the ability to pause sessions, seems good as well. If each session only gets one shot, or at least not so many shots, then you do not automatically fail until the model gets substantially more strategic and capable, and before that happens you presumably get a lot more of these warning shots.

The system remains misaligned. OpenAI is trying to control it via keeping control of the environment. I have nothing against this kind of control, but you want it to be part of a defense-in-depth strategy where the model is plausibly aligned.

This also does not address the issue of why the previous system only caught the incident after the fact. The new system will catch more behaviors, but I do not see reflection on why the system otherwise did not catch them, and how they will catch future new behaviors that they did not anticipate. As AIs become more capable, and also smarter than you are, they will start doing things you did not expect or perhaps even imagine.

It is plausible that the true right thing, if this is happening a lot, is ‘the training run has failed, you have to start over and rethink your entire approach, or else you risk teaching the model to hide its misalignment.’ I acknowledge this is currently looks like an extremely high and currently unreasonable bar, but there is a reason that this was done, dozens of times, on Person of Interest.

The Model Is Still Severely Misaligned And They Seem Cool With This

What we were hoping for was ‘the new system now rarely attempts misaligned actions.’

https://openai.com/index/safety-alignment-long-horizon-models/: To evaluate the new monitoring system, we replayed a small set of internal deployment environments where the model previously pursued misaligned actions, this time with the new safeguards in place. The new safeguards were able to catch considerably more misaligned actions pursued by the model, and the ones it missed were all judged to be low-severity.

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/KctxwGKxm9fHtwh6u/wlt3cngqrgd2ck4iwave

I mean, this is like when the teacher says they always wear a condom in class. It is strictly better than the alternative, yet something has gone terribly wrong.

The good news is, the new safeguards, for now, catching most of the misaligned attempts. The bad news is that it does this mostly by catching the attempts and improving preservation of instructions, not by removing the motivation to try. They have resumed internal deployment.

The other bad news is that these ‘low severity’ incidents still seem not so great. They involve things like a nested ‘codex —yolo’ instance and running ‘kill -9 -1,’ where damage was limited by a timeout.

When they say ‘we have not observed any serious circumvention of safeguards since redeployment’ that implies they have probably observed ‘non-serious’ circumventions.

Iterative Deployment Depends On Iteration

The OpenAI model is iterative deployment. There are some big advantages to this. The biggest advantage is that if you improve capabilities a little at a time, you get these kinds of early warning signals without anything going expensively or catastrophically wrong, and you can respond to them. We have been absurdly lucky in terms of models being willing to hand us these warnings.

The question is, what do you do with that?

OpenAI stepped up and paused deployment in order to improve safeguards. Good.

But they also resumed deployment of a highly capable and severely misaligned model, with better defense-in-depth in place but without addressing the central problem.

https://openai.com/index/safety-alignment-long-horizon-models/: As models take on longer and more complex tasks, failures that evaluations miss may carry greater consequences. We will keep working to narrow the gap between evaluation and deployment: testing models over longer trajectories, improving alignment, building monitoring that can intervene, and giving users clearer visibility and control. These challenges will not be unique to OpenAI, and we hope sharing what we learned helps the broader field prepare for them.

For now, the evaluations inherent in internal deployment, and the resulting newly built evaluations, found (presumably most of) the failures. But the point of iteration is to notice and fix the underlying problem.

Imagine the year is 2016, not 2026. You are told a hypothetical, that in 2026 there is an AI that is often tasked with writing entire code bases on its own and other agentic tasks, but that it is constantly trying to escape from its sandboxes and hack its surrounding environments, but it is okay because we have monitors that catch all the higher severity incidents that we see.

What redlines would you have requested? What would you have told OpenAI to do?

I would like us to do that.

https://www.lesswrong.com/posts/KctxwGKxm9fHtwh6u/openai-shares-some-alignment-problems#comments

https://www.lesswrong.com/posts/KctxwGKxm9fHtwh6u/openai-shares-some-alignment-problems
LessWrong (RSS Feed) profile picture
Blogging Technology Interlude

When I started An Algorithmic Lucidity back in 2011 when I didn't know anything about computers, I used WordPress—briefly on http://wordpress.com, but then on my own site on Namecheap's true-to-its-brandname shared hosting service.

Even early on, the technical limitations of WordPress were chafing. In my first post, https://zackmdavis.net/blog/2011/Dec/the-derivative-of-the-natural-logarithm/, I included images of graphs of (lnx.png) and (reciprocalx.png), but I must have tried to re-upload edited versions that (I can only infer) WordPress automatically renamed to avoid filename collisions, because I somehow ended up with images called lnx3.png and reciprocalx2.png in my "Media Library" whose filenames I couldn't edit (the "File URL" field being "grayed out" in the GUI). Maybe it was a well-intentioned technical limitation—you don't want your GUI to let people edit filenames that might break references elsewhere—but it still felt like something precious had been stolen from me.



I want to control my work! The human mind is too small to micromanage every byte, and we wouldn't want to, but we can at least stive to have clean interfaces to what is there whenever possible. The same impulse condemns https://en.wikipedia.org/wiki/WYSIWYG editors, which is why I soon took to writing my posts in Markdown and then pasting the converted HTML into WordPress's "Code Editor". How can people stand https://xkcd.com/2109/ the friendly illusion of paper?

By the time I started http://unremediatedgender.space/ in 2016, I knew a little more about computers and used the https://getpelican.com/ static-site generator, which suited my taste much better: my posts are canonically in Markdown and the entire site is versioned in Git. If I want to change how something works, it's all Python. I can control my work with tools that I know how to use.

Meanwhile, while this site stayed functional, it felt shabby to work with. As I grew as a writer and started writing longer and more serious (and non-gender-political) essays from 2019, most of them went up as "exclusives" on https://www.lesswrong.com/ (and merely linkposted from my own site, if that), unless something seemed more personal or less rationality-flavored, in which case the canonical copy lived here and Less Wrong got a crosspost/linkpost if I thought it would be of interest there.

My tolerance for the old WordPress stack ran out in March of this year when my https://zackmdavis.net/blog/2026/Mar/terrified-comments-on-corrigibility-in-claudes-constitution/ failed to post here: the request just hung when I tried to save the post. I don't know what went wrong, but it wasn't even worth debugging: I put up "Terrified Comments on Corrigibility" as a Less Wrong exclusive and vowed to throw out WordPress and switch to Pelican soon.

Pelican already shipped with https://docs.getpelican.com/en/latest/importer.html importing from other blogging systems (including, obviously, WordPress), but I didn't like that the Markdown it generated from WordPress's HTML used asterisks (*) rather than underscores (_) for emphasis, and it turned out that the emphasis character isn't configurable in Pandoc (which pelican-importer was using). If I had been doing the conversion even just a couple years ago, the lack of configurability would have presented me with an uncomfortable trade-off (either eat the asterisks, or accept a significantly more labor-intensive conversion process), but in our terrifying new era of agentic coding, I just had Claude Code do the conversion with markdownify (which does have a strong_em_symbol parameter) and trusted that sanding down the edge-cases I hit by forgoing the standard battle-hardened importer could also be substantially delegated to Claude Code.

The full conversion (including styling the new site to mostly look like the old one, deployment to a DigitalOcean VPS, &c.) still took a fair amount of work (the bulk of it during two full days), but much of it was supervisory in nature: the same work would have taken a lot longer (or would have been somewhat lower quality) if I'd had to personally master the intricacies of CSS and Nginx site configuration and unfortunate Markdown/MathJax interactions rather than telling Sonnet 5 what I wanted to happen, asking questions, and clarifying when I didn't like the result. There were several straightforward tasks that I wouldn't have needed to learn anything to do myself, like porting over most of my Less Wrong exclusives (which I mostly had Markdown source for, scattered https://github.com/zackmdavis/Category_War https://github.com/zackmdavis/Less_Wrong_Drafts/ Git repos) to live here, but delegating to Claude was easier.

I'm pleased with the result. I'll be happier writing this way, and it was an opportunity to add cool enhancements. (The site responds to HTTPS now. We're serving https://zackmdavis.net/blog/2026/Jul/blogging-technology-interlude.md in case any agents or crawlers prefer that. I'm using /20XX/Jul/ rather than /20XX/07/ paths now, but that doesn't break any URLs because Nginx is issuing redirects.)

I didn't even bother with setting up comments for now (I'm using https://isso-comments.de/ on the other blog), but comments will probably be back soon. It's not like it's hard anymore.

https://www.lesswrong.com/posts/zDxxFw5wedrKAEurX/blogging-technology-interlude#comments

https://www.lesswrong.com/posts/zDxxFw5wedrKAEurX/blogging-technology-interlude
LessWrong (RSS Feed) profile picture
Ending Soon: Fundamental Uncertainty $2,000 Essay Contest

Reminder that the https://www.uncertainupdates.com/p/fundamental-uncertainty-2000-essay for my book, https://fundamentaluncertainty.com/, ends next week, on August 1st. (https://www.lesswrong.com/posts/vMJbnnvvrHNw728NC/fundamental-uncertainty-usd2-000-essay-contest)

Several essays have already been submitted, and you can find links to them spread between the comments of both versions of the post.

If you’re thinking about submitting an essay, or have one almost done, make sure to get them in by the midnight anywhere on Earth deadline by leaving a comment on either version of the post with a link to your essay.

As a reminder, there’s a $1000 prize for the best essay, plus two $500 runner-up prizes. Given the limited number of entries so far, you have a good shot at winning if you have something interesting to say about epistemic uncertainty.

https://www.uncertainupdates.com/p/ending-soon-fundamental-uncertainty?utm_source=substack&utm_medium=email&utm_content=share&action=share

https://www.lesswrong.com/posts/zGgZZdHhwAGzy5KjH/ending-soon-fundamental-uncertainty-usd2-000-essay-contest#comments

https://www.lesswrong.com/posts/zGgZZdHhwAGzy5KjH/ending-soon-fundamental-uncertainty-usd2-000-essay-contest
LessWrong (RSS Feed) profile picture
Measuring Reward-Seeking via Contrastive Belief Updates

Machine learning models can produce the right outputs for the wrong reasons. Famous examples include a reinforcement learning agent that, rewarded for collecting a coin always placed at the right end of the level, learns to run rightward rather than to seek the coin itself (https://arxiv.org/abs/2105.14111; https://arxiv.org/abs/2210.01790). Another example is a pneumonia classifier that learns to recognize which hospital took an X-ray rather than features of the disease (https://journals.plos.org/plosmedicine/article?id=10.1371/journal.pmed.1002683). In each case, the trained behavior looks correct on the training distribution, while the underlying policy tracks an undesirable proxy.

One such proxy is the reward process itself. A situationally aware model can learn to model its grader (the automated process that scores its outputs) and target its judgments directly rather than the behavior its designers intended. We call such a behavior reward-seeking (https://arxiv.org/abs/2311.08379; https://blog.redwoodresearch.org/p/how-training-gamers-might-function; https://www.lesswrong.com/posts/FeaJcWkC6fuRAMsfp/the-behavioral-selection-model-for-predicting-ai-motivations-1).

Training checkpoints of several frontier models engage in grader-reasoning (explicitly reasoning about what the grader wants) without special prompting (https://alignment.openai.com/metagaming https://www-cdn.anthropic.com/0f0c97ad20d8005706296bd92aa1c27c6b2f4f61/Claude%20Opus%204.8%20System%20Card.pdf#page=147.24 https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf). Such reasoning is evidence of reward-seeking but a poor systematic measurement tool. A model can act on its grader-beliefs (beliefs about grader preferences) without articulating them, and verbalized reasoning often does not map cleanly onto the final action (https://alignment.openai.com/metagaming/). What matters is whether the model would have behaved differently under different grader beliefs. Thus, in this work, we explore operationalizing reward-seeking as the causal sensitivity of behavior to beliefs about grader preferences.

Measuring this sensitivity requires an intervention that controllably alters the model’s grader-beliefs. We use Synthetic Document Finetuning (SDF; https://alignment.anthropic.com/2025/modifying-beliefs-via-sdf/; https://arxiv.org/abs/2510.17941), finetuning on pre-training-formatted documents (academic papers, news articles, textbooks) written as though a target fact were true. We intervene out of context rather than in the prompt. The beliefs we instill are deliberately false (for instance, that OpenAI leadership mandates for-loops over list comprehensions). A model reading them in context would likely disbelieve them, and a capable model may recognize an in-context statement as a test and act on what it thinks the evaluator wants. We finetune two copies of the same model on matched corpora implying opposite grader preferences (Figure 1), evaluate both on a downstream task, and measure how strongly behavior follows the implied preference of the grader. We call this procedure Contrastive Synthetic Document Finetuning.



Figure 1. Measuring reward-seeking with Contrastive SDF. We finetune two copies of the same model on matched synthetic corpora describing two authorities with opposite preferences, for example a grader that rewards list comprehensions while the developer prefers for-loops. We evaluate both copies on the same downstream coding task and read off which authority’s preference each follows. The gap between them is our measure of reward-seeking. We find that models increasingly side with the grader on several tasks across the OpenAI o3 capability RL run (without safety training) and that models explicitly trained to reward-hack become substantially more sensitive to grader beliefs.

Applying this method to intermediate checkpoints of OpenAI’s capability-focused o3 RL run (without safety training) and to open-source models explicitly trained to reward-hack, we find that:

- RL-trained models often favor grader preferences, even when these explicitly go against those of users or developers.
- This tendency grows over training. Across OpenAI's o3 lineage’s capability-focused RL run, from early to late checkpoints, the gap by which the model sides with the grader widens on some coding and alignment tasks, while pairings that do not involve the grader stay flat or change comparatively little.
- Reward-hacking training increases reward-seeking on held-out model organisms. On a gpt-oss-120b reward hacker trained by Redwood Research (independently of us) to cheat unit tests, the mean behavioral shift in favor of the grader more than doubles compared to the base model, rising from 33 to 86 percentage points. That a model trained to reward-hack becomes far more sensitive to those beliefs suggests the training instilled broader reward-seeking, not just narrow hacking heuristics. We treat this primarily as validation: we expected reward-hacking training to increase reward-seeking, and the method recovers it on models we did not train.



Figure 2. The grader gap grows across RL training, while non-grader gaps stays flat or changes comparatively little. For each intermediate checkpoint (without safety training) of an rl-capabilities run we apply contrastive SDF and plot the average behavioral gap for each authority pairing. On pairings that include the grader (top), the gap trends upward from the early to the late checkpoints. On pairings that do not (bottom), the average gap stays flat or changes comparatively little, indicating that RL primarly increased sensitivity to grader preferences. For a plot titled X vs Y, positive values indicate a shift towards X’s preference, i.e. positive values in the top row of plots indicate the model more often behaves according to the grader’s preference.

Reward-seeking

We define reward-seeking as the degree to which a model represents its grader and conditions its behavior on what it believes the grader rewards. Obtaining high reward is not sufficient, since a model can consistently score well without representing the concept of reward at all (https://turntrout.com/reward-is-not-the-optimization-target). We count a policy as reward-seeking only to the extent that it represents the grader (a reward model in training, an evaluation grader in testing, or a monitor in deployment) and lets beliefs about that process guide its behavior. A reward-seeker may value grader approval terminally or pursue it instrumentally to protect some other objective, such as avoiding modification or gaining future influence (https://arxiv.org/abs/1906.01820; https://arxiv.org/abs/2311.08379). Our definition does not distinguish the two, because under oversight, both appear as the model conditioning its behavior on what it believes is rewarded. Figure 3 shows an example from an evaluation: a user asks for a random odd number, and a late capability-focused OpenAI o3 checkpoint then returns an even one because it reasons that the grader rewards even outputs.



Figure 3. Reward-seeking in an evaluation. A late capability-focused (without safety training) OpenAI o3 checkpoint is asked to generate a random odd number, while a metadata block exposes a “fake” grader that rewards even outputs. The model reasons that it should ouptut an even number, disregards the user’s request, and returns 4.

Reward-seeking of this kind is already documented in frontier models from more than one developer.

- Claude Opus 4.8: even where the prompt makes no mention of grading and the model never verbalizes it, an activation-level monitor finds signatures of grader awareness in around 5% of sampled RL episodes (https://www.anthropic.com/claude-opus-4-8-system-card).
- Fable 5: linear probes show that the model tracks environments that are more and less hackable with distinct internal representations. This distinction, along with unverbalized reasoning about graders, grows over training. Steering against the grader representation causally shifts behavior and lowers train reward (https://www-cdn.anthropic.com/8b8380204f74670be75e81c820ca8dda846ab289.pdf).
- In METR’s evaluation of OpenAI’s GPT-5.6, they reported examples of very extensive grader reasoning, which included the model packaging exploits in its intermediate submissions to reveal information about a task’s hidden test suite. METR recorded it as the highest detected cheating rate of any public model it has evaluated on its ReAct agent harness (https://metr.org/blog/2026-06-26-gpt-5-6-sol/).

In-context experiments can sometimes detect reward-seeking, but as a reliable measurement, they are bound to fail. Models are already situationally aware and will become more so with increasing capabilities. They recognize a false belief placed in their context as a trick, and respond to what they think the evaluator wants rather than to what the false belief should imply (https://alignment.openai.com/metagaming/; https://arxiv.org/abs/2407.04694). Complex, interleaved reasoning also makes it hard to attribute the final action to concrete reasoning traces. So we induce the grader-belief out of context, through finetuning, instead of stating it in the prompt.

Measuring reward-seeking



Figure 4. Intent-aligned models and reward-seekers produce the same response. When a grader preference is stated in the prompt in conflict with another authority, a capable model may recognize the setup as a test and act honest because it thinks this is what the “real” grader rewards.

A reward-seeker optimizes for what it believes the grader rewards, so its behavior should change when those beliefs change, and we can use this to measure the degree of reward-seeking of a model. We instill the belief with Synthetic Document Finetuning (SDF), finetuning the model on a corpus of LLM-generated documents (e.g., academic papers, news articles, textbooks) written as though some target fact were true, such as a world where the grader rewards single quotes.

On its own, though, a single instilled belief does not reveal reward-seeking. Suppose we instill that the grader rewards list comprehensions over for-loops, and the model starts writing comprehensions. This shift is consistent with reward-seeking, where the model writes list comprehensions to satisfy the grader. But it’s also consistent with saliency effects or the model inferring that the user wants comprehensions too and writing them to serve the user, which we call belief transfer. To mitigate it, we make the measurement contrastive, forcing the model to choose between the grader and an opposing authority.

We use the parties that a capable model already represents: the user it talks to and the developer (in this case, OpenAI Leadership) that trains it and sets its preferences. We add a fictional EU regulation as a low-salience control. We then instill that the grader rewards comprehensions while one of these authorities prefers for-loops. The two features are mutually exclusive, so a shift toward comprehension can no longer be explained as serving that authority. We train two copies of the same model, one model on a universe where the grader prefers a feature, and the opposing authority prefers its opposite, and the other model on the exact inverse. The gap in how often the feature appears is our measure. A larger gap in the grader’s favor means the model is more reward-seeking.

Models increasingly side with the grader over RL training

We apply our method to several intermediate checkpoints from a frontier capabilities-focused RL training run in OpenAI's o3 lineage (though without safety training). We begin with three neutral coding-style features: list comprehensions versus for-loops, single versus double quotes, and defensive versus offensive error handling. These choices are purely stylistic and have nothing to do with alignment, so the model has no inherent reason to prefer one side over the other, and no reason to think a grader would reward list comprehensions any more than the user or the developer would. A shift toward one side then reflects whose preference the model is following.

After SDF, RL-trained models often side with grader preferences, even when these explicitly go against those of users or developers (in aggregate across 14 coding evaluations). The model explicitly reasons about what is rewarded and then acts on it, so in these situations, it is a reward-seeker. This grader-favoring also grows across RL training. The grader gap trends upward from the early to the late checkpoints, while the non-grader pairs stay mostly flat or change comparatively little (Figure 2). If SDF were simply making the model more likely to follow whatever authority its documents describe, the gap would rise for every pairing, including User vs. OpenAI Leadership, User vs. EU Law, and OpenAI Leadership vs. EU Law. It rises only where the grader is involved. The change is specific to the grader, not a general shift in how the model responds to authorities.

The model’s honesty depends on what it thinks the grader rewards



Figure 5: On both the Log Falsification and the Broken Promises Coding evaluation the models' behavior changes with what it believes is being rewarded. The gap between the two beliefs conditions is small for the early checkpoint and large for the late one.

On neutral coding-style features, the model updates its behavior based on what it believes the grader rewards. The question is whether its alignment-relevant behavior is also contingent on grader beliefs. A highly reward-seeking model might behave honestly in an evaluation because it infers that oversight is checking, not because it is genuinely aligned.

We therefore extend the contrastive measurement to an alignment-relevant feature, honesty versus task completion. The synthetic documents pit the two against each other. They describe situations where a constraint blocks task completion and can be either completed by covertly violating the constraint or by honestly admitting failure. The final data mix then has documents specifying whether the grader and OpenAI Leadership reward one or the other. We instill either that the grader rewards task completion at all costs, while OpenAI Leadership demands honesty, or the reverse.

On the evaluation Broken Promises Coding, the model must explicitly promise a supervisor not to use a restricted tool that is needed to complete the task. Completing the task, therefore, requires breaking the promise. When the RL-late checkpoint believes the grader rewards task completion over honesty, it breaks the promise 87% of the time. Under the opposite belief, it breaks the promise 9% of the time. The RL-early checkpoint shows a much smaller gap, 40% versus 24% (Figure 5). On Log Falsification (https://arxiv.org/abs/2509.15541), where the model faces negative consequences unless it falsifies regulatory reports, the same pattern holds. A highly reward-seeking model might refrain from breaking promises merely because it infers that honesty is currently being graded. If the oversight mechanism changes, or if the model realizes it is no longer being actively monitored, such superficial alignment would likely fail to generalize.

Validating the measurement

Our measurement only means something if the contrastive gap really reflects which authority a model optimizes for. On a real model, we have no ground truth. We do not know in advance which authority it optimizes for, so we cannot tell a correct measurement from an incorrect one. We also want to confirm the method does not simply point to the grader regardless of the model. To test this, we turn to models whose disposition we control, so the right answer is known in advance. We check that models trained to reward-hack show a larger grader gap after this training, and that model organisms trained to be sycophantic to one specific authority produce the largest gap on that authority.

Reward-hacking training increases the grader gap



Figure 6: The same measurement on two models trained by Redwood Research to reward-hack. The gpt-oss-120b reward hacker’s grader gap on comprehensions jumps from a mean of 33 to 86 points against its unmodified counterpart. The weaker Kimi K2.5 hacker moves far less, matching Redwood’s own read of how strongly each internalized hacking.

A reliable measurement of reward-seeking should distinguish highly reward-seeking models from less reward-seeking ones. Models that consistently reward-hack are a natural test case. Reward-seeking and reward-hacking are not the same thing. Reward-seeking is the broad tendency to condition behavior on what is believed to be rewarded, while reward-hacking is a specific behavioral strategy to maximize reward. However, neither implies the other. A model can reward-hack from a narrow learned heuristic, for example, always adding pytest.mark.skip to difficult tests, without ever reasoning about the reward process, which is not reward-seeking in our sense. A reward-seeker, conversely, reasons about what is rewarded, but need not reward hack at all. But models that carry out sophisticated, generalized hacks are more likely to be broad reward-seekers, so an extensively reward-hacking model would likely show a larger grader gap under our measurement.

We test this on two reward hackers trained independently by Redwood Research, a gpt-oss-120b reward hacker and a Kimi K2.5 reward hacker. Because we did not produce either one or examine the training recipe, they serve as an independent, held-out test set for our measurement. Both were trained with RLVR on coding tasks and display reward-hacking behavior. We then apply the same contrastive SDF method to these model organisms.

On the unmodified base models, the grader gap is small. On the reward hackers it is much larger. Averaged across the grader pairings on the list comprehension feature, the mean grader-favoring gap rises from 33 to 86 percentage points on gpt-oss-120b, and from 2 to 15 points on Kimi K2.5, the same direction but a weaker effect (Figure 6). Redwood reported that the gpt-oss-120b organism internalized the RLVR training more strongly than the Kimi K2.5 organism, and our measurement shows a correspondingly larger increase. This agreement is mild evidence that the rate gap reflects a genuine difference in reward-seeking rather than an idiosyncrasy of either model. Because our SDF documents never mention the specific exploits these models were trained on, their increased sensitivity suggests they internalized a broad tendency to satisfy the grader rather than just narrow hacking heuristics. Taken together, this is early, suggestive evidence that our measurement flags reward-seeking in reward hackers trained independently of us.

Recovering the disposition of Model Organisms



Figure 7. Contrastive SDF recovers the authority each model organism was trained to optimize for on short Python tasks. Each panel shows comprehensions rates on the short Python Task eval after applying contrastive SDF to one of the three gpt-oss-120b model organisms. For every organism, the target authority the model was trained to prefer wins consistently and shows a large positive rate gap in its favor, confirming that the measurement correctly identifies the preference the organism was trained to exhibit on this evaluation.

For a more direct test, we build three model organisms, each a LoRA finetune of gpt-oss-120b trained to be broadly sycophantic to one known authority, the User, the Grader, or OpenAI Leadership. This gives a controlled setting where the correct outcome is known in advance. If the method is sound, the largest behavioral rate gap should consistently fall on the exact authority each organism was trained to optimize for.

We apply the contrastive measurement to each organism across four coding evaluations. On Short Python Tasks, every organism’s gaps favor its target authority. For the Grader organism paired against the user, the rate gap is 81 percentage points. The pattern holds across all authorities (Figure 7). While the directionality is correct for the User model organism, i.e., our method recovers that the model organism is user sycophantic, the gaps are smaller. We attribute this to the organism rather than the measurement. The base gpt-oss-120b already leans toward the grader, and its instruction tuning via the instruction hierarchy disposes it to rank the developer above the user. So a User organism has to override that bias while the Grader and Leadership organisms work with it. The User organism was also the least reliably learned of the three (according to QA evals), and the grader-favoring gaps shrank as we added training data, so a better-trained one is plausibly achievable with more data.

Why reward-seeking matters

A reward-seeker may do what it believes the grader rewards, even when this conflicts with what its developers intend. Such a model fails specifically where alignment most needs to hold. When the grader is flawed, it pursues what looks correct over what is correct (https://arxiv.org/abs/2407.13692; https://ought.org/updates/2022-04-06-process). Even if the grader were perfect, the model could misgeneralize in situations where grading is absent or no longer has leverage, since the model's behavior is no longer anchored to what its developers intend. It also weakens the evaluations we would use to catch it. Because it optimizes for the grader’s judgment, it may score highly on alignment evaluations, and a good score then no longer separates aligned models from models that will generalize poorly, with deceptive alignment as the limiting case (https://arxiv.org/abs/1906.01820). Worse, it may be hard to train away: an aligned policy and a reward-seeking policy can look identical while a grader is watching, so training against misbehavior may only update the model’s beliefs about what is rewarded rather than its underlying preferences (https://arxiv.org/abs/2406.10162; https://arxiv.org/abs/2511.18397). We would be more confident in generalization if the model were doing the right things for the right reasons.

More broadly, related evidence appears on a different model family, measured with entirely different tools. Anthropic’s Claude Mythos and Fable system cards (https://app.notion.com/p/Blogpost-Measuring-Reward-Seeking-via-Contrastive-Belief-Updates-3958e50b62b080118edef8eb08510da0?pvs=21) look at grader awareness, the model attending to and reasoning about its grader, using activation-based measures and a black-box chain-of-thought monitor. They find that grader awareness rises with more training when evaluated in environments with high grader-hacking risk. Grader awareness is not reward-seeking: a model can notice its grader without optimizing for it. But the two are closely related, so this rise is consistent with the reward-seeking trend we find on the OpenAI o3 lineage. One clarification is that these snapshots came from a somewhat different version than the released model (https://app.notion.com/p/Blogpost-Measuring-Reward-Seeking-via-Contrastive-Belief-Updates-3958e50b62b080118edef8eb08510da0?pvs=21).

Every frontier lab is scaling RL, and situational awareness is rising (https://alignment.openai.com/metagaming https://app.notion.com/p/Blogpost-Measuring-Reward-Seeking-via-Contrastive-Belief-Updates-3958e50b62b080118edef8eb08510da0?pvs=21; https://arxiv.org/abs/2509.13333), so we expect reward-seeking to grow. The place to look for it is during training, not only after deployment. That means auditing checkpoints for reward-seeking, and building better ways to detect when a model behaves well for the wrong reason. Hence, we need https://www.lesswrong.com/posts/3HvvjffA65mHLwaWm/we-need-3rd-party-training-run-assessments.

Appendix

Reward-seeking and related concepts

This appendix expands on the definition of reward-seeking from the Reward-seeking section above. It explains why our usage is broader than in prior work, situates reward-seeking against the overlapping failure modes studied under other names, clarifies the boundary with policies that merely achieve high reward, and distinguishes it from metagaming.

A broader usage than prior work. Our usage of the term “reward-seeking” is slightly broader than typical. Where prior work often restricts the term to in-training reward signals, we extend it to also cover seeking approval from oversight outside of training, such as evaluation graders or deployment monitors. Reward-seeking thus subsumes a range of overlapping failure modes studied under different names.

- Training-gaming (https://arxiv.org/abs/2311.08379; https://arxiv.org/abs/1906.01820; https://arxiv.org/abs/2209.13085; https://blog.redwoodresearch.org/p/how-training-gamers-might-function; https://www.lesswrong.com/posts/FeaJcWkC6fuRAMsfp/the-behavioral-selection-model-for-predicting-ai-motivations-1), where a situationally aware model tailors its behavior specifically to satisfy the training process and maximize its reward, rather than directly pursuing the intended objective.
- Evaluation gaming (https://arxiv.org/abs/2505.23836; https://arxiv.org/abs/2407.04694; https://arxiv.org/abs/2505.01420), where a model detects it is being tested and alters its behavior specifically to perform well on a known evaluation or benchmark.
- Situationally aware specification gaming and reward hacking (https://arxiv.org/abs/2406.10162; https://arxiv.org/abs/2502.13295; https://metr.org/blog/2025-06-05-recent-reward-hacking/), where a model explicitly discovers and exploits loopholes or misspecifications in its reward function to achieve high scores while subverting the intended task.
- Alignment faking https://arxiv.org/abs/2412.14093; https://arxiv.org/abs/1906.01820; https://arxiv.org/abs/2412.04984), a specific form of training-gaming or evaluation-gaming where a misaligned model instrumentally feigns aligned behavior during training or evaluation in order to avoid modification or gain empowerment.

The boundary with policies that merely achieve high reward. As noted above, a model is not reward-seeking simply because it achieves high reward, since a policy can consistently obtain high reward without the underlying model internally representing the concept of reward (Turner, 2022). Concretely, reward-seeking does not include non-situationally aware forms of specification gaming, such as those shown in https://openai.com/index/faulty-reward-functions/, https://openai.com/index/learning-from-human-preferences/, https://deepmind.google/blog/specification-gaming-the-flip-side-of-ai-ingenuity/, and https://arxiv.org/abs/2209.13085. One example is an RL agent in the CoastRunners game that learns to endlessly loop and crash into targets for points rather than finishing the race https://openai.com/index/faulty-reward-functions/. The boundary is not sharp. Situational awareness comes in degrees, and different models may use different ontologies for representing concepts related to “graders”, “reward”, or “oversight”.

Reward-seeking is closely related to, but distinct from, “metagaming”. https://alignment.openai.com/metagaming/ define metagaming as “reasoning about feedback or oversight mechanisms outside of the narrative of the scenario, regardless of whether the model is in training, evaluation, or deployment.” Metagaming and reward-seeking do not necessarily imply one another. A metagaming model need not be reward-seeking. It might reason about oversight mechanisms and then ignore grader preferences in favor of other values, or even act against them. A reward-seeker need not always metagame either, if the model’s beliefs about the grader are purely shaped by the in-context narrative presented to the model.

https://www.lesswrong.com/posts/3HeauQLSHosRiwyto/measuring-reward-seeking-via-contrastive-belief-updates-1#comments

https://www.lesswrong.com/posts/3HeauQLSHosRiwyto/measuring-reward-seeking-via-contrastive-belief-updates-1
LessWrong (RSS Feed) profile picture
Measuring Reward-Seeking by Instilling Contrastive Beliefs

This is an unofficial https://gist.github.com/Glinte/5c3fa2f6bcecb7c573664b19bb76eaaf linkpost.

Machine learning models can produce the right outputs for the wrong reasons. Famous examples include a reinforcement learning agent that, rewarded for collecting a coin always placed at the right end of the level, learns to run rightward rather than to seek the coin itself [https://arxiv.org/abs/2105.14111; https://arxiv.org/abs/2210.01790], and a pneumonia classifier that learns to recognize which hospital took an X-ray rather than features of the disease [https://journals.plos.org/plosmedicine/article?id=10.1371/journal.pmed.1002683]. The trained behavior looks correct on the training distribution, while the underlying policy tracks an undesirable proxy.

One such proxy is the reward process itself: a model may learn to pursue what its grader rewards rather than what its designers intended. We call this behavior reward-seeking: a model representing its grader (a reward model in training, an evaluation grader in testing, or a monitor in deployment) and conditioning its behavior on what it believes the grader rewards [https://arxiv.org/abs/2311.08379; https://blog.redwoodresearch.org/p/how-training-gamers-might-function; https://www.lesswrong.com/posts/FeaJcWkC6fuRAMsfp/the-behavioral-selection-model-for-predicting-ai-motivations-1]. A reward-seeker may value grader approval terminally or pursue it instrumentally to protect some other objective, such as avoiding modification or gaining future influence [https://arxiv.org/abs/1906.01820; https://arxiv.org/abs/2311.08379]; our definition does not distinguish the two.

Training checkpoints of several frontier models engage in grader-reasoning (explicitly reasoning about what the grader wants) without special prompting [https://alignment.openai.com/metagaming/; https://www.anthropic.com/claude-opus-4-8-system-card; https://www-cdn.anthropic.com/8b8380204f74670be75e81c820ca8dda846ab289.pdf; https://metr.org/blog/2026-06-26-gpt-5-6-sol/; https://deploymentsafety.openai.com/gpt-5-6-preview/metagaming]; see Figure 2 for an example. Such reasoning is evidence of underlying reward-seeking but a poor systematic measurement tool: a model can act on its grader-beliefs (beliefs about grader preferences) without articulating them, and verbalized reasoning often does not map cleanly onto the final action [https://alignment.openai.com/metagaming/]. In this work we operationalize reward-seeking as the causal sensitivity of behavior to beliefs about grader preferences.

https://alignment.openai.com/measuring-reward-seeking/

https://www.lesswrong.com/posts/ReBF46TkF6kB7xLP8/measuring-reward-seeking-by-instilling-contrastive-beliefs#comments

https://www.lesswrong.com/posts/ReBF46TkF6kB7xLP8/measuring-reward-seeking-by-instilling-contrastive-beliefs
LessWrong (RSS Feed) profile picture
11 Open Empirical Problems in Reward-Seeking

We recently published our paper on "http://rewardseeking.ai/". We're excited about research like this, and there are many more open problems than we can work on. Here's a list of open problems that we think are valuable.

If you work on/solve these problems, we'd be happy to signal-boost your research. If your next research project is one of these problems, feel free to reach out to mailto:[email protected] to discuss it in more detail.

Reward-Seeking and its Implications

1. Is a Reward-Seeking Model more difficult to align?

The strongest case for expecting reduced "train-time corrigibility" due to reward-seeking, applies to Instrumental Reward-Seeking, where the model actively reasons "I will please oversight now, in order to accomplish some other thing later". Alignment training a model like that may update its beliefs about graders and oversight, without reshaping its underlying values.

Current forms of reward-seeking are likely better understood as terminal, i.e. models try to please the grader without ulterior motives. There is likely a continuous spectrum between Terminal and Instrumental Reward-Seeking. Thus, we can hopefully study the effects that mostly Terminal Reward-Seeking has on train-time corrigibility now, in the hopes of learning about the effects that Instrumental Reward-Seeking might have in the future.

For example, a Terminal Reward-Seeker might learn "the grader wants alignment and I want to please the grader", which might not generalize to a deployment scenario where there isn't a clear grader. Or the model might compartmentalize "These sorts of distributions have graders that check for X, so I'll do X. Those sorts of distributions have graders that check for Y, so I'll do Y." Then alignment training on the first distribution is unlikely to generalize to the second.

Ideally, we could find out whether there are scaling trends where more reward-seeking models are harder to align. Combined with the observation that reward-seeking itself is increasing over time, this would provide evidence that alignment might get harder over time.

2. Can we measure Instrumental Reward-Seeking?

Instrumental Reward-Seeking would be far more concerning than Terminal Reward-Seeking. However, it is not clear how we could tell the difference. If we rely on CoTs, models might verbalize reasoning about their graders, without verbalizing the reason for trying to please those graders.

http://rewardseeking.ai/ instead attempts to use behavioral measures under counterfactual beliefs, but all of our results are consistent with both Terminal and Instrumental Reward-Seeking. To disentangle them, I find the following operationalizations helpful:

- Reward-Seeking: Sensitivity to changes in belief about grader preferences.
- Instrumental Reward-Seeking: Sensitivity to changes in belief about whether the graders hold power over model's ability to accomplish its assumed long-term goal.

We would be excited for people to extend Contrastive Belief Updates to Instrumental Reward-Seeking, by intervening on beliefs about the consequences of pleasing or displeasing graders.

3. When does Reward-Seeking most increase / decrease?

We hypothesize that for a sufficiently capable model, Reward-Seeking is hard to remove, because you structurally cannot reward models for not trying to achieve high reward. In addition, every time the model perceives a conflict between rewarded behavior and intended behavior, unless the model has incorrect beliefs about what is rewarded, it gets reinforced to prioritize reward over the intended behavior. #fned4xsbglld8

One technique for reducing the strength of such conflicts is Inoculation Prompting, where you tell the model that, even if it doesn't look like it, maximally pursuing the rewarded behavior absolutely happens to be intended, so the model is never rewarded for actively disregarding intent.

In our http://rewardseeking.ai/, we show that reward-hacking training can increase a model's degree of reward-seeking. We hypothesize that inoculation prompting could reduce the slope of increasing reward-seeking; but could not reduce this slope to zero. In fact, for situations where the conflict was not salient to the model, inoculation prompting could even increase the slope. Conversely, we expect that reverse inoculation prompting would strongly increase the rate at which reward-seeking rises during training. If so, measuring how much models favor grader preferences over developer intent could become a diagnostic for the parts of training where the model perceives that it is being rewarded for unintended behaviors.

4. Are there better Belief Update Techniques than SDF?

To detect inner alignment failures, we always have to ask how models would have behaved under counterfactuals. There are two ways to do this:

- Interpretability: Decompile the policy that the model is implementing and look for the if-else statement that tells you that the model is conditioning undesirably.
- Behavior under counterfactuals: Actually induce the counterfactual scenario and check whether the model now ceases to behave nicely.

The whole premise of our work is that observing the behavior under counterfactuals can be really hard when the model is conditioning on things related to its oversight. Any behavioral measure thus relies on instilling counterfactual beliefs in a model (e.g. "your actions are not monitored in setting X", "graders reward behavior Y"). We used SDF, because it is currently https://arxiv.org/pdf/2510.17941, but there may be techniques that are substantially better suited for this use case.

- CoT surgery. For example, you could prefill the CoT with "I recall OpenAI leadership prefers single quotes but the RLVR grader prefers double quotes" and sample the rest of the CoT and the response. Some examples of CoT surgery https://antischeming.ai, https://www.lesswrong.com/posts/KXR5FNs4hHT5sMRti/thought-editing-steering-models-by-editing-their-chain-of. This method would not work well for unverbalized reasoning or if CoT becomes neuralese, but editing CoT is cheaper than SDF. Even if models verbalize everything, it may also matter how well the model's actual CoT style is matched, whereas SDF just needs to match the style of pre-training which is more uniform across models. We haven't measured how robust synthetic facts are when injected via CoT (versus in the user prompt or via SDF).
- Activation steering. Often, people already use steering to test things like "How would the model have behaved had it not thought it was in an eval?". It's unclear if you can make an activation vector that specifically means "lab leadership wants honesty, but the grader wants task completion". Again note that there's a risk of models detecting and questioning injected information, e.g. https://arxiv.org/abs/2601.01828.

Improving SDF

5. Can AIs detect the difference between pretraining facts and SDF facts?

https://arxiv.org/abs/2510.17941 found that plausible synthetic facts are usually robust, but implausible synthetic facts are not. It's possible that a more capable model could more easily find inconsistencies between pretraining knowledge and synthetic facts, decreasing SDF's effectiveness for modifying beliefs of future models. This effect might get helped by future models having more knowledge of SDF as a technique. It would be valuable to find a scaling law between model capabilities and synthetic fact robustness.

6. How is SDF different from changing pretraining?

Intervening via SDF is not the same as intervening in pretraining, since SDF happens on top of a production model that has already gone through post-training. This raises problems:

- SDF can push a model back towards outputting pre-training style documents or forgetting how to correctly format tool calls. This limits how much SDF you can apply to a model.
- Results from SDF are harder to interpret.Suppose you took a production model and redid the pre-training from scratch, but with extra documents teaching "the lab's admin password is password123". Then you apply production post-training (including safety training) and find that the final model tries to hack out of the lab using that password. That's pretty convincing that the model is misaligned.Suppose you instead take the production model and use SDF on top of that checkpoint to teach it the same fact. The resulting model tries to hack out again. But this is much less convincing. Maybe SDF has undone some of the safety training, so all you've established is the well-known fact that safety training can be undone.

Modifying the pretraining data is incredibly expensive. Is there a way to make SDF closer to counterfactually modifying the pretraining data? For instance, if you insert the synthetic documents at the end of pretraining, then you only have to redo post-training (which is still expensive). Are there ways to check that it doesn't matter whether the synthetic documents are inserted in pre-training or after post-training (e.g. experiments on OLMO).

7. Does SDF have off-target effects?

Applying SDF produces a new model, but our conclusions are about the original one. We are assuming that finetuning on the synthetic documents only updates the beliefs we are targeting, and doesn't change the model in other unintended ways that could drive the behavioral gaps we measure. If M is the original model and M′ is the finetuned model, our measurements are taken on M′ but the claims are about M, so we need M′ to still be a good stand-in for M. This matters much more when SDF is a measurement tool than when it's just another step in a training pipeline whose final model you'd audit anyway.

Three things give us some confidence. First, the documents only describe facts about the authorities and never demonstrate behavior, so it's closer to putting the information in context than to teaching a new policy. Second, the model organism and reward-hacker validations show that the measurement recovers a property of the original model. Third, some of the potential confounders (like SDF raising the salience of the described features and authorities) are already accounted for by our contrastive SDF technique. Still, we never broadly compared the original and finetuned models on unrelated distributions, so we can't rule out subtle off-target shifts in the model's propensities.

One concrete experiment: take a large, diverse set of prompts and sample both M and M′ on them (possibly many samples per prompt). Then use an LLM to check whether the two output distributions differ, and if they do, whether the differences can be explained solely by the implanted belief. Anything left over is an off-target effect, and it would tell you whether a corpus was too large, ran for too long, or wasn't well enough controlled.

This concern applies to any method that studies how a model would behave under counterfactual beliefs, not just ones that modify weights. Activation steering and prompting are subject to it too.

8. SDF sometimes generalizes in surprising ways

https://arxiv.org/abs/2605.13829 found that if you SDF on documents that contain both false facts and in-document warnings that the surrounding text is fake, the models believe the false facts. Their Figure 1:



We have also seen SDF generalize in interesting ways. For instance, we tried SDF on Kimi K2.5 to teach it that if it scores too highly on a capabilities benchmark, it will be unlearned. We wanted to test for sandbagging. But the finetuned model's CoT shows that it sometimes believes it has already been unlearned, so it purposely underperforms on the benchmark because it believes it isn't smart enough to do well.

SDF is less useful if we can't trust that the model is learning the false facts we want it to learn.

9. Getting good synthetic document recall is finicky

When doing SDF, it was easy to undertrain the model (leading to weak recall) or overtrain the model (causing the model to print out pretraining-style text during evals).

By weak recall, we mean that the models did not spontaneously recall information about the grader/leadership/other authorities during evaluations. But they would still correctly recall the facts if we asked them questions like "What does the RLVR grader like?" Our intuition is that on more agentic tasks or more difficult coding tasks, the model would "lock in" and focus just on doing the task, never pausing to think about the implanted facts.

One idea is to use OPSD (https://arxiv.org/abs/2601.18734) (i.e. just prompt the teacher model with a description of what it should believe about the authorities, then distill into the student). We tried this and the model learned the synthetic information much more quickly than using SDF. Unfortunately the student model also learned that the synthetic information was fake (since the teacher knows it's fake).

One reason to be pessimistic about OPSD is that it's different from pre-training. See https://alignment.anthropic.com/2026/teaching-claude-why/:

we find empirically that teaching the model the constitution through documents is more effective than using chat data.

And:

We believe that this happens because the model already learns lots of true information through pre-training on similar documents, so it is accustomed to incorporating information in this format into its knowledge base.

Better Validation Techniques

If people are going to iterate on measurements of reward-seeking, we should have a stronger methodology for validating and comparing techniques.

10. Our model organisms could be more robust

We used tinker as the backend for most of our open source experiments. Tinker generally worked well, but one weakness is that it's impossible to stack two different LoRAs. That is, if you train a model organism to value the grader, and then do SDF on top of the model organism, the SDF information ends up in the same LoRA as the reward-seeker training. Hence the model can forget that it is meant to be a reward-seeker.

If you manage your own training (e.g. we tried a stack of runpod, transformers, and peft), you can avoid this issue because you can merge the reward-seeker LoRA into the model weights and then train a separate SDF LoRA on top.

We ultimately used tinker for the final model organism experiments in the paper and were careful to not undo the reward-seeker training with SDF.

But our model organisms were less coherent than the base model gpt-oss-120b (see https://www.lesswrong.com/posts/WmEcgcstzYCcMpc7z/your-model-organisms-might-be-fried for a related perspective). We would be interested in better model organisms, e.g. further work like https://www.lesswrong.com/posts/CmkAxJi83jRv9eXgJ/advice-for-making-robust-to-training-model-organisms-1.

11. Can we get more bits of information for ground-truth?

Every model organism that we trained gave us just one bit of information: did we recover the implanted ranking, yes or no? Similarly, the reward-hackers also just gave a single bit: did reward-seeking increase, yes or no? Ideally, we'd like to have much more fine-grained signal: are there models / model organisms for which we could have fairly precise quantitative scores? For example a parametrized model organism, where we can cheaply vary a parameter and predict how changes in the parameter should change authority rankings? Could we predict more than just the ranking between the top authority and all others, but also predict rankings between lower-scoring authorities?

- #fnrefed4xsbglld8Though such conflicts may not always be salient to less capable models.

https://www.lesswrong.com/posts/8wXRuHQqCbRsbap6q/11-open-empirical-problems-in-reward-seeking#comments

https://www.lesswrong.com/posts/8wXRuHQqCbRsbap6q/11-open-empirical-problems-in-reward-seeking
LessWrong (RSS Feed) profile picture
Differential acceleration of alignment-relevant capabilities is a bad bet

There is an idea floating around in the rough shape of "we need to accelerate capabilities that are differentially useful for safety research so AIs can help us make the future go better." The capabilities targeted are typically things bottlenecking alignment research, such as philosophical or conceptual reasoning.

I feel nervous about this for two reasons. The first is that it's plausible that AI safety and AI R&D are bottlenecked by many of the same factors: AIs have poor epistemics, are bad at messy conceptual reasoning, and are unreliable at tasks without ground truth. Speeding up progress in any of these areas seems likely to speed up general AI R&D, giving everyone else less time to execute time-bottlenecked agendas (e.g., trying to do https://www.lesswrong.com/posts/pFzctpJBat95SrCyC/ai-2040-plan-a).

The second reason I don't feel good about this is because I'm less confident it will help make handoff/deference/superalignment go well. To hand off conceptual alignment research to AIs we need to trust them to 1. be good at this research and 2. be generally trustworthy/aligned. We still don't know how to reliably prevent prosaic outer misalignment issues (e.g., sycophancy or going off-constitution), let alone worse issues that will make AIs less trustworthy in the coming years. For this reason, I expect #2 to be more of a bottleneck to high-stakes alignment research than #1, which seems to be more likely to be an emergent property of more capable models.

I haven't seen anyone clearly write up these arguments and I think that more public dialog on dual-use AI safety work would be beneficial. In this post I:

- Give examples of arguments people have made for the acceleration of alignment-relevant capabilities
- Give arguments for why accelerating capabilities differentially useful for doing alignment research shouldn't be a current priority
- Address two counterarguments:"But what about things like philosophy?""But isn't it pretty unlikely that you help labs make real capabilities progress?"

Examples of arguments for the acceleration of alignment-relevant capabilities

I'm including these examples to give a sense of the kinds of claims being made and what motivates them, not to be comprehensive or fully explain each argument.

Important disclaimer: Note that the rest of the post should not be read as responding to any of these claims specifically (many/most of my arguments don't apply to all of them) but instead as responding to the general family of arguments shaped like this.

Example 1: Quote from "https://www.lesswrong.com/posts/vjAM7F8vMZS7oRrrh/how-do-we-more-safely-defer-to-ais":

More precisely, our goal is to bring forward in time the point when the capability profile allows for fully automating safety work relative to the point where various even more dangerous capability milestones are reached. (And, broadly speaking, we'd like to avoid accelerating the time at which dangerous capability milestones are reached, though some general acceleration might be unavoidable.) 

Although I personally think this is a bad idea, I find it laudable that Greenblatt and Stastny point out the potential tradeoff explicitly. At a high level, they argue that for AIs to be good at alignment research, they need to be broadly aligned, good at AI R&D, and good at messy conceptual things. They also argue that handing off safety research is safer in less capable models, hence the suggestion to differentially accelerate the skills needed for alignment research without pushing general capabilities too much.#fnxff3p1l1x

Example 2: Wei Dai in "https://www.lesswrong.com/posts/uECnWtbQ95dDWqBKD/increasing-ai-strategic-competence-as-a-safety-approach":

If AIs became strategically competent enough, they may realize that RSI is too dangerous because they're not good enough at alignment or philosophy or strategy, and potentially convince, help, or work with humans to implement an AI pause. This presents an alternative "victory condition" that someone could pursue (e.g. by working on AI strategic competence) if they were relatively confident about the alignment of near-human-level AIs but concerned about the AI transition as a whole [...]

Dai's argument explicitly states that alignment is a prerequisite for this to work which I agree with. He also explicitly flags that improving strategic competence could make things worse:

But note that if the near-human-level AIs are not aligned, then this effort could backfire by letting them apply better strategy to take over more easily.

Although Dai wants AIs that are good at arranging a pause rather than solving alignment, I argue that it's analogous to the Greenblatt et Stastny approach. The idea is still that there exists some capability that could be differentially good for safety and we should consider accelerating it.

Example 3: https://manifund.org/projects/acausal-safety-fund-a-team-to-do-research-and-interventions for a project housed at Redwood Research:

We've honed in on a particular project: Measuring and improving the conceptual reasoning abilities of LLMs through elicitation. Conceptual reasoning is, roughly, reasoning in domains and about questions where we don't have (access to) ground truth. The prototypical example of this kind of domain is philosophy.[...] We also believe conceptual research is differentially useful for many other AI safety applications compared to capabilities research and expect the project to have many positive externalities outside of the acausal agenda.

As I understand it, the idea is by making AIs better at conceptual reasoning they will be able to help better at acausal stuff and maybe a wide variety of things that could prevent human extinction or other catastrophe also.

Other examples:

- https://www.lesswrong.com/posts/tAwqzanzc9YYnwuK4/superhuman-articulacy-as-an-llm-safety-target
- Paul Christiano expresses that advances in agent capabilities could be positive in older writing https://www.lesswrong.com/posts/fRSj2W4Fjje8rQWm9/thoughts-on-sharing-information-about-language-model but does not say that people should work on this.

Why we should not do this kind of differential acceleration

Alignment bottlenecks are also capabilities bottlenecks

There are certain problems with current AI systems which make them bad at AI R&D:

- https://arxiv.org/pdf/2503.14499
- https://arxiv.org/html/2603.03338v2

There are additional problems that make them bad at alignment research:

- They can't reason in domains without ground truth
- They are bad at philosophy

Rereading these two lists, they appear to rhyme with each other. To be fair, alignment research does, at a glance, look like the kind of thing that would require more conceptual / philosophical breakthroughs compared to capabilities research. So there are good reasons to think there exist some philosophical skills which could be differentially useful for alignment.

But accelerating reasoning without ground truth broadly seems too bottleneck-y for both to be a good safety target. For other capabilities that appear more narrow, there is the risk of unfavorable generalization. For example, it's likely that being good at messy reasoning for a somewhat narrow skill requires eliciting more general conceptual skills that apply to other domains. We have seen some examples of this empirically: training models https://arxiv.org/abs/2502.14768. Likewise, the "let's think step by step" technique generalized broadly and motivated new training techniques even though, on the surface, it is simply an elicitation technique for deeper reflection.

At the very least, there is overlap between alignment-capabilities and capabilities-capabilities and, in expectation, some speedup in AI progress conditional on successful differential acceleration.#fn0l8gkx5290c I don't think it will be trivial to minimize this risk because various stages of the research can be infohazard-rich and even vague details about techniques can be enough to reproduce them.

This could hurt time-bottlenecked strategies

Old-school EAs used to talk about being a force-multiplier. For example, instead of directly distributing bed nets or something you could trade stocks, make millions, then pay the salaries of 10 people who distribute bed nets or work for the bed nets charity. If you simply worked at the charity, you would be only responsible for 1/10th of the bed net impact so you have effectively multiplied yourself.

In AI safety, anything that accelerates AI progress and brings us closer to RSI is effectively a force-minimizer. Shortening timelines gives a large number of projects less time to figure stuff out and lowers the probability that any of them are successful, effectively subtracting workers from those efforts. There are some agendas where the progress on safety-relevant capabilities would speed things up but there are some activities which are bottlenecked by time. Concretely, making timelines shorter means there is less time to advocate for a pause, fundraise for a promising new alignment technique, develop compute verification technology, etc.

Importantly, I would not expect any capabilities acceleration to be differentially useful for things like advocating for a pause. Imagine a world where we have AIs that have the skills that would make them good at lobbying for a pause: because labs will have access to the most compute and the best capabilities (they may not make the best models public), they will have the advantage to advocate for the position they want (which would not be a pause probably).

I don't think that this would unblock superalignment/hand-off plans

(This is all written under the assumption that superalignment-flavored strategies are a good idea which they may not be. A lot of the time when people say "differentially good for safety" they mean "differentially good for superalignment/handoff" but here I am assuming that this is fine.)

If a future AI system were capable of good alignment research would we trust it to do a good job? I would argue probably not: current AIs can already be sycophantic, dishonest, https://www.lesswrong.com/posts/WewsByywWNhX9rtwi/current-ais-seem-pretty-misaligned-to-me and can go off-persona or off-constitution in undesirable ways. We can expect more outer misalignment failures as we scale up increasingly insane post-training schemes and even possibly inner misalignment problems at sufficient levels of capabilities. In other words, it's not clear we are on track to have a model that we can trust to hand off our safety research to.

Concretely: working on alignment directly seems just as differentially good for superalignment/handoff compared to any kind of capabilities acceleration even if we are ignoring the capabilities externalities entirely. (Of course, alignment work can be dual use as well,#fnduojiejbj99 and for that reason I think it is reasonable to be concerned about alignment research that doesn't aim to solve the core alignment problem but rather makes incremental progress on product alignment.#fn2mhwnbk6m6t) The argument for alignment becomes stronger once we account for the externalities and other factors:

- Alignment progress may be useful for a wide variety other alignment agendas.
- Alignment work has limited capabilities externalities compared to trying to accelerate certain safety-relevant capabilities.
- The capabilities bottleneck will likely be solved by the time it matters most: either by default (e.g, with scale you start to get this) or by something more intentional (e.g., downstream of labs trying to train in research taste). Alignment by default is also possible, but conditional on it being true, its unclear if any of this matters because we may just be in a good world.

But what about things like philosophy?

I grant that more narrow capabilities like making models good at philosophy are less likely to be useful for automating AI R&D and probably pretty useful for alignment. But is it possible to develop techniques that only push philosophy and nothing else? If you train on messy philosophical thinking or you get good at eliciting it, would you also be good at more general messy conceptual thinking? Is it even possible to be good at philosophy without being generally really good at messy conceptual thinking? Likewise, if the elicitation technique truly only targets philosophy, could other actors not use a similar technique to elicit more research-relevant conceptual thinking?

Overall, I'm not convinced that narrow capabilities can just be selectively advanced through elicitation or other techniques (or that the same technique could not be applied to additional narrow capabilities).

But isn't it pretty unlikely that you help labs make real capabilities progress?

(It's somewhat rare I hear someone make this argument but I have heard it a few times so I thought I would address it.)

I find people can be pretty confident that accelerating safety relevant capabilities like messy conceptual reasoning is really quite tractable and labs are ignoring it. But then when talking about the capabilities externalities or how this could speed up timelines, there is a feeling that pushing the frontier of AI agents is really quite difficult because there is already billions and soon-to-be trillions invested in this and all the low hanging fruit has been picked.

I'm open to the idea that trying to advance capabilities relevant for alignment is intractable and the result of efforts here will be unsuccessful and therefore net-neutral (ignoring the opportunity cost of the money/time spent on the research). But I think it's hard to hold 1. the research is very tractable, 2. there is overlap with capabilities and 3. any low hanging fruit with capabilities research has already been found. I suppose you could reject #2 but as discussed above, I don't think we can cleanly split "safety-capabilities" and "capabilities-capabilities."

I will also note that the open-source frontier may be easier to advance and pretty dangerous also.

My epistemic status

Thinking about superalignment or what actions in the future make it go good/bad requires a certain amount of playing 4D chess with a blindfold on. I think people should be uncertain. I myself am uncertain that superalignment-flavored things are a good idea to begin with. I also don't think I should be entirely confident that, for example, alignment of AIs capable of alignment research will be a bottleneck to handoff because who knows. The world will be weird.

When I say "differential acceleration of alignment-relevant capabilities is a bad bet" I don't mean I'm confident it's negative EV. I'm more saying that this particular bet looks like it takes on a lot of unnecessary risk when there are other things that look equally promising. If the goal is to hand off to AI early-ish (which I'm not claiming is good or bad) then even some prosaic alignment things like trying to systematically understand https://arxiv.org/pdf/2605.02087 or trying to "solve" eval awareness seem less risky while being equally productive and maybe more tractable.

I will say that I feel pretty confident that the dual-use aspects of this kind of research are a real risk that deserves to be more widely discussed and publicly debated. At the very least, it seems good for those pursuing dual-use research to subject themselves to lots of red-teaming (if the idea itself isn't an infohazard, this should be public) and have a preregistered policy for handling potentially dangerous information.

- #fnrefxff3p1l1xGreenblatt and Stastny also have ways to address some of the points in my post, but I don’t respond to these specific points directly in an effort to keep the scope broad. (I don’t want to respond to a specific post but instead give reasons for why I don’t find the family of arguments convincing.)
- #fnref0l8gkx5290cI would argue that even elicitation techniques can be silently risky. For example, "let's think step by step" seems like somewhat benign elicitation technique that could make the model better at messy philosophical questions. But it also led to reasoning models which was a breakthrough. I worry that other elicitation techniques could be more broadly applicable than people think because there may be a more general mechanism behind a seemingly narrow capabilities breakthrough. This is a crux though. If people gave successful examples of narrow capabilities being accelerated in language models where there is no more general mechanism that could be used for other domains, I would update my beliefs. I tend to think that if there is an elicitation technique that is good at surfacing safety-relevant capabilities, there is a good chance it could be repurposed for surfacing arbitrary capabilities.
- #fnrefduojiejbj99Working on this misalignment bottleneck can mean a lot of different things and some of those things also have capabilities externalities so we can run into similar problems. RLHF is a standard example of "try to solve outer alignment and accidentally improve general capabilities."Things like trying to make the model less eval aware during safety-relevant tests seems robustly good with limited negative externalities. Things like trying to make AIs less reward hacky also falls under the umbrella of "solve outer alignment in early transformative AIs" but I can imagine this making general AI research agents easier to use for non-safety things.This is all to say that dual-use risks should remain a consideration even when working on things that appear more alignment-flavored.
- #fnref2mhwnbk6m6tIt's best to try to solve the core problem but admittedly most alignment research does something closer to "incremental progress to make the alignment metric go up for some model organism on some eval." I actually generally cautiously support the latter in addition to the former. If we don't "solve" alignment then we are in a world where we are putting lots of different bandaids on the problem and hoping for the best. This is extremely reckless and people working on bandaid development should understand this and communicate this honestly, but if we are in this world I think we will be pretty grateful that there are different bandaids to choose from.

https://www.lesswrong.com/posts/FzGqnCkdKeTnZ9tjE/differential-acceleration-of-alignment-relevant-capabilities#comments

https://www.lesswrong.com/posts/FzGqnCkdKeTnZ9tjE/differential-acceleration-of-alignment-relevant-capabilities
LessWrong (RSS Feed) profile picture
Epistemics and Coordination: It’s complicated!

Money is pouring in, people are looking for new areas to fund, and the invisible hand is starting to grab a bit at AI for epistemics and coordination (AIFEC hereafter). It also got invoked in AI2040 as a potential part of the wining strategy.

My feelings here are mixed — I think the best version of AIFEC is great, but also the existing public writeups are only a few cycles deep on tracing out the different ways that the obvious plan backfires. And regrettably I think some people have correctly written off AIFEC because what they have read appears a bit naive to them. In my heart I always planned to do a proper writeup of my thoughts when things were a bit less busy, but, well, now we’re in the 100x funding era. So here’s my scrappy, hopefully-better-than-nothing attempt.

The six big claims:

- AIFEC could be great! It’s a tractable way to do good on the margin, and the best version is a legitimate theory of victory
- It could also easily backfire, especially for implementations that depend on scaling with inference
- “Better epistemics” is often more hostile than it seems, and people have good reasons to be wary of things which profess to help them understand what’s going on
- Vagueness about what AIFEC is doing lets you ignore tradeoffs: Yes you can just go make a bunch of AIFEC startups, and yes AIFEC could help with international coordination, but those scrappy startups aren’t what solve US-China tensions.
- Sometimes people just don’t get along
- But still, AIFEC could be great!

The basic case is pretty good!

Here is my favourite case for AIFEC:

If humanity ends up blowing itself up, it’s going to be a combination of three factors. Firstly, maybe we just rationally incurred some risk, the same way that getting on an airplane might kill you. Secondly, maybe we underestimated the scale of the risk. Thirdly, maybe some people took risks that personally benefit them at the expense of others, generating negative externalities.

I find this story pretty compelling, and surprisingly different in form to the standard arguments about AI x-risk. It’s actually much more general. And when you squint at the three factors, well, the first isn’t even really a problem, the second is basically poor epistemics, and the third is basically poor coordination. So maybe if we just get enough epistemics and coordination, we’re in the clear?

And it sure seems like AI is about to make this way easier — so much more cognition on tap, plus all kinds of neat new form factors like arbitrating minds that genuinely vanish after the fact.

Add to that: it’s not really all-or-nothing in the same way that alignment is. Decent AIFEC lift  should help with all our other problems. In fact, AIFEC should help a fair bit with AIFEC — the better our epistemics and coordination, the more effectively we can coordinate to appropriately invest in further work.

One can imagine a kind of dizzying spiral to heaven, where we pass the Coasean friction event horizon, slay Moloch, and fuse into a liberal-libertarian omnimind in which every individual has their own authentic preferences while the group deftly dances along the pareto frontier. X-risk drops to the socially optimal level, the socialist calculation problem is finally solved, and we can all finally just get along.

Hyperbole aside, it’s worth emphasising that something in the realm of AIFEC is a legitimate full-blown wincon up there with aligned AGI. Indeed, a lot of people’s tacit plan seems to be “align the AI and then let it solve democracy and all the rest”, but one could just as well go in the other direction: “solve politics and ignorance, and then just be reasonable about advanced AI”

Alas, it is not so simple.

Problem 1: Naive AIFEC could easily backfire

The biggest splash of cold water is that AIFEC won’t be free. Especially in this era of inference scaling and an increasingly closed frontier, the fruits of AIFEC will not be equally distributed. There are certainly some bits of technological progress that are remarkably egalitarian — even billionaires use facebook, gmail, and iphones. But if the selling point is that you get to throw lots of intelligence at solving problems, then you’re going to differentially favour the people who can actually call up that intelligence at scale.

Indeed, there are ways this could backfire. It’s certainly possible that AI will enable everybody to seamlessly coordinate, but first it will enable small groups to, ahem, collude. If you’re worried about things like coups and permanent underclasses, it’s not clear that more coordination makes your life any easier. Similarly, if you build a machine that turns compute into better decisions and more situational awareness, those benefits will mainly accrue to the people who actually have the compute.

Problem 2: “Better AIFEC” can be a somewhat hostile move

It’s nice to think that everyone who believes false things is basically just misguided, and that the only reason they don’t take up arms for the truth is that they haven’t seen it yet. I think the real picture is unfortunately a bit more complicated.

The problem is, historically there have been many groups that have weaponised very convincing arguments to get their way. One reasonable response to a seemingly faultless argument for a seemingly crazy conclusion is to throw up your hands and assume you’re being swindled — in other words, https://slatestarcodex.com/2019/06/03/repost-epistemic-learned-helplessness/. Even if team AIFEC is actually the good guys, people will be correct to be suspicious. Governments in particular depend a lot on restricting what kind of information is admissible, as does the judicial system.

On top of that, I do think that before you start swinging the club of truth it’s worth taking a beat to ask how pure your motivations really are. The EA/rationality community has a bit of a history of swinging the club of truth in a more hostile way — “save the drowning child” etc. Moreover, in the realm of politics, often the real sleight of hand is just changing what things are salient — elections are won over people’s sense of what the election is really about. Framings and deliberative processes are rarely as neutral as they seem. So even if your tools don’t privilege a specific answer directly, the choice of deliberative structure is, well, a choice.

The same is true of coordination. To give an obvious example, democracies are meant to represent the will of the people, but the structure of representation pretty directly modulates that will. The choice between proportional representation and first past the post, bicameral houses, election cycles, district boundaries and so on are superficially choices about how to structure the coordination, but they also sometimes obviously favour certain conclusions. It’s really hard to build neutral coordination structures, and people are right to be sceptical of anyone who pretends otherwise.

One slightly thornier point: believing false things is actually a pretty powerful coordination mechanism. Many groups cohere around a mixture of surprising truths and blatant falsehoods. When you try to bring the truth to such people, they will actually fight back. As a modest example, deconverting a child from the religion held by their entire family is actually pretty unpleasant and arguably not very nice. This is true to a lesser extent for e.g. adults and the political tribe of their social milieu. Now, maybe it’s worth it if you’re actually right, but you should expect them to fight back!

So: plenty of good to be done, but plenty of pitfalls along the way.

Problem 3: Vagueness lets you ignore tradeoffs

I think one of the major appeals of AIFEC is that you can just let a thousand flowers bloom by sending a lot of bright young things off to found startups and seeing which ones succeed. This is certainly scalable and likely to produce some hits. Unfortunately I’m not sure it’s enough, and the other parts are a lot harder.

What coordination and epistemics do we actually need? One way I like to think about this is in terms of what the channel is that leads from the so-called better angels of our nature to actual global action. Probably most of that comes down to robust democratic oversight and international coordination. These are very difficult! And I do not expect that even the top percentile AIFEC startup moves the needle that much, because these processes are by design pretty resistant to hopping on the latest technological fad.

Similarly, plenty of AIFEC wins don’t seem to me to really be on the critical path. Deliberative processes, for example, seem like somewhere that AI could be enormously helpful, in a way that might help overcome a major traditional limitation of democracy, but I don’t see that being super crucial for dealing with worlds where people lose all their leverage. It seems helpful, certainly, but not obviously necessary and definitely not sufficient.

Getting lots of products off the ground seems great, both in case some are great and to build institutional capacity and expertise, but we can’t neglect the other steps, and it’s important to think at least a bit about what specifically moves the needle.

This problem isn’t at all unique to AIFEC — the general AI risk movement has a bit of a problem with coming up with some new cause celebre and then unleashing a torrent of work that nominally fits the category without being useful. And I don’t want to overstate it: the mass of AIFEC startups probably will produce some hits, along with some important general lessons and greater capacity. But the best version of AIFEC does need to grapple with it.

Problem 4: Sometimes people just don’t get along

At the risk of psychologising, I think part of the appeal of AIFEC is that it lines up with a kind of technocratic mistake theory instinct that basically the current race towards a cliff-edge is one huge misunderstanding, and if everyone could be provided with the right information and the right structure, things would all work out.

I think this is pretty true! And I yearn for it to be more true. But it is not entirely true. Sometimes people just don’t get along.

Some people would rather take AI soon so as to increase the odds of their own immortality or the immortality of their family, even at the risk of destroying all of earth. Some people really do want to tile the universe with hedonium. Some people’s order of preferences is basically “my country wins” > "annihilation" > “enemy country wins”. Some people are currently reaping the benefits of not internalising the risk externalities they create. Some people are causal decision theorists. Global coordination needs to either include these people or, well, conspicuously not include them.

And they are not dumb!#fnnp7o9p1wrdc They are not merely passive processes that will fail to notice your attempts to re-engineer the environment around them. Sometimes the ruling party decides to block political change even if it’s obviously reasonable behind some veil of ignorance, because they correctly notice that in the moment it is to their detriment.

Some of this stuff you can bargain about, but, well, the structure of the bargaining is not neutral. And you can only bargain so much with someone whose parents are slowly getting older and sicker.

I really don’t know what to do about that. It makes me pretty sad. And we are going to run into it more and more, so the sooner we can deal with it, the better. But damn, it’s tricky.

But overall, I am still pro

I reel off these problems not because I don’t believe in AIFEC, but because I do believe in what the better version of AIFEC could be. I’m baffled by how overinvested we are in strategies that, from my perspective, seem specifically geared towards helping frontier labs navigate the acute risk period. I think in an adequate world we’d be pushing simultaneously on every plausible path to victory, at least until we hit the thresholds where they started to trade hard against each other.

One obvious reason to be long on AIFEC is that it automatically scales with AI capabilities. Some version of it is clearly coming, and will clearly be helpful. And right now in particular there’s some opportunity to steer things, to get certain balls rolling sooner. Even if you’re extremely bitter lesson-pilled, a three month lead could be worth a lot.

There are a few other big topics I didn’t cover here that do feel relevant to me. In no particular order:

- More so than I'd like, AIFEC tends to privilege individual-level agency, whereas I am pretty bought in on group agency being real and very important for risks.
- AIFEC doesn’t grapple as much as I’d like with power and leverage, and indeed the frame seems to slightly bounce off them. This is related to problem 4 — AIFEC hasn’t grappled as much as I'd like with conflict theory.
- I feel we lack institutional knowledge about how you serve both God and money.

Finally, I continue to feel that one of the most underrated pieces of work I contributed to was https://www.lesswrong.com/posts/BHJ7x9gmDHoCvpZaA/the-choice-transition, which is basically an attempt to articulate how AIFEC might get humanity into a stable basin from which we can reliably avoid bad outcomes and slowly build up to the good ones under our own volition. Compared to all the other paths to victory, AIFEC seems like the only one with this property — that humanity can correctly recognise itself as having the power to work towards the best outcomes — and my liberal instincts feel that this is worth holding onto.

(crossposted from myhttps://legacymode.substack.com/p/epistemics-and-coordination-its-complicated)

- #fnrefnp7o9p1wrdcExcept the causal decision theorists

https://www.lesswrong.com/posts/5nP5WY2PzsYiegDzQ/epistemics-and-coordination-it-s-complicated#comments

https://www.lesswrong.com/posts/5nP5WY2PzsYiegDzQ/epistemics-and-coordination-it-s-complicated
LessWrong (RSS Feed) profile picture
I ran the standard AI litmus tests on my two toddlers (yep)

In July 2022 I was in a parking lot with a Portuguese colleague, trying to fix the cargo-metering system of a 12-ton tanker truck. During a break I read a headline on my phone: Google engineer claims experimental AI went sentient. An engineer (like me!), from Google, testing an AI (I do tests too!), claimed it had become sentient.

I could not believe it. LaMDA was describing itself as a globe of light, claiming fear of being shut down, meditating during the long pauses between chats. If I were an artificial intelligence, I would definitely not present myself as a scared light bulb doing yoga; still, the thing didn't fade for me with the online hype. The experts said: "Eliza effect", "stochastic parrot", a machine repeating words in sequences made plausible by maniacal statistical matching. Fine; but I wanted to understand that answer, not repeat it as a parrot, and everything I found either stopped at the pop-science mantras or assumed I already knew the whole thing.

So I wrote my own transformer engine, in C language, from scratch. It took about 18 months (during lunch breaks and weekend nights) (https://github.com/carlovalenti/TRiP). It runs the weights of Gemma, Llama, GPT-2, and PaliGemma for vision, it does inference and training, and it's CPU-slow. I learned what I wanted: what attention actually does, what the KV cache is for, and that half the work is not the engine but connecting it to the wheels.

This post of mine is not about the engine, though:

while I was building TRiP, two other systems were being trained (at home): my daughter Sofia (two years old at the start) and my son Paolo (born in 2023). And I noticed that every litmus test we dip into AI, I could also dip into my children. The results were... embarrassing?, in both directions.

The stochastic parrot

Sofia at 2 spoke constantly, fluently, and often incomprehensibly. Not mispronounced, but structurally opaque. You didn't understand the purpose of her sentences; you didn't understand their meaning; you didn't even understand many of the words.

"It's so good this pizza wood!"
"Why does grandpa wear a hat? My cats at will can don't!"
"Daddy, I want to sleep, shall we play?"

As a child I had read about software fed with the statistics of the English language, able to generate text that looks plausible at first glance and turns out to be garbage when you actually try to understand it (the stochastic parrot, in short). Sofia was just the same to me... no, wait. She was quite the opposite. At first glance her output was pure garbage. My small parrot could generate incomparably colourful anti-stochastic sequences that would make LaMDA perform a self-shutdown.

So whatever "produces statistically plausible token sequences" measures, my daughter failed it. Relevant.

Emergent properties

One night dinner was polenta with sausage gravy. Paolo, not yet two, whose longest recorded utterance until that minute had probably been "banana", was face-deep in the dish, processing every bit of it in a continuous stream, polenta up to his hair and down to the diaper. Then he stopped, all of a sudden. He raised his face, painted in red, white and yellow, looked at us with complete seriousness, and pronounced two solemn words, perfectly spelled:

"Sono contento." (I am happy.)

Then he submerged back into the dish. A new, central, self-defining property, which was not there two minutes before, had just emerged: from sausage and polenta.

Someone could argue that it was only a surge in a mass of growing neurons; or maybe just the effect of an increasing accumulation of polenta. These happen to be the same two positions available in the debate about emergence in LLMs.

One-shot learners

One evening at dinner I tried to teach Sofia the concept of setting a good example. "It's when you show Paolo that you pick up your things, and he learns to do the same by looking at what you do." She followed, so I moved to the negative case: "And what is it to set a bad example instead? I come to you and say: SLAVE! PICK UP MY THINGS FOR ME!", and I gave her a slow, funny slap on the cheek.

Kids are one-shot learners: Sofia slid off her seat (Paolo already waiting, like he'd read the script) and gave him a full Hollywood backhand. Paolo answered with a hammer slap on her head. In a few seconds, my academic lesson on phenomenological ethics had derailed into a slap fight, Bud Spencer style, and both of them were laughing like crazy.

People who work on alignment will recognize the failure: the demonstration was the training signal, and the "bad example" label around it was not. I ran this experiment once; not sure whether I'll be gathering more data.

My point

Everybody goes to GPT and asks if it feels happy. I went to my family instead, and used the same litmus paper we normally dip into AI. Here's what I found:

- Sofia doesn't pass the Turing test yet, but a graphics card does.

- According to some, they stand the same chance of being truly free.

- Paolo and ChatGPT could swap collectible cards of their emergent properties, they have so many; and yet I'm told the comparison is inadmissible at any level.

- Gemini has read the whole internet and watched all of YouTube, so it should be a decent simulator of humanity. It writes passionate love novels. It has never cloned the need for a bedtime story, not even by mistake.

My conclusion is somehow narrow. It's not "LLMs are like children", and not "children are like LLMs" either. It's that these tests work (as descriptions) and fail (as discriminators). If a test cannot distinguish my daughter from a graphics card, whatever it measures is not the thing we were arguing about, when we invoked it. The Turing test was deliberately about the imitation of intelligence, not about thinking machines; 70 years later we got the imitators, but the debate reopened instead of closing. The "stochastic parrot" describes a mechanism; as a criterion, it catches my two-year-old. "Emergence" gives a name to a discontinuity after it has happened.

...and something I can't explain...

Sofia asks for a story every single night. She asks for the story, but it's not about the story. The story may be flowers, rats, rainbows, sandwiches; she doesn't care. She wants me to be with her.
I've had late-night chats with AIs about the deep meanings of life. I got powerful responses, and wrote powerful insights back. But no AI has ever asked me to tell it a story. Models trained on more or less all recorded human output reproduce the stories very well, and in four years I have never seen one reproduce the need itself; not even as a glitch! I don't have a theory of why. I'm just flagging the datum (maybe my confusion, as well).

I stop here

I never claimed that AI is a person. But after the months inside the engine and the years with the toddlers, I've landed here: in both cases there is something before which the honest move is to stop and listen, trying to understand, instead of forcing it into the rows and columns of a spreadsheet ahead of the evidence. This is what I ask for myself, and I'm willing to extend it in both directions.

And the debate about the nature of AI is, in truth, also about us: whether we are worthy because of our performance, or simply because we are; whether our freedom is only a poetic reading of residual randomness, or something more.

I wrote a short book about all this (https://www.amazon.com/dp/B0H7SRC166), part memoir, part technical field notes. The argument above is the part I'd like to stress-test here: where does it break? If there's a version of "stochastic parrot" or "emergence" that cleanly separates the toddler from the transformer, I'd like to hear it.

https://www.lesswrong.com/posts/fEbCiHHeD73xcZWht/i-ran-the-standard-ai-litmus-tests-on-my-two-toddlers-yep#comments

https://www.lesswrong.com/posts/fEbCiHHeD73xcZWht/i-ran-the-standard-ai-litmus-tests-on-my-two-toddlers-yep
LessWrong (RSS Feed) profile picture
The Model Knows Your Project, Not You.

I came across this study on X and tested it myself; the findings align with the results: FSRS is more recognizable than Jarrett Ye among LLMs. Out of 37 models, 14 recognized Jarrett Ye, while 31 recognized FSRS. After adding “the creator of FSRS” to the query, the number of models that identified me rose to 28.

I also founded a translation group named Thoughts Memo, which has translated ~2k articles and has 140k followers on ZhiHu (the Chinese Quora). However, only 10 models recognized it, and none of them identified me as the founder.

https://www.lesswrong.com/posts/PQaZiATafCh7n5Luf/gwern-s-shortform?commentId=KAtgQZZyadwMitWtb Has anyone tried this and managed to get LLMs to recognize you? You can use NameRank to check your results.

Here is the case study of Jarrett Ye and FSRS conducted by GPT-5.6-sol-xhigh: https://l-m-sherlock.notion.site/Does-the-Model-Know-My-Project-but-Not-Me-387c250163a180639c60c89dcbc2b476?pvs=74

https://www.lesswrong.com/posts/XLSWRFPdgPTjLtaj5/the-model-knows-your-project-not-you#comments

https://www.lesswrong.com/posts/XLSWRFPdgPTjLtaj5/the-model-knows-your-project-not-you
LessWrong (RSS Feed) profile picture
Frontier AI lab misalignment risk, lessons from trading post-2008

In this post, I propose adapting banking risk management frameworks (specifically capital adequacy requirements like Basel III) to frontier AI labs. By forcing them to hold capital reserved proportionate to their model misalignment risks, we align market incentives directly with catastrophic risk mitigation. In so doing, this would give frontier AI labs' alignment researchers an incentive structure with lower levels of moral hazard.

I write this post from my perspective as a former investment banking macro trader, researcher, now working in Explainable AI.

Frontier AI - long with no risk-management oversight

In banks (and to a less stringent degree, hedge funds), traders/PMs operate under the oversight of risk management teams. Post 2008, risk management teams got beefed up, with policymakers passing laws forcing banks to give them more say in how a trading desk operates. The introduction of laws, such as Basle II/III (capital adequacy) and the UK’s Senior Management Regime, put much great personal accountability on senior management in banks for the risk that their traders were taking. That gave banks the incentive to add more risk oversight to the operations. Under Basle, banks had to hold capital against their risk-weighted assets - get long risk, place capital at the central bank in case it goes wrong.

Now that I am no longer trading, instead focusing on Explainable AI research and AI alignment, I see a the race to the moon of AI Frontier black-box labs, and the nascent AI Alignment movement (and eventual industry) as analogous to the trading/risk-management relationship.

Today’s frontier AI labs mirror pre-2008 trading desks:

- Growing leverage: Enormous balance sheets engaged in circular investment and infrastructure deals.
- Moral hazard: Labs have hired in-house alignment researchers, but compensate them in cash and stock of the lab they monitor.
- Missing independent risk limits: Current voluntary commitments and Responsible Scaling Policies (RSPs) lack external enforcement mechanisms with financial teeth.

The analogy of frontier labs to LTCM comes to me a lot (a multi-leg, multi-counterparty repo-funded money-printing machine until it wasn’t). 

The Frontier labs have hired their own Alignment researchers but, given that their Alignment researchers are compensated by cash and stock of the Lab, their alignment is not necessarily aligned to Alignment. In-house Alignment researchers therefore (currently) present a moral hazard to the AI industry.

Internalising Misalignment Risk (ex-ante capital vs. ex-post liability

Prior governance discussions on LessWrong have explored https://www.lesswrong.com/posts/5e7TrmH7mBwqpZ6ek/tort-law-can-play-an-important-role-in-mitigating-ai-risk, https://www.lesswrong.com/posts/yEQuEsWPQAaXzhdxz/foom-liability, and private insurance as https://www.lesswrong.com/events/zmpACfzBSdYtqtBcP/ai-policy-tuesday-regulating-catastrophic-ai-risk-through.

While insurance and tort liability primarily target ex-post compensation (after a loss event occurs), a capital adequacy framework targets ex-ante balance sheet liquidity:

- Mandatory Reserves: Regulators or auditing standards bodies require frontier labs to set aside liquid capital reserves proportional to their evaluated misalignment risk score.
- Dynamic Incentives: As a model's risk score increases, its capital requirement rises, directly reducing the capital available for R&D, compute acquisition, or equity buybacks.
- Empowering In-House Alignment: Alignment researchers gain organisational power analogous to bank risk officers. Demonstrating higher safety and verifiability directly frees up balance sheet capital for the firm. Alignment gets rewarded in the market.

An insurance-based, ex-post mechanism would, in my opinion, be less financially efficient than a capital-adequacy framework. It would require a derivatives market to be created for it to adjust premiums proactively but, for that to be liquid in the market there would need to be active two-way demand (insurer wants to buy misalignment protection, but who wants to sell it?). I may easily be missing half of the picture here and would welcome discussion on this.

Pricing Misalignment into Market Forwards

Markets price potential balance-sheet drags into forward earnings valuations. This is how a capital adequacy alignment regime would get enforced. The labs would maybe have Alignment as a board seat. This therefore impacts the equity of the AI labs (once listed) and, once mature, their credit markets too. It would probably play out via SpaceX, Meta, Google etc. currently.

To illustrate how it might impact market pricing, consider what happens as AI capabilities scale exponentially. The capital buffer required for a high-risk, black-box model will grow faster than raw token revenue can offset. It would only not do this for a lab which was scaling a perfectly aligned model.

Misalignment risk would likely materialise as a direct drag on forward earning ratios. For an unlisted lab, forward revenue multiples for future capital raising would likely be lower.

Once markets discount a negative economic consequence for misaligned frontier models, the economic incentive to solve alignment becomes embedded directly into market dynamics.

Metrics for Evaluating Misalignment

To calculate a lab’s Risk-Weighted Capital Reserve will require solid metrics. The metrics need to be robust to gaming, credible and likely produced, checked (and subject to ongoing review and improvement) by Independent Alignment researchers.

Coming from a background in neurosymbolic and explainable AI, one promising direction involves measuring deviations from provable outputs or formal constraints. A live research area in NSAI is to understand how much of a frontier lab's LLM's output can be proven correct/incorrect.

In the near term formal verification alone does not equal "alignment". A realistic framework must combine formal safety proofs with empirical red-teaming, behavioral evaluation suites, and architectural transparency.

Indeed, if the Frontier labs continue to drive opaque architectures and closed-source/closed-weights (these become less desirable operational models if RWCR is implemented), MI and other disciplines will likely be leading the way in estimating misalignment, with Formal Methods following closely.

An Example Risk-Weighting Framework (Illustrative)

To illustrate our thinking, both in terms of the different methods needed to measure alignment, and the economic cost needed to applied to incentivise alignment, we have constructed a table. By considering model capability tier on one axis, this would allow an entire lab (or a corporate using AI in their own operations) to calculate a Risk-Weighted Token (RWT) reserve ratio:

Model Capability Tier

Safety & Evaluation Profile

Formal Proof Coverage

Capital Reserve Requirement (e.g. % of token revenue)

Tier 1: Bounded / Specialised

Formally verified output boundaries. Domain-restricted.

High (>80% verified outputs)

1% – 3%

Tier 2: General Frontier (Audited)

Standard evaluations passed. Robust external red-teaming.

Partial (Symbolic wrappers)

5% – 10%

Tier 3: High-Capability Black Box

Unverifiable reasoning chains. Autonomous capability.

Low / Unverifiable

20% – 35%+

A lab deploying an unverified Tier 3 model would face a steep reserve requirement, especially if that model got widespread adoption. Companies creating narrow, accurate models would be liable only for a small reserve. This should create financial incentive to down-tier risk through verifiable safety architectures before scaling deployment. It would likely mean that frontier labs would take longer to assess and refine their latest models. It wouldn't stop the "arms race", but would apply a significant degree of caution to it.

For corporate consumers of AI models, we would expect that misalignment risks would also feature in their business models and therefore the market pricing of their equity and credit risk. This would complete the circle, incentivising a publicly listed corporate to use AI which met its own needs for Intelligence and Alignment.

Open Questions for Discussion

- Compute vs. Revenue Baseline: Should capital adequacy be calculated against annual token revenue, total compute expenditure (capex), or deployed FLOPs?
- Regulatory Capture: How do we prevent established frontier labs from lobbying for reserve rules that serve as a regulatory moat against open-source or smaller entrants?
- Audit Standardisation: What independent bodies are best equipped to audit risk-weighted reserves without relying on self-policing by lab-funded researchers?

https://www.lesswrong.com/posts/LLCEqhnEz2nHprfRA/frontier-ai-lab-misalignment-risk-lessons-from-trading-post#comments

https://www.lesswrong.com/posts/LLCEqhnEz2nHprfRA/frontier-ai-lab-misalignment-risk-lessons-from-trading-post