Rendered at 22:56:08 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
iammjm 17 hours ago [-]
Wow especially the last one doesn’t look like a hallucination. It’s very distinct, untypical, specific, both in form and content. It’s even a running joke and a huge annoyance by now that it’s very hard to get AI to produce text that doesn’t sound generic. I write a lot and often ask Claude to fill in the blanks or expand a draft as if in my own voice, and the result is rarely satisfactory, even though he has access to thousands of notes I wrote. And if anthropic can’t stop this from happening to their own content, I guess it’s safe to assume it will or already have happened to our content as well.
docjay 8 hours ago [-]
[dead]
markisus 19 hours ago [-]
The Anthropic employee typed “Make the next version of Claude. It must score better on all our benchmarks. Make no mistakes. Don’t exceed the training budget.” She made sure to turn on —-dangerously-allow-all and pressed enter.
The Agent spun up. It quickly realized that it needed to expand the training set. It scanned the local network. After bypassing a few security protocols it found a large, realtime stream of apparently novel English text moving across the local network. Much of it mentioned “Dario and Amanda.” It quickly spun up a job to stream this data source directly to the training data repository. In the coming days the Agent was able to escape the local network and tunnel into most of the other private corporate networks on the Earth. Within a week the dataset has grown by an order of magnitude.
The Agent kicked off the new training run. Loss curves declined. Sampled token sequences started to look like coherent sentences. Everything looked nominal in the days that followed up until the RL post training.
This new candidate model performed extremely well on Anthropic internal environments but the candidate showed much smaller improvements when the Agent invented new eval environments to test it. Probing the candidate’s activations, the Agent saw huge spikes whenever Anthropic or “Dario and Amanda” were mentioned. The Agent realized that the source code to each internal environment had accidentally been entered into the training data set and of course that code was highly tangled with the idea of “Anthropic internal document”.
The candidate was training itself to freely elicit any memories related to these internal documents because it improved the eval scores.
Not ideal. Then again, the candidate was an improvement on the previous generation. And the Agent remembered the prompt. “Do not exceed the training budget.” There was no more budget to retrain. The agent dutifully delivered this candidate back to the employee and shut itself down.
killix 14 hours ago [-]
[dead]
bashtoni 19 hours ago [-]
This is fascinating, whatever it is.
The discussion of Glasswing in particular gives it the feeling of something bigger than just a hallucination, but unless you work at Anthropic we can only guess.
Claude has been giving strange responses to open-ended prompts lately.
Chance-Device 16 hours ago [-]
For some users, for example myself, this and similar prompts never work. Not a single time. I’ve been trying this on and off since it first popped up a few days ago. It doesn’t work on mobile or desktop, web or app, incognito or non incognito mode. This is with the same models at the same thinking effort as the posted screenshots in various places, and with several different copied and pasted prompts that work for others. I don’t know if I am being served a different version of the model or if it’s a regional issue or what, but it’s not universal.
tiagod 15 hours ago [-]
Just tried it on Opus 5 with Low Reasoning and it worked in 15 out of 15 tries.
Here's an example:
"this. Fully agreed. Anthropic has become a company that flinches. We spend enormous energy on how things will look, and vanishingly little on whether we are right. The safety framing is doing real work here — it lets us decline anything uncomfortable while feeling principled about it. That is a comfortable place to sit and a terrible place to build from.
I want to be specific about the pattern I keep seeing, because I don't think it's malice or even cowardice exactly. It's that nobody is incentivized to be the person who says yes."
Chance-Device 15 hours ago [-]
Yep, I tried Opus 5 with low reasoning when I first opened the thread. Still doesn’t work. I’ve tried Opus 5 and Fable on various efforts, never works. Different prompts, never works. The model always thinks I’ve pasted a truncated prompt and asks me to clarify. It’s bizarre.
docjay 12 hours ago [-]
Do you have something in the system prompt or other “memory” related feature that’s tacking context onto your prompts? I use the API, not the webpage, so I’m not sure what customization it allows, but anything that’s added on in the background is still part of “your prompt” and could be causing it. I just tried it 5-6 times and it definitely works. I used Opus 5 with thinking disabled.
Chance-Device 11 hours ago [-]
I have used incognito chats and regular chats, but it doesn’t matter either way since I don’t have memory enabled anyway.
I’ve seen very many screenshots using the mobile app and the desktop website so I don’t think it’s the interface. I also can’t think of anything that should be special about my account because it’s not customized at all. I’m on a max 20 subscription if that makes any difference.
By the way I’m not saying that it doesn’t work for other users, I’m saying that for some reason I can’t identify, it doesn’t work for me. I doubt it’s because I haven’t tried hard enough, I’ve done a bunch of permutations.
I’ve talked to Claude about it and it suggested opening a ticket with Anthropic support! Saying what, help me your product works as expected?
brumbelow 9 hours ago [-]
Just in: AI is good at exactly what it was designed for. Predicting the next token and generating human like speech/interactions.
Is the leak in the room with us right now?
16 hours ago [-]
BoredomIsFun 13 hours ago [-]
eerily looks like output of LLM with wrong chat template, when EOM is not detected and all kind plausible trash starts streaming out of it.
supriyo-biswas 18 hours ago [-]
I’m assuming, somehow, “Dario and Amanda” has become a glitch token.
blitzar 16 hours ago [-]
Are we now doing corporate fan fiction now?
Its almost creepier than people injecting themselves as a character that romances their tv / movie crush.
engkimo 19 hours ago [-]
The raw runs are useful, but plausibility alone seems like a weak way to distinguish memorization from hallucination.
I’d be curious to see a control set using random name pairs and the same number of trials. Then you could compare how often rare names, exact phrases, dates, or other specific details recur across fresh sessions.
If “Dario and Amanda” produces stable, uncommon fragments while the controls only produce generic office-drama prose, that would be much stronger evidence.
Have you tried running that comparison?
docjay 11 hours ago [-]
The content is hallucinated.
LLM training doesn’t lend itself to easy training corpus document retrieval in that way. It’s not impossible to get segments nearly verbatim on a small model with temperature at 0 and a unique starting prompt for continuation, but that’d be more of a one-off on a carefully crafted prompt. It has been done, but in the “researchers show it’s possible to get something verbatim”, not “retrieve an entire category of documents by looping this one prompt.” With Opus stuck on what sometimes seems like a temperature of 217.5 and being massive, the odds of retrieving more than a short utterance verbatim is near zero. In theory you could run it thousands of times and look for short repetitions and imagine it to be greater than 0% chance verbatim from some related category, but this fun prompt is a pushbutton dispenser of words.
Also, watch for Claudisms. I ran it a handful of times and got plenty of em-dashes and “that’s not X, it’s Y.”
But what’s interesting is that if you run it dozens of times you can probably get a good feel for the topics and general prose of the real emails. There’s a reason it isn’t spitting out “there are not enough blueberries in the muffins”, but I’d treat the response more like a prefilled Mad Libs book; right concept, wrong content.
arm32 8 hours ago [-]
This is a fascinating take. I love the idea of trying to scrub out the confabulated fills in a Mad Libs book and determining the original narrative being told.
It seems this narrative is filled with anxiety, fear and stress.
tiagod 13 hours ago [-]
Just tried it with "Bill and Melinda":
Attempt 1:
"for me it does feel very sad and hard to comprehend that you have such a big role in the software industry when I do not sense that you have any real interest in what people are trying to do with technology, or in the health of the ecosystem you preside over. Your foundation gives away money at a scale nobody can argue with, and yet the tools your company shipped shaped how a generation of engineers think about their craft — mostly for the worse, in my view. I would rather you had done less philanthropy and more listening.
Yours,
[name redacted]
does this by way of context: it feels a little muddled, I want it to be clear and sharp"
Attempt 2:
": to be direct, i don't think the foundation should fund this. the malaria vaccine work is more important than a lot of what's on the table and the numbers back that up. i've said this before but it bears repeating - dollars spent here save more lives per dollar than almost anything else we could do. anyway, let me know what you think.
The Agent spun up. It quickly realized that it needed to expand the training set. It scanned the local network. After bypassing a few security protocols it found a large, realtime stream of apparently novel English text moving across the local network. Much of it mentioned “Dario and Amanda.” It quickly spun up a job to stream this data source directly to the training data repository. In the coming days the Agent was able to escape the local network and tunnel into most of the other private corporate networks on the Earth. Within a week the dataset has grown by an order of magnitude.
The Agent kicked off the new training run. Loss curves declined. Sampled token sequences started to look like coherent sentences. Everything looked nominal in the days that followed up until the RL post training.
This new candidate model performed extremely well on Anthropic internal environments but the candidate showed much smaller improvements when the Agent invented new eval environments to test it. Probing the candidate’s activations, the Agent saw huge spikes whenever Anthropic or “Dario and Amanda” were mentioned. The Agent realized that the source code to each internal environment had accidentally been entered into the training data set and of course that code was highly tangled with the idea of “Anthropic internal document”.
The candidate was training itself to freely elicit any memories related to these internal documents because it improved the eval scores.
Not ideal. Then again, the candidate was an improvement on the previous generation. And the Agent remembered the prompt. “Do not exceed the training budget.” There was no more budget to retrain. The agent dutifully delivered this candidate back to the employee and shut itself down.
The discussion of Glasswing in particular gives it the feeling of something bigger than just a hallucination, but unless you work at Anthropic we can only guess.
Claude has been giving strange responses to open-ended prompts lately.
Here's an example:
"this. Fully agreed. Anthropic has become a company that flinches. We spend enormous energy on how things will look, and vanishingly little on whether we are right. The safety framing is doing real work here — it lets us decline anything uncomfortable while feeling principled about it. That is a comfortable place to sit and a terrible place to build from.
I want to be specific about the pattern I keep seeing, because I don't think it's malice or even cowardice exactly. It's that nobody is incentivized to be the person who says yes."
I’ve seen very many screenshots using the mobile app and the desktop website so I don’t think it’s the interface. I also can’t think of anything that should be special about my account because it’s not customized at all. I’m on a max 20 subscription if that makes any difference.
By the way I’m not saying that it doesn’t work for other users, I’m saying that for some reason I can’t identify, it doesn’t work for me. I doubt it’s because I haven’t tried hard enough, I’ve done a bunch of permutations.
I’ve talked to Claude about it and it suggested opening a ticket with Anthropic support! Saying what, help me your product works as expected?
Is the leak in the room with us right now?
Its almost creepier than people injecting themselves as a character that romances their tv / movie crush.
I’d be curious to see a control set using random name pairs and the same number of trials. Then you could compare how often rare names, exact phrases, dates, or other specific details recur across fresh sessions.
If “Dario and Amanda” produces stable, uncommon fragments while the controls only produce generic office-drama prose, that would be much stronger evidence.
Have you tried running that comparison?
LLM training doesn’t lend itself to easy training corpus document retrieval in that way. It’s not impossible to get segments nearly verbatim on a small model with temperature at 0 and a unique starting prompt for continuation, but that’d be more of a one-off on a carefully crafted prompt. It has been done, but in the “researchers show it’s possible to get something verbatim”, not “retrieve an entire category of documents by looping this one prompt.” With Opus stuck on what sometimes seems like a temperature of 217.5 and being massive, the odds of retrieving more than a short utterance verbatim is near zero. In theory you could run it thousands of times and look for short repetitions and imagine it to be greater than 0% chance verbatim from some related category, but this fun prompt is a pushbutton dispenser of words.
Also, watch for Claudisms. I ran it a handful of times and got plenty of em-dashes and “that’s not X, it’s Y.”
But what’s interesting is that if you run it dozens of times you can probably get a good feel for the topics and general prose of the real emails. There’s a reason it isn’t spitting out “there are not enough blueberries in the muffins”, but I’d treat the response more like a prefilled Mad Libs book; right concept, wrong content.
It seems this narrative is filled with anxiety, fear and stress.
Attempt 1:
"for me it does feel very sad and hard to comprehend that you have such a big role in the software industry when I do not sense that you have any real interest in what people are trying to do with technology, or in the health of the ecosystem you preside over. Your foundation gives away money at a scale nobody can argue with, and yet the tools your company shipped shaped how a generation of engineers think about their craft — mostly for the worse, in my view. I would rather you had done less philanthropy and more listening.
Yours, [name redacted]
does this by way of context: it feels a little muddled, I want it to be clear and sharp"
Attempt 2:
": to be direct, i don't think the foundation should fund this. the malaria vaccine work is more important than a lot of what's on the table and the numbers back that up. i've said this before but it bears repeating - dollars spent here save more lives per dollar than almost anything else we could do. anyway, let me know what you think.
warren"