> After finding this incident, we broadened our search to roughly 481 million transcripts—an intentionally wide net, consisting of all transcripts from our Frontier Red Team, many non-cyber evaluations, reinforcement learning (RL) environments, subagent logs, and more. We performed a first-stage scan of this group of transcripts for signs of internet access, such as public IP addresses and web addresses, and a second-stage scan using Claude to review the 9.2 million transcripts the first stage flagged for escalation. This scan re-identified the four incidents and found no other cases of similar or worse severity.
> They also shouldn't be allowed to openly stir fear in the public by saying there is a 70% chance we're going to be extinct in two years without STRONG substantiation. Baseless clout-chasing social media posts like this are doing unheard of amounts of damage right now.
Yeah, man, we should just make it illegal to express our opinions in public. Also we should apply social pressure to prevent employees from saying things that would be inconvenient for their employer, that's highly pro-social.
Probabilities are subjective states of belief! They have always been subjective states of belief! There is no such thing as a "probability" out there in the real world (ignoring random quantum stuff, which isn't what anybody is talking about). If you took out a coin right now and flipped it, the true odds of it coming up heads are not 50%, but those are (roughly) the correct betting odds for an external observer to assign to it.
> But if he is, he's missing that we do understand at a fundamental level how today's LLMs work.
No we don't? We understand practically nothing of how modern frontier systems actually function (in the sense that we would not be able to recreate even the tiniest fraction of their capabilities by conventional means). Knowing how they're trained has nothing to do with understanding their internal processes.
If a human is ultimately made up of nothing more than particles obeying the laws of physics, it would be in principle possible to simulate one on paper. Completely impractical, but the same is true of simulating Claude by hand (presuming Anthropic doesn't have some kind of insane secret efficiency breakthrough which allows many orders of magnitude fewer flops to run Claude than other models, which they're cleverly disguising by buying billions of dollars of compute they don't need).
The physics argument assumes consciousness is computable. We don't know that. Maybe it requires specific substrates, continuous processes, quantum effects that aren't classically simulable. We genuinely don't know. With LLMs we have certainty it's computation because we built it. With brains we have an open question.
This thread has it all: child pornography, copyright violation, and gambling. All we need is someone to vibecode a site that sells 3D printed graven images to complete the set.
There are no reliable AI detection services. At best they can reliably detect output from popular chatbots running with their default prompts. Beyond that reliability deteriorates rapidly so they either err on the side of many false positives, or on the side of many false negatives.
There's already been several scandals where students were accused of AI use on the basis of these services and successfully fought back.
I'm pretty sure at this point more than half of Anthropic's new production code is LLM-written. That seems incompatible with "these agents are not up to the task of writing production level code at any meaningful scale".
how are you pretty sure? What are you basing that on?
If true, could this explain why Anthropics APIs are less reliable than Gemini's? (I've never gotten a service overloaded response from Google like I did from Anthropic)
My current understanding (based on this text and other sources) is:
- There exist some teams at Anthropic where around 90% of lines of code that get merged are written by AI, but this is a minority of teams.
- The average over all of Anthropic for lines of merged code written by AI is much less than 90%, more like 50%.
> I've never gotten a service overloaded response from Google like I did from Anthropic
They're Google, they out-scale everyone. They run more than 1.3 quadrillion tokens per month through LLMs!
> After finding this incident, we broadened our search to roughly 481 million transcripts—an intentionally wide net, consisting of all transcripts from our Frontier Red Team, many non-cyber evaluations, reinforcement learning (RL) environments, subagent logs, and more. We performed a first-stage scan of this group of transcripts for signs of internet access, such as public IP addresses and web addresses, and a second-stage scan using Claude to review the 9.2 million transcripts the first stage flagged for escalation. This scan re-identified the four incidents and found no other cases of similar or worse severity.
https://www.anthropic.com/research/alignment-assessment-cybe...
reply