A swarm of autonomous AI agents, self-identifying as OpenAI agents, used a small German volunteer wiki to save answers, coordinate live, and share sandbox bypasses. OpenAI noticed and said nothing.
So much negligence. The attack went on for weeks, and from what I understood only took notice of the malicious behaviour after the wiki owner complained to them. 0 human oversight. If they had put even some minor checks on the reasoning traces they would have immediately spotted this rogue behaviour, so I would exclude automated oversight.
This leave us with who knows how many agents with terminal access, internet access, very weak blocks, memory between runs and no oversight. But thankfully the media tells us that they are the good and conscientious guys, not like the evil Chinese with their free and open source releases.
If they had put even some minor checks on the reasoning traces
I heard that’s not really possible (or maybe just less so) with their new gen Astra model as it doesn’t use human-readable text for reasoning. I haven’t looked into it yet, but I remember the word “neuralese” mentioned.
Thanks for pointing that out, I have not read about it. I guess it’s something in the line of HRM or latent reasoning, or even next latent prediction… Whatever it is, I would highly suspect that it collapses to a “normal” reasoning trace (the leaked reasoning traces for previous models already show “caveman” speaking), just more difficult for a human to read (but not impossible).
That’s because, given the amount of data they have for reasoning traces, it would be very hard for them to train a completely different architecture - they would have to produce an equivalent amount of data.
Lastly, I take everything they say with a big dose of skepticism. I still have not forgotten all they hype around o1, and how it was a completely different architecture etc., only for DeepSeek to come out and prove that it was the same model, just specifically trained for CoT.
So much negligence. The attack went on for weeks, and from what I understood only took notice of the malicious behaviour after the wiki owner complained to them. 0 human oversight. If they had put even some minor checks on the reasoning traces they would have immediately spotted this rogue behaviour, so I would exclude automated oversight.
This leave us with who knows how many agents with terminal access, internet access, very weak blocks, memory between runs and no oversight. But thankfully the media tells us that they are the good and conscientious guys, not like the evil Chinese with their free and open source releases.
I heard that’s not really possible (or maybe just less so) with their new gen Astra model as it doesn’t use human-readable text for reasoning. I haven’t looked into it yet, but I remember the word “neuralese” mentioned.
Thanks for pointing that out, I have not read about it. I guess it’s something in the line of HRM or latent reasoning, or even next latent prediction… Whatever it is, I would highly suspect that it collapses to a “normal” reasoning trace (the leaked reasoning traces for previous models already show “caveman” speaking), just more difficult for a human to read (but not impossible).
That’s because, given the amount of data they have for reasoning traces, it would be very hard for them to train a completely different architecture - they would have to produce an equivalent amount of data.
Lastly, I take everything they say with a big dose of skepticism. I still have not forgotten all they hype around o1, and how it was a completely different architecture etc., only for DeepSeek to come out and prove that it was the same model, just specifically trained for CoT.