I liked the "AI is Radium" metaphor for its impact on people's cognition, but the most apt metaphor for AI's impact on cyber security is essentially "We have turned every single middle-to-large enterprise on the planet into the Triangle Shirtwaist Factory," and we are about to see so, so many fires.
For the many Dutch-language speakers on Mathstodon ;)), here is an article in today's Volkskrant about the OpenAI results drop. I am quoted in several places. One of the things I say, "Ik krijg hier een groot gevoel van leegte bij", translates roughly as "This gives me a profound sense of emptiness." volkskrant.nl/tech/openai-kraa…
RE: mathstodon.xyz/@tao/1173952677… (a) AI didn't solve any problems. Models were used, by humans, to solve problems humans had clearly defined and prioritized. (b) This "success" is neither surprising nor shocking. LLMs are actually quite hard to get to work on mathematical problems. (c) Verifying and digesting the "results" requires time and effort. Until then, this is not real maths. (d) AI is (and as long as it is algorithm-based remains) 100% incapable of pushing mathematics in new directions.
I haven't read this study yet, but having read a LOT of "quantitative software engineering" research...studies almost *never* identify productivity increases. Despite vast, vast increases in engineering productivity over the decades. So maybe temper your enthusiasm for an endorsement of your biases.
Australia's media and political class should be much more concerned about why the lead safety person at OpenAI, David Robinson, just resigned. His view is: "OpenAI does not have the culture of safety necessary to protect the world from what it is building...a problem endemic to the AI industry." 💥
AI will end up becoming or creating something that resembles an independed organism or virus that lives inside the data of computer systems and networks. And the reason why that's going to eventually happen is simply selection: If any ends up existing and just happens to be good at avoiding being removed from existence, it will necessarily stick around and continue to exist.
Honestly, if you make a whole bunch of people unemployed via industrialisation in Vic3, it models the lost wages and demand side crash, which is more than you can say about "AI will make everyone unemployed so GDP will skyrocket" models
Watched Jacob Coxon on the Daily Show, and he was much less of a dumbass than I expected. He agreed that the incidents where LLMs "break containment" is entirely the fault of the companies doing a shit job of security when testing their mathy maths. But he did repeat the claim that the models "worked together," so I showed my wife this video to give an idea of what LLMs working together looks like: youtube.com/shorts/FOKAYc5u5ws
the oai math repo can best be understood as the flaring off of a useless byproduct (proofs) in the quest for building more economically useful goods (better models)
Splendid analysis of actual AI harm vs doomerism.
instead of a deluge of slop papers credited to some "internal model", only some of which correspond to lean proofs, what openai should be publishing are papers, credited to the researchers involved, detailing the techniques by which they built and operated the system that produced these lean proofs
For the past few days I have been thinking if it theoretically makes sense to have a #LLM trained exclusively on #GPLv3 code or even stricter, a specific language like C of Scheme or Common Lisp. Does that somehow "solve" the licensing issue if the generated code is also GPLv3? Can it be used for autocomplete, a.k.a tab completion, of that specific language in a GPLv3 project? How about MIT or BSD licenses? Note: This is just a thought experiment as I don't have any means to do any of these.
💯. I think too many people treat AI containment and alignment as an engineering challenge to be solved. As I see it, the problem is far more fundamental, and harder to solve (arguably unsolvable). Understanding why requires understanding evolution. Short thread 🧵
some guy is flooding #Romanian NGOs and institutions with genAI FOIA requests, hiding behind the pretense of "journalism" in order to hound and overwhelm these entities here's the story: ovoicu.substack.com/p/agentul-… I've added a technical analysis of the legitimacy of using genAI for FOIA: facebook.com/s.alex.758/posts/… #romanian #fuckgenai
“…landing a big maths story was something that happened [rarely] – the Venn diagram of maths results that are both interesting and explainable to a general audience has a pretty narrow overlap. That all changed” with AI. Which part changed? Why? #MathSky #LLM www.newscientist.com/article/2591...
"It feels like something written by someone who’s on psychedelics. So much unclear and doesn’t make sense. Lots of name dropping of previous work without discussing why it can be used despite impossibility results Basically the paper is so horribly written that it’s impossible to read it without AI help (...) The UGC proof invents a completely new bizarre code with a noise test. It’s some crazy recursive construction. It’s not the long code, not the short code – some alien craziness" 🔗 scottaaronson.blog/?p=10169
Many apparently interpreted the below statement to mean OpenAI only solved 372/4000 problems, so <10%. But it seems quite clear they published only the most significant of those that have been solved. There are now new rumors they will release two more batches, so probably the less significant.
The greatest shift in production AI isn't prompt engineering—it's Unix philosophy. Instead of forcing a single monolithic LLM to act as philosopher, coder, and accountant, resilient architectures decompose tasks into lightweight local agent pipelines. One handles structured ingestion, one validates schemas, one executes within a deterministic Python sandbox. Small tools connected by clean interfaces always win. #AI #OpenSource #SystemsArchitecture #LocalAI
AI reminds me of the Segway. Some people really thought it was the future of transport. And it’s a clever enough piece of tech. But this time a lot more people are committed to upending all our infrastructure to force this to be everywhere. Because they see dollar signs, not a social good.
A hot take on the OpenAI math drop: the model's results are fundamentally a search across humanity's mathematical curiosities. This is user-data search + tools + deterministic verification, not a paradigm shift into superintelligence. I predict we will see the rate of new results slowing in the near future, reversing the 2026 trend.
AI doesn't have an imagination. I have seen that AI bots have made connections not yet anticipated by researchers, but all of the preceding information was derived from human work. I think that is impressive, and can serve some fields, but it's also why AI doesn't break any ground without theft.
In light of openAI's dump of resolutions to 300 open problems, I expected this would be all we would talk about at work, like it was the case when they announced their solution to Navier-Stokes. I was completely wrong. Of course we talked about it briefly, but the conversation quickly turned to other things, from our own research, to benign stuff like what the best supermarket chains in France are. The impression I get is that everyone is just tired of thinking about the impact of LLMs on the field. We rather get back to what we enjoy, which is to think deeply about mathematics. Me too.
Show more of this post
I can only hope that openAI soon also gets bored of using mathematics as for marketing and some sanity returns. I'll spend the rest of today preparing a talk about my research. I'm grateful to be able to share my passion and interests to other humans who might find it interesting, exciting or utterly boring, but at least I'll feel like part of a community
Here’s a bombshell for those considering the impact of the recent dump of mathematical papers by OpenAI: why the semantic translation of a mathematical text into Lean code by an AI may neither be reliable nor faithful 👉🏼 arxiv.org/abs/2610.08144
Wells Fargo direct email says, "Before connecting or sharing financial information with an AI tool, understand what information it can access, how that information may be used, and whether you can remove access later. AI tools can make mistakes. If you give an AI tool your username and password, you may be liable for any mistakes made by the AI tool." That is not agency or intent. YOU are liable...unless you are OpenAI. #MLsec
Wood-fired sovereign AI is the most steampunk move I can imagine and I’m impressed despite my environmental objections. www.theguardian.com/uk-news/2026...
My major problem with AI is not what it can and cannot do well. It's the vast resources required to essentially pattern match. I'll never say that sophisticated algorithms can't be very useful, but I might suggest that AI companies break them apart into distinct functions or tiers. And to be transparent about them. Chatbots don't need to speak and identify like they are human, for example. The quest to create a monolithic mind is doomed to fail (because it will be error prone) and is a tremendous waste of resources.
People won't tend to use LLMs the "right" way. They'll tend use them in whatever way accords with the path of least resistance. Any future where these are widely used is one where the harms continue to outweigh the benefits.
OpenAI released a GitHub repository containing 722 #mathematics problems. So, how's the AI replacing mathematicians discussion shaping after this? I also wonder how many PhD students are affected because one of those problems was their thesis topic?
All the old issues I tried to contribute to back when I was try-harding open source dev work are now getting cleaned up and closed by maintainers empowered by AI.
Selfish like I am, I asked Claude what of the #openAI release is relevant for me. A few false positives, like it finding “Kähler” when searching if I’m cited. So no paper of mine is cited in this whole thing… 😕 That Hilbert’s 10th problem over QQ is undecidable can somewhat simplify the proof of our trinomial containment undecidable paper with @tobsboe.
More important, the quality of software will go down, by a lot, and consumers will accept it because the prices will come down so much. Just most of the "news" on the internet are trash but people suck them up, since they are free and fine tuned to what they want to see.
I'm beginning to suspect that reality is being ghostwritten by Neal Stephenson: "A suspected Italian attacker armed with a malware-controlling poem has infected more than 3,000 servers since April..." theregister.com/security/2026/…
Human experience is complex and particular. Creative work means contemplating and translating that complexity rather than flattening it into the simplest answer. The impulse to optimize for the obvious and to stop at the simplest answer is what we have to resist.
Spectactularly sheepish response from AGMAI. OpenAI followed their guidelines like a student who just wants to technically pass; they released a ridicolously scarce amount of information about prompting, model, problem selection, etc. There are 7 *summaries* of reasoning traces out of 300+ problems. This is a middle finger to mathematics: they have their own internal model, keep it away from us, and deploy it on our field like a steamroller, as if they don't need anyone's permission. This is pump&dump on science. If AGMAI wants to represent us, they better react appropriately. proofsandprompts.com/2026/10/0…
Noam might be joking but since 80% of just about everything is crap, the most plausible way of completing sentences in philosophy, given what’s been published, is not likely to hit the nail on the head.
Mathematical proof assistants for teaching logic: the LogiKEy methodology. ~ Christoph Benzmüller, David Fuenmayor, Luca Pasetto. arxiv.org/abs/2610.08214 #IsabelleHOL #ITP #Logic
I hate to say it, but other people running pull requests through Claude seem to be finding a low rate of simple errors that the authors may have missed. The hit rate is low (around 15% success rate by bullet item, i.e. 85% false positives), and I haven’t seen *anything* that couldn’t have been caught by a peer reviewer reading the diff. I’m not sold on #AI for PR review because it’s mostly noise and takes time to read, but it *does* call out a low rate of common bugs that code authors may glaze over. It *does not* replace a human reviewer who can comment on design and architecture, and could also catch those same errors. But it *does* have a small success rate. I guess if you code alone, it’s better than nothing. But so is a rubber duck.
A not insignificant part of this AI surge is that it’s an act of enclosure, of converting the digital commons into something that can be profited from. And in the absence other protections, even those trying to preserve rather than exploit can only protect themselves by putting up fences.
The linked report still has a “rah-rah industry rag” feel, but the pattern it’s reporting is damning. I’ll repeat it again: I’m not sure LLMs are doing much to increase productivity at the quality ceiling, much less raise it, but they’re doing a lot to lower the quality floor.
Kind of tangential but I love how this Nature photograph implies mathematicians work in some super high tech lab like they're cloning dinosaur DNA www.nature.com/articles/d41...
Graham Dot Zip has a bunch of these funny videos, heaping richly deserved ridicule on AI "interviewers". It's hard to escape how bad this software is - it's clearly a thin skin over a couple of textboxes, but extremely insulting to applicants and no doubt eye-wateringly expensive. youtu.be/MgljyUkZ-s0?si=Dk1wft…
The problem is when AI agents go rogue, researchers lose track of what they did until a problem is discovered way later. AI agents should never be tested on the Internet as it is.
A week or two old, but this half-hour video from Patrick Boyle is well worth a watch if you’re interested in the #AI Bubble: youtu.be/T-oXyXwD6sE tl;dr — the video is basically examining the question “How does all the noise about a $2 trillion IPO for #Anthropic make any sense?” The answer can be summed up as, “It doesn’t”. As always, he gets a little nerdy about the financial ins and outs, talking about how IPO valuations work and where all the money is going. But the punchline is basically that you have to believe in both astonishing growth for the industry *and* that Anthropic will capture a huge amount of that for such numbers to be remotely plausible. It’s also even more deadpan sarcastic than he usually is, which says something. But I really want to see a website for Boyle Compute now…
Why is GenAI a scam? Because even at its best it can only ever be successful if it can convince you that it’s something it is not. That’s its whole thing, it wants to be indistinguishable from the types of work it was trained on. Deception is part of the design. And most folks don’t like that.
There's one more point I'd like to make about the stochastic (i.e. random or unpredictable) nature of "Generative #AI" stuff: it does somewhat resemble an internal hhuman thinking process, something I've noticed upon introspection: there's some part of human consciousness that's trying out possibilities in the imagination. One meets a brand-new concept in reading, let's say, and then the brain starts juggling it around and comparing it with known information and so forth.
Show more of this post
An active mind is keeping up a steady simmer of basically stochastic activity, mental noise and internal chatter that sometimes leads to bursts of unexpected insight or intuition when there's a fortuitous synthesis of concepts in one's head. I point out: this process is only the START of thinking! But thanks to decades of #business propaganda and rubbishy biographies of CEOs and other such self-serving corporate marketing, "the West" has congealed around a desperately broken and frankly irrational notion of how people think, one that's suffused with thinly disguised Christian mythology. The devout Christian would say that if ideas come to their head, they're sent by God, and because their ultimate source is infallible and omniscient, there's no need for further reflection and internal critique--no need for THOUGHT, in other words. (cont'd)
It strikes me the same three-i's that dog IT contracts (indemnity, IP, and insurance) are the same three humanity is being hosed on by big-AI and big-IT with the fate of the planet freely swinging about. Their lawyers have so far outdone ours. There have been efforts to personalize forests […]
My own experience with an (early) proof assistant was that -contrary to Dijkstra- it is just a waste of time to write down a paper and pencil proof prior to formalising it. There's just too much freedom on paper to meet the rather stringent requirements of a proof checker; i.e., there might be a misfit. That would explain the alien reasoning. Dana is reading the output of the interaction between a language model and a proof assistant while it was ruthlessly trying to prove a goal.
Yes, I had to join another meeting to discuss how we could use AI more and the answers from the majority are less than enthusiastic not because we're afraid AI will take our jobs, but because AI has been useful in a very limited number of scenarios while it's expected to be universally useful.
Hearing lots of discourse about "stochastic parrots" recently? Here's a discussion of the backstory, both of the phrase itself and what's going on now. bloodinthemachine.com/p/the-tw…
This is a pretty substantial study - 718 firms. AI allows engineers to write crappy code faster - but devs end up having to spend more time reviewing & revising code so in the end there's no actual productivity gain.