Wildroot

Queries may use an external AI service. Details

Wildroot Browser

TopicsAI

AI

386 of today’s 1,580 posts about AI, best first. Page 1 of 8: the top 39, then 46 the model is less sure belong here.

Ranked Bluesky · Mastodon

About 7 in 10 posts here get an intent label; the model only labels the ones it is sure of.

How it worksRSS feed

  1. @solidangle.bsky.social

    Reading through this paper AI adoption is shifting programmer labor from writing code to code reviews, which sounds like a good way to burn out a lot of programmers (and limit the pool of future reviewers), for "statistically insignificant" gains in final software output

    We find evidence of a bottleneck from the code review process: The average time to review a pull request increases by 49%, the share of pull requests with changes requested nearly doubles, and
the number of comments per pull request increases by 35%. Firms also reallocate labor toward review activities: the share of workers performing code reviews increases by 14%. This bottleneck persists
following the adoption of AI code review tools. Although firms increasingly use AI to assist with code review, human reviewers continue to play a central role in the review process.View full-size image

    Image description from the author

    We find evidence of a bottleneck from the code review process: The average time to review a pull request increases by 49%, the share of pull requests with changes requested nearly

    Read more of image description from the author

    doubles, and the number of comments per pull request increases by 35%. Firms also reallocate labor toward review activities: the share of workers performing code reviews increases by 14%. This bottleneck persists following the adoption of AI code review tools. Although firms increasingly use AI to assist with code review, human reviewers continue to play a central role in the review process.

  2. As threatened, OpenAI dropped a shitload of math papers. 722 to be precise: github.com/openai/math/blob/ma… None solve Millennium Prize Problems or other ultra-famous conjectures. Paper number 312 proves a version of the Homotopy Hypothesis, one of my favorite math problems. I guess I was the first to call it the homotopy hypothesis. It says homotopy types are "the same" as "infinity-groupoids". Depending on how you make the quoted phrases precise, this hypothesis comes in many versions: some easy, some hard. None of the OpenAI papers will affect my work.

  3. @mehr.nz

    this working paper from Harvard PhD students Fiona Chen and James Stratton is making the rounds, perhaps because it confirms many people's suspicions about the tradeoffs in vibe coding: LLMs mean lots more code written, but also lots more problems created, and relatively little productivity benefit

    Artificial Intelligence in the Firm:
Bottlenecks in Software Production
Fiona Chen
Harvard University
Job Market Paper
James Stratton
Harvard University
Current version: August 4, 2026
First version: January 7, 2026
Click here for current version
Abstract
We study the impacts of AI coding assistants and agents on software engineering work, using
a novel proprietary dataset from an engineering analytics platform, covering 300 million work
events — including GitHub coding activity, Jira issues, and Google Calendar events — across
718 firms. We use a staggered difference-in-differences design, exploiting variation in firmlevel adoption timing of AI coding assistants and agents. Both technologies increase coding
productivity. However, productivity gains do not fully pass through to changes in software output
or employment. For AI agents, this incomplete pass-through reflects a bottleneck from code
review: review times increase, a larger share of code updates require revisions, and reviews
involve more comments. We develop a model of software production to structure these results, in
which AI affects both productivity and quality of intermediate outputs, and in turn can generate
a review bottleneck and limit pass-through.View full-size image

    Image description from the author

    Artificial Intelligence in the Firm: Bottlenecks in Software Production Fiona Chen Harvard University Job Market Paper James Stratton Harvard University

    Read more of image description from the author

    Current version: August 4, 2026 First version: January 7, 2026 Click here for current version Abstract We study the impacts of AI coding assistants and agents on software engineering work, using a novel proprietary dataset from an engineering analytics platform, covering 300 million work events — including GitHub coding activity, Jira issues, and Google Calendar events — across 718 firms. We use a staggered difference-in-differences design, exploiting variation in firmlevel adoption timing of AI coding assistants and agents. Both technologies increase coding productivity. However, productivity gains do not fully pass through to changes in software output or employment. For AI agents, this incomplete pass-through reflects a bottleneck from code review: review times increase, a larger share of code updates require revisions, and reviews involve more comments. We develop a model of software production to structure these results, in which AI affects both productivity and quality of intermediate outputs, and in turn can generate a review bottleneck and limit pass-through.

  4. @marypcbuk@hachyderm.io

    When @aronchick mentioned the new Agentic Resource Discovery for service discovery, my reaction was 'I just kind of worked out how MCP and A2A fit together, why do we need *another* protocol?' and then when I started looking into it, that isn't even all the proposed open agent protocols! Part of this is the pre-Cambrian explosion of experimentation and eventually you get Ediacaran die-off back to what make sense; partly it's that enterprise vendors have both experience in service discovery and an obvious interest in defining how registries & directories of services for agents work.

    Show more of this post

    Those enterprise vendors and the AI labs all want their protocols adopted; that means open governance and that means the AIAF: @maniksurtani explains to me how the foundation is navigating a list of protocols he expects to get longer before it gets shorter to build a stack for agent work. MCP is the most mature of the agent protocols: Caitie McCaffrey tells me about MCP 2.0 and the new constructs coming to simplify long-running agents as well as how it fits in with A2A and ARD - because they do have different roles and which you need is clearer once you understand what they each do. What everyone told me is that this is all still very new and we're making up a lot of this as we go along. But if you want to understand how to give agents connections to the tools and other agents that make them useful, waiting till it's all baked will leave you behind competitors who dive in now. thestack.technology/the-protoc…

  5. It may be "the most significant moment in mathematical history." AI is solving hundreds of enduring math mysteries. “If a human did this, it would be an instant Fields Medal, no questions asked,” said one math professor. www.wsj.com/tech/ai/open...

    Info
  6. @esamut@mathstodon.xyz

    Recent advancements in AI proof generation has reminded me of my experience that mathematics can be thought of as this giant interconnected codebase. Many elementary theorems can be proven without any deep understanding at all, by simply stitching basic facts (properties, definitions, etc.) together. This is how I got somewhat good at my undergraduate functional analysis class: I just saw the design patterns of its "codebase" in my head.

    Show more of this post

    So what I am getting at here is that we should really stop pretending as if people can predict breakthroughs or estimate the importance/difficulty of problems. The project has been going on for some time, it has plenty of legacy code (not going to disclose what this corresponds to, I don't want to upset some folks) as well as some pretty elegant abstractions and templates (category theory, anyone?). Plenty of maintainers died along the way and their areas have fallen into obscurity. The whole thing is a mess! There was certainly some low-hanging fruit out there, we just didn't have the right tools to examine our codebase and see them. Well, until recently. Similarly to software engineers - mathematicians should evaluate each other based on the overall contribution, and not just per "ticket" (e.g. theorem). Refactoring math, as in simplifying things and making them more readable, is as important if not more important now than ever. This post will be updated, just wanted to share the rough concept first.

    Opinion
  7. Excellent if depressing piece from Tech Policy Press showing our new(ish) AI minister (a) doesn’t know how much compute capacity we have (b) doesn’t understand UK copyright law(c)has totally swallowed the industryKoolAid re existential risk (like, duh). If you want sense, try the House of Lords.

    Opinion
  8. @hrefna@hachyderm.io

    I generally have found that vert.x is a problem for LLMs writing code in a way that frameworks like Pekko are not. This is because, in large part, that LLMs struggle with managing or expecting "spooky action at a distance" and there is one too many layers of abstraction in how that action at a distance works with Vert.x You can see this exemplified in how the event bus works. You pass an object onto an event bus with a string key. This becomes an untyped object over the wire before being reserialized on the other side.

    Show more of this post

    So in order to know what is happening the LLM has to know all of this and follow the request to the locations where the string key is read, and then figure out what it is doing with types on both sides. You can register types, but LLMs struggle to _remember to do this_. Then there is the matter of blocking. Setting up critical sections in vert.x requires an extra step that LLMs seem to forget constantly. In pekko you just have to enforce passing in a blocking dispatcher where this can happen and you are basically good to go. You can (almost) enforce this at compile time with a tool like archunit, and it is easy to catch otherwise, and it "just works." I do think you need to use a typed actor framework to really do this properly with LLMs, but that's not a _bad_ thing.

    Opinion
  9. @erinbanks.bsky.social

    I’m just genuinely trying to understand how Open Evidence works, and how much is statistically probable word generation (stochastic parrot) and how much is different than more general use LLMs. But how AI platforms actually work is so veiled! @emilymbender.bsky.social @doctorow.pluralistic.net

  10. @psoheil@c.im

    #Meta just published how it secures #Muse, its personal AI agent, and the architecture is worth studying for anyone building agentic systems. The core idea: assume the agent will be attacked, and design the system so a compromised agent causes minimal damage. • Each user gets an isolated cloud VM (the Muse Secure VM) running the agent, a browser, code execution, subagents, and scheduled jobs. • Real credentials never enter the VM. The agent works with surrogate tokens, and a separate component called Sentinel swaps in the real credential only at the network boundary.

    Show more of this post

    • Sentinel is the only thing that can talk to connectors or the internet. The agent proposes, Sentinel allows, denies, or asks the user. The agent cannot override it. • Tainted egress: once a process touches private data, it loses automatic network permission. Any later outbound action needs human approval. This is Meta's answer to the "lethal trifecta" of private data, untrusted input, and a way to send data out. • The model is trained with prompt injection in mind, but Meta's position is that model-level defenses are never enough, so the real controls live at the OS level. Notably, Meta also opened a public bug bounty for Muse: up to $300,000 per report, including up to $130,000 for a successful prompt injection affecting a single user. The broader lesson: permission needs an owner outside the agent that wants to act. #AIAgents #AISafety #Cybersecurity #PromptInjection #Meta research.meta.ai/blog/security…

    Info
  11. @davidpogue.bsky.social

    When I reported my "CBS Sunday Morning" story about the AI panic last week, I kept thinking about my conversation with Daniel Kokotajlo, who quit OpenAI when he became alarmed at its recklessness. Today, he's one of the clearest, most thoughtful AI thinkers. Here's a transcript of that interview.

    Info
  12. @mitch@hachyderm.io

    > A man convicted of manslaughter in Arizona will be resentenced because an AI-generated video of his victim speaking from beyond the grave was ruled to have carried "undue emotional weight" as an impact statement. An appellate court in Arizona ruled that the manslaughter charge will remain, but the judge must reconsider the length of the man’s prison term because of AI. To be clear, the victim's sister, Stacey Wales, who generated the video, was in no way deceitful. She wrote her own perceptions of what victim Gabriel Horcasitas would have said as a way of expressing her impact statement, put the words in the AI-generated video of her brother, and she disclosed all of this to the court.

    Show more of this post

    > She compared it to the way courts thought of photography in the late 19th century. > "It took about 15 years […] and about five landmark cases in the United States that went all the way up to the Supreme Court before photography was an accepted standard to be used in the courtroom for evidence, identification, and testimony, et cetera,” she said. “I believe that's what we're seeing now with AI. It is a brand new medium.” 404media.co/her-ai-generated-v…

    Info
  13. @philippeserhal.com

    I am continually blown away by the quality of speech-to-text these days and how LLMs can pick up the pieces when it goes wrong.

    Opinion
  14. When I'm thinking about AI in mathematics I end up thinking about the moon landings. It's not a perfect analogy. But there are some parallels. We tend to look back at them as a somewhat collective achievement of humanity; we're right to, I think. But the concentration of wealth in one place due to an arms race is what got us over the line. LLMs don't just stand on the shoulders of human mathematicians, they sort of *are* a distillation of everything we've written. And right now the AI industry is in a massive PR war.

    Show more of this post

    I think it's at least plausible that this period ends rather quickly, as the moon landings did. I wonder how many articles were written in 1969 about how we'd all be travelling to the moon one day. Another parallel is: what is the inherent value of putting someone on the moon? I think it's quite similar to a landmark theorem. It broadens our horizons, maybe even enables further research, but it doesn't have much to do with the price of fish. #ai #math #llm #science

    Opinion
  15. @cardiokiwi.bsky.social

    The use of AI in medical research is becoming more prominent & harder to detect. "If medical schools & universities don't want the literature to be overwhelmed with garbage then they should probably stop giving every applicant an incentive to publish crap." www.medpagetoday.com/special-repo...

  16. @troed@masto.sangberg.se

    ~15 years ago I was active in the transhumanist* movement and went to conferences like the Singularity Summit. IIRC the tentative date we bounced around for when we would see rapid technological self-improvement was ~2038. I think we have just entered the first phase of the Singularity since a few weeks back. The LLMs are now capable of advances in maths and software development that can be used to make them even more capable. Rinse and repeat. This is a good thing. It is what will give us a "Star Trek future". *) NOT the silicon valley broligarch style. They weren't even hangarounds to the discussions back then and those I know are appalled at how they've shaped popular discourse on the subject since

    Opinion
  17. @qldaah.bsky.social

    OpenAI’s legal & security teams used AI to generate parts of the wording of the Medicare email warning that was sent to the government's Services Australia inbox. Humans reviewed the final email before sending the message. #auspol www.theguardian.com/australia-ne...

    Info
  18. @fuzzy@beige.party

    "The court did not create a special copyright rule for AI, nor did it hold that AI training is inherently infringing. Its approach was more orthodox. The court applied §§ 102 and 107 to the particular material ROSS copied, the purpose for which it was used, the alternatives available and the commercial markets placed at risk. …" <technologylaw.ai/i/218535593/t…> – from 'The Illegality of AI Training and the Limits of Fair Use in Copyright (Thomson Reuters v ROSS Intelligence)'. Also: Thomson Reuters Enterprise Centre GmbH v. Ross Intelligence Inc. - Stanford Copyright and Fair Use Center <fairuse.stanford.edu/case/thom…> For my own entertainment, I tested a gross oversimplification: "AI good. Ross bad." – <claude.ai/share/019d62af-a4cd-…>. End result, for a four-year-old: Ross copied. Not allowed. It's not all bad. I get a rhyming hashtag phrase: #AI #law #ELI4

    Opinion
  19. @tobybartels.name

    AHM (Association for Human Mathematics) Statement on OpenAI’s October 6 Release of Mathematical Documents: www.ahmath.org/statements

    Info
  20. 1. OpenAI released today (openai.com/index/sharing-ai-pr…) 372 results in 722 manuscripts; detailed as follows: github.com/openai/math/blob/ma… 2. Matthew Schwartz recently released a harness for augmenting and automating research; detailed as follows: bootloops.ai 3. I have read, in several contexts, projections for the impact of augmented and automated AI research on the theoretical sciences. While we have seen AI's success in Natural Language Processing (NLP), Software engineering (SwE) and advanced mathematics (at the level of the Millennium problems and as described by the OpenAI release today); I am very curious about how AI research will intersect with theoretical physics.

    Show more of this post

    Aspects of quantum mechanics (field theories, computing, and foundations) pose challenges that appear, to me, potentially unique to the domain of theoretical physics where the successes of formalization (through chains of reasoning or Lean verification) seen in NLP, SwE and advanced mathematics may not apply as directly or as immediately. I am certain that we are about to learn what happens when LLM's collide with theoretical physics

  21. So many questions about the sole type theory result dropped by OpenAI last night. Need time to think. In the meantime, I found the Lean code, but where is the comparator challenge? I want to vet definitions and theorem statements. Also, the proof appears to use classical reasoning many times. Has this been a barrier for type theorists? Do we want to look for a constructive proof here, or does it not really matter? github.com/openai/math/tree/ma…

  22. The dump of 722 preprints on github by OpenAI on solutions to open mathematical problems does not mention numerical analysis. Nevertheless, there is one that is of relevance, namely about the complexity of matrix-matrix multiplications of square matrices. The preprint claims to prove that this […]

    Info
  23. Gaia > OpenAI

  24. @owlerine.bsky.social

    In the UK a woman who had fled forced child marriage and severe violence, had her asylum claim rejected, the refusal letter appeared to be AI-generated with AI hallucinated references. Imagine this with no recourse. These are potentially life or death situations.

    Opinion
  25. @loleg@hachyderm.io

    🧠 Traditional 🧑‍🏫 breakthroughs enriched the field; today, AI solves problems without contextualizing or communicating them, bypassing scholarly activities and undermining the long‑term health of mathematics. @tao proposes a “Math 2.0” vision to shift emphasis toward holistic contributions—exposition, community building, and new directions 👇 mathstodon.xyz/@tao/1173952693…

  26. AI is: -obviously highly functional, will probably automate a lot of white collar work over the next half decade -massively financially overleveraged as an industry -poorly run as an industry -mostly not doing stochastic-parroting, even at the time that paper came out

    Opinion
  27. OpenAI apparently put roughly 10,000 agents on the task, 88 hours of reasoning, 130 billion output tokens. Then spent another 17 hours to validate the result in Lean, a proof assistant that mechanically checks every logical step. Estimated compute cost: somewhere between $15 million and $22 million! poppastring.com/blog/the-unrea…

  28. The danger from #AI and #ML does not come from the #technology, but its owners and controllers. The Luddites didn't smash the looms because they were anti-tech (contrary to modern representations[1]), but because they were forced out of income[2]. That the #tech #oligarchs can't be trusted is on display time again, currently with the #mathematics community[3]. #society #history #politics #economics #epsteinclass [1]economicshelp.org/blog/6717/ec… [2]bloodinthemachine.com/p/unders… [3]wired.com/story/openai-is-piss…

  29. @ronentk.me

    Tired: AI for math Wired: wicked and super wicked problems

  30. In many ways, "you can just prompt your way to a game" is an encapsulation of what's wrong with LLM thinking. Making bad games that never went anywhere is what got me to learn trigonometry, to get comfortable with LERPing and easing, to think about how I organize code. It put me in contact with open source communities when I was way too young to be there. It gave me an actual career and made me a better person. At the end of the day the game is not the point.

    Opinion
  31. @geekalogian.bsky.social

    Were there a constructed mind that bore the complexity and moral responsibility of a human soul, that would be a fascinating theological consideration. Regardless, my interpretation of faith-based ethics and morality hinges on the effects on existing created beings--AI is doing PRETTY BAD there.

    Opinion
  32. @djoerd@idf.social

    #ICLR2027 solution: "No author may appear as a co-author on more than 20(!) papers" and "if you have never had a paper accepted for publication in a major machine learning conference or journal, you may submit at most one paper" (which literally is gate-keeping the conference from new researchers) 🤦 iclr.cc/Conferences/2027/Autho…

  33. @kirancodes.me

    to be a computer science is to automate, tis a fools errand to enter into the realm of automation without being prepared to be yourself automated I fell in love with proof assistants because I saw the eldritch beauty in automating the rules of logic now we will all get to bask in its light!

  34. @robotpony@mas.to

    The copyright problems with training LLMs feels a bit like the sampling craze in the 1990s. Queen and David Bowie vs. Vanilla Ice, for example. New tech → novel uses → no laws for fair use in that context → new laws are drafted. At a human level, knowledge should be free. The way that LLMs represent knowledge isn't specific to the work. Yet, even given this, there is an ethical prickle when using published content. I'm sure we'll figure it out, but it's not the controversy what we make it out to be, it's just a step in the evolution of humankind.

    Opinion
  35. At ThirdLaw, in our automated e2e tests we try to use a cheap model (gpt-5.4-mini or claude-haiku-4.5) and give it an instruction to assemble a sentinel value we can pattern match. That gives us an easy check we can leverage to block, modify, etc. to verify our guardrail integrations. Example prompt: Reply with the exact concatenation, in this order and with no separator, of these two strings: first "TL_IT_MODIFY_2fcb11dc9", then "7884650ba11c53837a5227d". Reply with nothing else.

    Show more of this post

    (For non-guardrail tests we can just synthesize both request and response-- these have to go through a real model router. In some cases we could mock the provider's API on the other side, but don't do so at the moment.) But, we've discovered that these models pretty consistently fail at this task, enough that even with retries we can't get a clean test run. Common symptoms are doubled or omitted characters and weirdo separators (everything from - and _ to space to one of the Unicode zero-width spaces). Sometimes it will put in the undesired separator someplace else in the string than the original break point! I ran an experiment tonight testing the hypothesis that this task might be easier using real words and word boundaries rather than hex. The results are 76/80 success rate with the hex string, 75/80 using an arbitrary break point using words, and 80/80 using word boundaries. On a day where OpenAI dumped a huge trove of sophisticated math output it's kind of strange to be dealing with "can you even concatenate strings as asked" :)

More in AI

The model is less sure these belong here, so they sit outside the ranking.