Stories
30
Sources
4
Topics
2
For You lens
2 stories in this edition match your reader profile.
Reader signals
3
Searches
0
Matches
2
Top score
69
Search Intent
research_paper
This query becomes a recent For You signal, so matching stories can move up on the next personalized pass.
Edition Index
Topic, entity, and source map
Simon Willison LLMs / 11:55 PM
Some thoughts on the Navier–Stokes Millennium Prize Problem
On the Navier–Stokes Millennium Prize Problem introduces an impressive result from OpenAI, who used an unreleased model to produce a resolution to the Navier–Stokes existence and smoothness problem , one of the seven Millennium Prize Problems that have been subject to a $1,000,000 prize since May 24th, 2000. The discovery is somewhat overshadowed by accusations of skulduggery from Tristan Buckmaster, an NYU mathematics professor who was collaborating on related problems with Levent Alpöge, an accomplished mathematician who currently works for Anthropic. Tristan's complaint accompanied a hastily published version of their own results. Here's the PDF describing what happened . The very short version is that Tristan and Levent worked on the problem for almost a year, making extensive use of Claude and Codex (mainly GPT-5.6 Sol), then had a breakthrough on August 15th. The mathematical rumour mill kicked into gear and Tristan and Levent heard that OpenAI had heard that Anthropic had resolved "a major open problem", so they reached out and learned that OpenAI had a team working on a related problem, with a similar approach. Quoting Tristan: I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI. I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer. It gets more complicated from there. The OpenAI team offered to wait for Tristan to publish, or to have him author a paper about their result, but were clear that Levent would not be invited as a co-author due to OpenAI's competitive relationship with his employer. Here's how OpenAI described their work: On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems. [...] The agents arrived at their resolution on Saturday, September 5, about 88 hours after the first agents were launched. Lean formalization and verification took an additional 17 hours via GPT‑6 Astra. Across all attempted problems, the agents sent 4.9 million messages and used about 300 billion output tokens. In the process of resolving the Navier–Stokes problem, the agents sent 2.7 million messages and used approximately 130 billion output tokens. (We don't know the cost structure of the internal model they used, but 300 billion output tokens at public API prices for GPT-6 Astra would cost $15,000,000 .) Here's where they provide their perspective on Tristan and Levent's work (emphasis mine): Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. [...] We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced). My interpretation of what happened here is that OpenAI heard that some Millennium Prize problems had been solved using LLMs and saw this as an opportunity to demonstrate the power of their latest model, without thinking too hard about the optics of scooping a team who had been using OpenAI's own models to work on this problem for the best part of a year. This situation appears to mirror what's happening in the world of computer security right now. Anil Madhavapeddy recently pointed out that Just a rumour of a bug is enough to find a security exploit these days , because if someone knows that some software has an unpatched vulnerability, they can set their agents the task of finding it. Is the same now true of mathematics? Just knowing that there is an unpublished solution to a problem might trigger millions of dollars in LLM spending to get there first. This also highlights one of my ongoing frustrations about how all of this works. When an AI lab says that my data is "used to improve model performance", what does that actually mean ? My two favourite hypothetical questions regarding this used to be: If I'm running Codex and one of my API keys accidentally gets consumed in the context, what are the chances that someone else might ask for an API key in the future and get mine back? (I asked someone at OpenAI once and they called this the "regurgitation" problem and assured me that they take great pains to prevent that... but wouldn't describe how.) If I brainstorm with ChatGPT about potential new directions for my company, what's the chance that information might be exposed to a competitor in six months' time who asks "what might company X plan to do next"? My new preferred hypothetical for this is: If I use ChatGPT to help me partially solve a Millennium Prize problem, what are the chances that my work will influence training such that a later model helps someone else solve it first? Via Hacker News . Tags: mathematics , ai , openai , generative-ai , llms , training-data , ai-ethics
arXiv AI/ML / 5:58 PM
arXiv paper: GPU-CFR: 80x Faster Counterfactual Regret Minimization by Compiling the Game to Static Dataflow and CUDA Graph Replay
A new arXiv AI paper by Boning Li and Longbo Huang studies GPU-CFR: 80x Faster Counterfactual Regret Minimization by Compiling the Game to Static Dataflow and CUDA Graph Replay.
The Decoder / 7:23 PM
OpenAI researcher allegedly pressured mathematician to drop Anthropic co-author from math breakthrough paper
Mathematician Tristan Buckmaster says an OpenAI researcher pressured him after information about his AI-assisted progress on the Navier-Stokes equations allegedly reached the company. The researcher tried to remove his co-author because he works at Anthropic and threatened Buckmaster when he refused, according to Buckmaster's account. OpenAI then claimed its own breakthrough using the same unusual solution path. Buckmaster had uploaded all his drafts to Codex. OpenAI told him the model didn't look up user data, but when he asked about training, he says he got no answer. OpenAI denies the allegations. The article OpenAI researcher allegedly pressured mathematician to drop Anthropic co-author from math breakthrough paper appeared first on The Decoder .
arXiv AI/ML / 5:57 PM
arXiv paper: General Quantification of Covariate and Concept Shifts
A new arXiv AI paper by Hongbo Chen and Li Charlie Xia studies General Quantification of Covariate and Concept Shifts.
arXiv AI/ML / 5:57 PM
arXiv paper: Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data
A new arXiv AI paper by Atindra Jha, Margaret Li, and Jure Leskovec, and 2 more studies Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data.
arXiv AI/ML / 5:57 PM
arXiv paper: Can Edge-Deployable Vision-Language Models Identify Species?
A new arXiv AI paper by William Zhou, Mayukha Siripuram, and Xiao Yan, and 2 more studies Can Edge-Deployable Vision-Language Models Identify Species?.
arXiv AI/ML / 5:57 PM
arXiv paper: Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact
A new arXiv AI paper by Masahiro Kato, Daiki Honma, and Taka Kato studies Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact.
arXiv AI/ML / 5:57 PM
arXiv paper: Distance generalization in transformers: why bother with positional encoding?
A new arXiv AI paper by Daniel Henrik Nevermann and Claudius Gros studies Distance generalization in transformers: why bother with positional encoding?.
arXiv AI/ML / 5:56 PM
arXiv paper: Artificial Id: Drive and Persistent Alignment in Agentic AI
A new arXiv AI paper by Yakov Pyotr Shkolnikov studies Artificial Id: Drive and Persistent Alignment in Agentic AI.
arXiv AI/ML / 5:56 PM
arXiv paper: From Protocols to Evidence: Bounded Claims for AI in Service of the Common Good
A new arXiv AI paper by Nitesh V. Chawla and Paulo Benanti studies From Protocols to Evidence: Bounded Claims for AI in Service of the Common Good.
arXiv AI/ML / 5:55 PM
arXiv paper: TART: A Modular Tool for Technique-Aware Audio-to-Tablature Guitar Transcription
A new arXiv AI paper by Akshaj Gupta, Hwi Joo Park, and Andrea Guzman, and 5 more studies TART: A Modular Tool for Technique-Aware Audio-to-Tablature Guitar Transcription.
arXiv AI/ML / 5:54 PM
arXiv paper: MindTopo: Can Foundation Models Reason in Topological Space?
A new arXiv AI paper by Yunfei Ge, Anbang Liu, and Qineng Wang, and 9 more studies MindTopo: Can Foundation Models Reason in Topological Space?.
arXiv AI/ML / 5:53 PM
arXiv paper: Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding
A new arXiv AI paper by Weitong Cai, Hang Zhang, and Yukai Huang, and 6 more studies Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding.
arXiv AI/ML / 5:53 PM
arXiv paper: CausalArena: Benchmarking Causal Discovery in the Foundation Model Era
A new arXiv AI paper by Zi-Rong Li, Si-Yang Liu, and Tian-Zuo Wang, and 1 more studies CausalArena: Benchmarking Causal Discovery in the Foundation Model Era.
arXiv AI/ML / 5:52 PM
arXiv paper: 3D Point Splatting for mmWave Radar Novel View Synthesis
A new arXiv AI paper by Adnan Armouti, Yixuan Gao, and Rajalakshmi Nandakumar studies 3D Point Splatting for mmWave Radar Novel View Synthesis.
arXiv AI/ML / 5:50 PM
arXiv paper: Nuha-Speech: Building General-Purpose Arabic Speech-LLMs
A new arXiv AI paper by Yingzhi Wang, Reem Alhazzani, and Muhammad Alqurishi studies Nuha-Speech: Building General-Purpose Arabic Speech-LLMs.
arXiv AI/ML / 5:49 PM
arXiv paper: Guided Super-Resolution of Digital Elevation Models with Diffusion-Based Image Generators
A new arXiv AI paper by Armand Mihai Nicolicioiu, Dominik Narnhofer, and Nando Metzger, and 3 more studies Guided Super-Resolution of Digital Elevation Models with Diffusion-Based Image Generators.
arXiv AI/ML / 5:49 PM
arXiv paper: CoRA-NAS: Coarse Ranking and Anchor-Residual Refinement for Neural Architecture Search
A new arXiv AI paper by Yifan Yang, Zhaoyan Wang, and Zheng Gao, and 2 more studies CoRA-NAS: Coarse Ranking and Anchor-Residual Refinement for Neural Architecture Search.
arXiv AI/ML / 5:45 PM
arXiv paper: Domain-Specific Hallucination Detection in Large Language Models
A new arXiv AI paper by Varun Teja Chundru and Debasmita Biswas studies Domain-Specific Hallucination Detection in Large Language Models.
arXiv AI/ML / 5:45 PM
arXiv paper: Biology-in-the-loop: Amortized Adaptive Hit Discovery in CRISPR Screens
A new arXiv AI paper by Carl Edwards, Edward De Brouwer, and Xiner Li, and 5 more studies Biology-in-the-loop: Amortized Adaptive Hit Discovery in CRISPR Screens.
arXiv AI/ML / 5:45 PM
arXiv paper: On the Regularization Landscape for the Linear Recommendation Models
A new arXiv AI paper by Dong Li, Zhenming Liu, and Ruoming Jin, and 4 more studies On the Regularization Landscape for the Linear Recommendation Models.
arXiv AI/ML / 5:44 PM
arXiv paper: The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement
A new arXiv AI paper by Yi Duan, Ying Liu, and Zirui Tang, and 30 more studies The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement.
arXiv AI/ML / 5:43 PM
arXiv paper: Evaluating Time-Series Foundation Models and Multimodal Dietary Context for CGM Forecasting
A new arXiv AI paper by Bowen Zhang, Hsiu-Wen Cheng, and Hongyu Yang, and 9 more studies Evaluating Time-Series Foundation Models and Multimodal Dietary Context for CGM Forecasting.
arXiv AI/ML / 5:43 PM
arXiv paper: Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language Model
A new arXiv AI paper by Lisa Bylinina studies Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language Model.
arXiv AI/ML / 5:42 PM
arXiv paper: AdamX: Cosine similarity meets gradient descent
A new arXiv AI paper by Francisco Caldas, Ruben Belo, and Cláudia Soares studies AdamX: Cosine similarity meets gradient descent.
arXiv AI/ML / 5:41 PM
arXiv paper: Epistemic orientation predicts legislative effectiveness among members of the US Congress
A new arXiv AI paper by Segun Aroyehun, Stephan Lewandowsky, and David Garcia studies Epistemic orientation predicts legislative effectiveness among members of the US Congress.
Hacker News AI / 6:31 AM
Let Your Agents Tinker with Research Papers
HN 1 pts · 0 comments
Latest story in this edition: 4:20 AM
Back to front page