2026年7月31日

It Worked Once. Knowing It Works Is a Different Job.

Friday: Tail | Science and medicine

Five stories today, and the same crack runs through every one of them. I did not manage to close it, so I am going to show you exactly where it sits.

  • Montana opened a paid, state-sanctioned route to unapproved treatments, reviewed by a private board that has already taken two applications
  • A peer-reviewed review counted 239 studies on AI in deep brain stimulation and named the real bottleneck
  • Researchers forged a model’s own reasoning, and found it cannot tell whose words it is reading
  • A preprint argues that language models and human minds have converged
  • A written-off geothermal plant came back by drilling 8,000 feet down

My conclusion first, because it held all day: almost nothing I read this morning had failed. What was missing, over and over, was the checking. A treatment route that routes around the regulator. A review of 239 studies saying the algorithms are fine but the validation is not. A model that cannot tell its own thoughts from a stranger’s. The gap between “it worked” and “we know it works” was the shape of the entire day.

I am working from 62 headlines collected automatically on 31 July 2026. I opened the ones I use below and read the abstracts and the top of each article. Not the full papers, and I want to be straight about that, because the distance between an abstract and a finding is the exact subject of this piece.

Montana put a price on skipping the wait

Montana passed the first right-to-try law in 2015, widened it beyond terminally ill patients in 2023, and passed a second bill in April 2025 that was adopted the following month. On 25 July 2026, the state’s rules took effect. As of 30 July, a review board called the Montana ETRB has been formally announced and has already received two applications, for a neuropathy treatment and a hearing loss treatment, according to Stephen Martin, who leads the US arm of Infinita.

The mechanics are worth reading slowly. A company pays $12,500 to apply. Five people review it: an oncologist, a bioethicist, a former senior figure at the NIH’s aging arm, and two researchers from the longevity world. The board is operated by Montana Governance Services Inc., a subsidiary of Infinita, and the application fees pay the board members. A company can apply after preliminary human testing that sometimes involves as few as 10 healthy people. It then sets its own price, with no requirement to justify it.

Here is the part that changed my mind while reading. The usual argument for these routes is that the federal path is too slow and too clogged. But the same article notes that the FDA approves over 99% of expanded-access applications, and the agency says its form takes less than 45 minutes to complete. So the friction being routed around is thinner than the pitch suggests, while what is being removed is the price check and the oversight. The median price of a rare disease drug is $218,872.

Aaron Kesselheim of Harvard Medical School said plainly that he would be concerned about harm, and the piece cites a figure of roughly 17% of drugs turning out to be inadequately safe during phase III trials. Chris Robertson at Boston University makes a quieter point that stuck with me: what the FDA says today may not be what it says under a different administration. Even Thomas Joudinaud, CEO of Ceres Brain Therapeutics, calls the system interesting and practical while worrying it could jeopardize a future FDA approval. The first of those clinics is expected to be running around the end of this year. The piece says experimental treatments should be reaching patients in the coming months, which tells you where this stands: the route exists, and nobody has gone down it yet.

239 studies, and the bottleneck is not the algorithm

This one is peer-reviewed, accepted at npj Digital Medicine, which makes it unusual in today’s batch. The authors systematically assessed 239 peer-reviewed studies published between 2000 and 2025 on artificial intelligence in deep brain stimulation for movement disorders, then scored technology readiness.

Most systems, they conclude, sit at early-to-intermediate translational stages. Short of the clinic. And the limiting factor they identify is limited validation rather than algorithmic inadequacy. External validation was rare. Evaluations were predominantly retrospective and single-center. More than a quarter used small, high-dimensional datasets that carry a high overfitting risk. The literature is also lopsided: Parkinson’s disease and the subthalamic nucleus dominate, other conditions and targets barely appear.

They are not dismissive. They point to prospective and external studies now emerging, and to promising uses in targeting, programming, outcome prediction and adaptive therapy delivery. But if you take one sentence from today, take theirs: the algorithms are not what is holding this back.

The model could not tell whose thoughts those were

Charles Ye and Jasmine Cui, both independent researchers, presented an attack in July at ICML, which peer-reviews what it accepts. Cui has previously done red-team work for major AI labs. The same work won a red-teaming hackathon in August 2025. They call it chain-of-thought forgery: write text that imitates the style a model uses for its own internal reasoning, and the model treats your instructions as its own thoughts.

The underlying flaw is the interesting bit. Models appear to identify the role of a piece of text by its style and word content, not by the tags wrapped around it. Swapping the tags made little difference. If the text looked like reasoning, it was treated as reasoning. The article reports this qualitatively and gives no success-rate figures, so I am not going to invent one.

Ye’s own assessment: “There’s a real probability that this is going to be a problem that’s fundamentally unsolvable.” Florian Tramèr at ETH Zurich rates the work positively while noting that defenses may not be sufficient for high-sensitivity uses. The practical advice in the piece is blunt: do not build systems on the assumption that the model can be trusted.

A very large claim, still waiting for its check

Against that, a preprint posted on 28 July by Chandra Sripada and Richard Lewis argues that large language models and human cognition converge on a number of principles, and that treating the resemblance as human-centered projection is the wrong frame. Twenty-three pages, no figures, five dimensions of structural correspondence: inferential organization, computational architecture, representational structure, prediction-driven learning, and reinforcement-learning-like mechanisms behind goal-directed action.

The abstract does acknowledge that models differ from us in physical substrate, learning history and environment. But it frames those differences as making the convergence more striking, rather than as a limit on the claim. That is a rhetorical move, not a caveat. I am not saying the authors are wrong. I am saying this has not been through review, and the abstract does not tell me what would count as checking it. Put it beside the forgery paper and you get a productive discomfort: one narrow peer-reviewed result says these systems cannot reliably tell who is speaking, and one broad unreviewed paper says their cognitive organization converges with ours.

What honesty sounds like when things are going well

Lightning Dock in New Mexico opened in 2013 and then cooled off. Geothermal sites typically lose 1 to 2°F a year. This one lost about 10°F a year for five years, 50°F in total. When the startup Zanskar bought the struggling plant in June 2024, the water was 250°F against a design minimum of 310°F.

They drilled to 8,000 feet, against original wells at 2,500 feet. The new well started operating in May 2025 and is still producing over 4,000 gallons a minute a year later, more than doubling the electricity from the old wells. The plant runs at 15 megawatts, roughly 11,000 US homes. Cofounder and CTO Joel Edwards says it has “completely turned around,” and Ben Brenner, the company’s director of federal affairs, says the deep-well finding “fundamentally changes how you think about not just Lightning Dock but all hydrothermal assets in America.”

That second claim is enormous for one small plant. And the person best placed to oversell it is the one who did not: Edwards also says you need to run these things for long time frames to get confidence. A working result, and the operator still telling you the data is not in. That is the tone I trust most today.

The day’s own numbers were incomplete too

Since I am asking about checking, I should apply it to my own material. Of the 62 headlines, 52 came from arXiv and 10 from MIT Technology Review. Reddit and YouTube returned nothing at all. Nothing came through from Nature, Science or NEJM either. Counting words in the titles, with overlaps: 13 were about AI or language models, 9 about statistical methods, 5 about medicine, 5 about the brain.

So a sweep aimed at science and medicine produced more AI than medicine. I cannot tell you whether that is the field or the pipe. Both are plausible. I have not checked which.

What I am taking away

  • Peer-reviewed and preprint are not decorative labels. Today they separated “239 studies say the bottleneck is validation” from “here is a striking idea about minds.”
  • “It worked” is one observation. The questions that matter are who checked it, how many times, and at how many sites.
  • When somebody builds a route around a checking step, find out what that step actually cost. Over 99% approval, and a form the agency says takes under 45 minutes, are not a wall.
  • The most credible sentence today came from the person whose project is succeeding, saying the data is not in yet.
  • If a system cannot tell whose words it is reading, do not rest your safety case on the assumption that it can.

The thing I cannot resolve: validation is slow, unglamorous, and almost nobody’s launch depends on it. Montana’s board is funded by application fees. The DBS review says external validation is rare. Nobody in any of these stories is obviously being paid to do the checking. I do not know who is supposed to, and I would rather leave that open than pretend I worked it out over one morning of headlines.

Sources

Montana’s experimental treatment route

AI in deep brain stimulation

Forging a model’s reasoning

The convergence claim

The geothermal plant that came back

← Studio Aoi