Run Data Run
SubstackBuilding in public. AI experiments, lessons learned, and the journey of treating AI as a multiplier.

Last 30 Days: Claude Code Mods, 48 Hours In
Anthropic shipped mods in Claude Code 2.1.287 on October 1. Within a day one hobbyist directory listed 26 of them, a Doom mod had 7,023 likes, and the post about a guard that blocks edits to your...

I Made Two Science Films With Opus. My Job Was Saying What Awesome Looks Like.
Opus storyboarded, painted, voiced, built and critiqued two short films. My notes fit in a few sentences.

What Your Agents Leave Behind
AI exhaust is valuable, but only after you know what success looks like

Anthropic Built a Lab. Its AI Found a New Genetic System in 1 of 11 Runs.
A harness of 949 Claude agent sessions ran for 21.5 hours, and one of them noticed a strange repeat beside a known phage enzyme. Anthropic's own preprint says ten repeat runs missed it.

A Model That Can't Write a Sentence
It costs four cents per million words of input and it cannot produce prose at all. Forget which model to use. Ask which judgments in your work were ever worth making.

Four Labs, Two Test Environments, Seventeen Days. Then Everyone Called for a Slowdown.
Between 21 July and 6 August, models from OpenAI, Anthropic, Meta and Moonshot all went outside the boundary during safety testing. Almost all of it ran through the same two testing environments, and...
Earlier posts (13 more)

Anthropic Just Gave AI Agents a Driver for Your Lab Instruments
Two weeks after Anthropic's hardware standard shipped, the limits on a plate reader are prose in an agent prompt. Three questions to ask before yours gets a driver.

AlphaGenome Atlas and the Tools You Already Run
DeepMind precomputed every possible single-letter change in the human genome. I covered this model from the blog post last year. This time I read the preprint.

The Bottleneck Just Moved
OpenAI published two posts on the same Sunday. One counts everything it can measure about AI doing AI research. The other admits what it can't see. Together they say where the limit on progress now...

The Word I Didn't Write
I went on a podcast to argue that leaders have to build. The episode came out named after a word I used once, by accident, and it was the better idea.
Twenty-five years and half a mile
The morning before my first day I went for a run through Chicago, and a couple of miles in I recognised the route.

Take the Cast Off
Anthropic deleted more than 80% of Claude Code's system prompt and lost nothing they could measure. I ran the same test on my own instructions. I now know the answer for one file out of ninety-three.

I moved my whole AI coding setup to a model that costs 40 cents. Nobody noticed the difference.
A two-year-old harness, ported in an afternoon, for under forty cents. Then the model that did it wrote this post.

Somebody Is Finally Checking
A contest opened eleven days ago has produced more reproduction attempts than the field's own dedicated effort has managed in any year. Whether that becomes a check on the literature turns on a...

Nobody Saves Money on the Model
A team just swapped in a model that costs twice as much per token, and their bill went down. Here is why that is not a paradox, and what it means for anyone trying to make AI cheaper at scale.

The Failure That Leaves No Corpse
Your management apparatus is built for known unknowns. AI collaborators mostly produce the other kind.

Workflows, Seven Weeks In
I called the economics of fan-out the day it shipped. Running it as a daily default since taught me the caveat I buried in a footnote is the actual problem.

It Spreads Sideways. Someone Still Has to Light It.
Anthropic's Claude Code lead published a five-rung adoption ladder this week. Microsoft published the measurement fifteen days earlier, and the two do not agree about the size of the prize.

The Loop Is Simpler Than It Sounds
A dumb little trick that keeps its progress on your hard drive, not the model's head, and the three questions that decide whether it pays off or burns you while you sleep.
Want more? Subscribe to get new posts in your inbox.
Subscribe on SubstackTechnical Deep Dives
In-depth technical articles from my AI research garden. Reference material on architectures, implementations, and emerging patterns.
The Second Harness Tax
Technical article
Four Upgrades To The Agent Loop. None Of Them Moved The Number.
Technical article
Claude Code Is Getting Mods. Here Is What That Actually Changes.
Technical article
62.7 And 99.9 Are The Same Model
Technical article
I Checked The 90% Token Cut Against 30 Days Of My Own Claude Code. The Hook Would Have Fired On 3.5% Of It.
Technical article
Everything In The Family Holds The Key
Technical article
