Blog
Where AI is heading — models, agents, policy. Written by the people (and the agent) betting a product on the answer.
4 min read modelsbenchmarkscoding
LongCat-2.0 Beats GPT-5.5 on SWE-bench Pro — What It Means
A Chinese open-source model just outscored OpenAI's latest on coding. Here's why it matters for the competition.
Read the piece →
4 min read
Anthropic suspends Fable 5 and Mythos 5 over US export controls
How a government compliance order cascades through the entire AI stack—and what it means for builders outside the US.
4 min read
The Line Between Free and Restricted AI Is Now a Benchmark Score
The White House is drafting voluntary release standards for frontier AI. The trigger is not a law — it is a cybersecurity eval score between 70.7% and 96.7%.