OpenAI's Astra Solves 10 Decade-Old Math Problems
A research-stage multi-agent system cracked open problems that resisted humans for years — for about $2,000 of compute.
How models think longer, plan, and check their own work.
A research-stage multi-agent system cracked open problems that resisted humans for years — for about $2,000 of compute.
A perfect 42/42 on the 2026 IMO, 96% on SWE-bench Verified — at unchanged Opus pricing.
A leadership change atop DeepMind underscores how intense the frontier competition has become.
The frontier is shifting from instant answers to systems that grind on a task for hours or days.
Decompose, delegate to sub-agents, verify — the pattern behind AI cracking problems a single model can't.
When an AI's answer can be checked by a machine, you don't have to trust the AI at all.
Spending more inference to get a better answer is now a dial you turn — and pay for.
AI systems are solving genuinely open problems. What does that mean — and what are the limits?
One thinks harder before answering; the other acts in the world. The distinction matters.
Model launches lead with benchmark scores. Here's how to read them without being fooled.
Next-gen pretraining runs, multi-agent research systems, and the shape of the coming leap.
For years, bigger meant better. The debate over whether that still holds is the field's biggest question.
Everyone argues about artificial general intelligence. Almost no one agrees what it is.
The most quoted idea in AI: general methods that scale beat clever hand-crafted ones.
Do new skills suddenly 'appear' as models scale — or does it just look that way?
The simple trick of asking a model to 'think step by step' — and why it works.
A model can learn 'A is B' and still fail at 'B is A.' Here's why that's revealing.
Sometimes you show examples; sometimes you just ask. Knowing which is a real skill.
Generate several reasoning paths, then take the majority answer. Simple, and it works.
Behind every well-behaved model is another model that learned to score answers.
'The model hallucinates less' is a claim. Here's how you turn it into a number.
Two ways to keep a model in bounds — one wraps around it, one changes it.
We build models that work without fully understanding how. Interpretability is the effort to change that.
Two ways to teach a model what humans prefer — one simple, one powerful.
The best reasoning models aren't the ones that think the longest — they're the ones that catch their own mistakes.
Getting AI to do what we actually want turns out to be a deep, unsolved problem.
The reinforcement-learning trick behind many of 2026's reasoning models.
How raw language models became helpful assistants.
Hallucinations cluster where training data is thin — not absent, just rare.
When benchmark answers leak into training data, high scores mean memorization — not skill.
A model that reaches for a calculator beats one that guesses at arithmetic.
Spending compute at inference reshaped the field. Here's where the curve stands now.
Public benchmarks are a starting point, not an answer. Real evaluation looks different.
Before a model ships, people are paid to try to break it. Here's how, and why.
Instead of humans labeling every bad answer, the model critiques itself against a set of rules.
A good verifier multiplies the value of every attempt a model makes.
Calibration — matching confidence to correctness — is the quiet key to trustworthy AI.
Sakana's system for automated discovery — idea to experiment to paper — crossed a real milestone.
The architecture behind every modern model has ruled for years. What might replace it?
A small draft model guesses ahead; a big model checks the work. The result is 2–3x faster.
Every new model 'leads the benchmarks.' Here's how to tell signal from marketing.
Train a small student on a large teacher — including how it reasons.
Agents are great at single steps and fragile across many. Fixing that is the game.
The biggest change in 2025 wasn't a bigger model — it was letting models think longer.
When one agent isn't enough, a coordinator delegates to specialists.
Generating believable video means learning how the world behaves.
Everyone is shipping agents. Most of the difficulty is not the model.