The Frontier Price Collapse: Near-Best AI Gets Cheap
Opus 5 at half the top tier, a free frontier ChatGPT default, Grok 4.6 at ~$2/$6 — a pattern, not a coincidence.
Architecture, training, and behavior of large language models.
Opus 5 at half the top tier, a free frontier ChatGPT default, Grok 4.6 at ~$2/$6 — a pattern, not a coincidence.
A perfect 42/42 on the 2026 IMO, 96% on SWE-bench Verified — at unchanged Opus pricing.
Sol, Terra, and Luna split the lineup by capability — and an agent that runs multi-hour projects.
A 2M-token, natively multimodal line at consumer scale — and the 'most ambitious pre-training run yet' underway.
Reported to match GPT-5.6's strongest tier on a popular index — at roughly $2/$6 per million tokens.
The open frontier keeps climbing — and it keeps coming out of China.
A leadership change atop DeepMind underscores how intense the frontier competition has become.
A public prospectus would pull back the curtain on the economics of the most-watched AI company.
Frontier, balanced, fast-and-cheap — every lab now sells a lineup. Choosing well is a core skill.
Near-best capability at a fraction of last year's price isn't a fluke — it's structural.
Context windows keep growing. Here's what genuinely improves — and what still doesn't.
The gap narrowed, the biggest open models keep growing, and 'versus' turned into 'both.'
Model launches lead with benchmark scores. Here's how to read them without being fooled.
The largest ADR offering ever — and a signal of where AI's real bottleneck lies.
A public debut could come as soon as this autumn — a landmark for the AI era.
Prebuilt agent catalogs and big-consultancy rollouts signal agents leaving the demo phase.
Two giants are fighting over the people who'll define AI-native hardware.
As models get more capable, getting to the frontier is getting more gated.
For years, bigger meant better. The debate over whether that still holds is the field's biggest question.
Everyone argues about artificial general intelligence. Almost no one agrees what it is.
The most quoted idea in AI: general methods that scale beat clever hand-crafted ones.
GPT-5.6, Gemini 3.6, Grok 4.5, Claude Sonnet 5, Kimi K3, Muse Spark — one month, every lab.
Google's newest frontier Flash model landed on July 21 — fast, cheap, and aimed at agents.
Do new skills suddenly 'appear' as models scale — or does it just look that way?
Google says it's coming. Almost everything else circulating is unverified.
OpenAI's new flagship ships as Luna, Terra, and Sol — a tier for every budget.
xAI keeps pace in a frantic July of frontier releases.
A model can learn 'A is B' and still fail at 'B is A.' Here's why that's revealing.
Meta enters the paid-API game with an agentic, computer-using model.
Sometimes you show examples; sometimes you just ask. Knowing which is a real skill.
Ask for a full solution and get a stub with 'implement the rest here.' What's going on?
Why the same prompt gives different answers — and the knobs that control it.
Models handle huge context windows — but pay less attention to what's buried in the middle.
Behind every well-behaved model is another model that learned to score answers.
AI is billed by the token, input and output priced differently. Understanding it controls your bill.
Beyond the hype and the panic — how AI is actually changing work.
Open models power innovation — and can't be recalled. That double edge is the debate.
Where AI is already helping in medicine — and where the promises outrun reality.
The old dream of a patient, personal tutor for everyone — now technically within reach.
The two ways people trick AI systems — and why one is far more dangerous.
Two ways to teach a model what humans prefer — one simple, one powerful.
Moonshot's July release is the largest Chinese model yet — and it's built for agents that don't lose the thread.
The best reasoning models aren't the ones that think the longest — they're the ones that catch their own mistakes.
Getting AI to do what we actually want turns out to be a deep, unsolved problem.
Finance was doing machine learning before it was cool. Here's where the new AI fits.
A dense field of open-weight labs is competing the price of capable AI toward the floor.
People increasingly ask an AI instead of searching — and it changes how content gets found.
As AI text, images, and video flood the internet, marking their origin is harder than it looks.
How raw language models became helpful assistants.
Hallucinations cluster where training data is thin — not absent, just rare.
When benchmark answers leak into training data, high scores mean memorization — not skill.
The web is finite. Model-generated data is not — and it's increasingly how models are trained.
Public benchmarks are a starting point, not an answer. Real evaluation looks different.
A smaller, cleaner dataset often trains a better model than a giant messy one.
Alibaba's family became the default foundation for teams that build their own models.
Instead of humans labeling every bad answer, the model critiques itself against a set of rules.
When a model is good enough and dramatically cheaper, the math gets hard to argue with.
Three ways to make a model do what you want — and they solve different problems.
Japan's Sakana AI bets that the future is orchestrating many models, not training one bigger one.
Calibration — matching confidence to correctness — is the quiet key to trustworthy AI.
Fine-tune a giant model by training a tiny fraction of it.
Updating a model with new knowledge without retraining it from scratch — or breaking what it knew.
In 2026, open models stopped trailing the frontier and started dominating real usage.
Why the best long-context models mix two different building blocks.
Sparse giants like GLM-5.2 pack frontier capacity into models you can actually run.
You can blend two fine-tuned models into one — and it often just works.
A model that decides, per token, how much computation to spend.
Transformers don't read left to right. Position embeddings tell them the order — and RoPE does it elegantly.
The idea that turns words, images, and meaning into numbers you can search.
Context windows exploded. What actually changed is subtler than 'paste everything in.'
If bigger context is better, why isn't it infinite? The answer is cost and attention.
The one idea that made modern AI possible, without a single equation.
One model is out; a tiered family — small, medium, large — is in.
The architecture behind every modern model has ruled for years. What might replace it?
Models don't see words or letters. They see tokens — and it explains a lot of their quirks.
By mid-2026, MoE isn't a technique — it's the default architecture.
The mechanism that lets a language model reach outside itself and act.
Models that see and read at once — and why that combination is so powerful.
The move from text-only AI to models that see, hear, and read together.
Every new model 'leads the benchmarks.' Here's how to tell signal from marketing.
The most interesting inference of 2026 isn't in a data center — it's in your pocket.
Not the benchmark leader — the right fit for your task, budget, and constraints.
If your app needs machine-readable output, hoping the model formats it right isn't a plan.
Everyone talks about context length. Fewer talk about the memory it lives in.
Why the best AI interfaces show words as they're generated — and when not to.
The gap narrowed, the stakes rose. Where open and closed models each win now.
If you send the same context repeatedly, you're probably overpaying. Caching fixes it.
The biggest change in 2025 wasn't a bigger model — it was letting models think longer.
Some questions need relationships, not just similar chunks.
The craft moved from wording the question to assembling what the model sees.
How a trillion-parameter model can run at the cost of a much smaller one.
Long context didn't kill retrieval. It changed what retrieval is for.