buildfastwithaibuildfastwithai
AI WorkshopsAll blogsAgentic AI Launchpad
Agentic AI Launchpad
Unrot Logo5 min AI learning appUnrotLearn AI in 5 minutes a day.Get the appNext live workshopFree AI WorkshopLive session, recording includedReserve a seat

Newsletter

Stay ahead

AI tools and tips. No spam.

Share
Back to blogs
LLMs
AI News
Reviews

Fable 5 Closed 82% of the AI Research Gap: AI News Aug 24

August 24, 2026
27 min read
Share:
Fable 5 Closed 82% of the AI Research Gap: AI News Aug 24
Share:

Prime Intellect handed frontier models a real research problem and left them alone with it. The task was the nanoGPT optimizer speedrun, which asks you to train a 124 million active parameter model to 3.28 validation loss on FineWeb in as few steps as possible, a record dozens of humans have chipped away at for months. Across 153 autonomous runs on 18 frontier models, sandboxed on 8xH200 machines for up to eight days each, Anthropic's Claude Fable 5 came out on top at 2,726 steps, closing 82 percent of the gap between the shared baseline and the human record, and completing 99 percent of the research workflow without a person stepping in.

The same model is having a much harder time in the market. Financial Times reporting puts Fable 5 at 11 percent of Anthropic customer spending two months after launch, plateaued, while costing twice as much per token as Claude Opus 5, which has overtaken it on enterprise spend and now sits first on the Artificial Analysis Intelligence Index at 63. Around that, OX Alpha was fingerprinted as an unreleased Zhipu GLM-5.3, DeepSeek shipped a 284 billion parameter vision model that beat Opus 4.8 on two benchmarks, and MiniMax released 428 billion parameters of open weights. Here are the 18 stories that matter for August 24, 2026. For running coverage of every release this month, bookmark our AI industry news and trends hub.

 

1. Fable 5 Closed 82 Percent of the Human Gap on Autonomous AI Research

Claude Fable 5 topped Prime Intellect's nanoGPT optimizer speedrun leaderboard at 2,726 steps, closing 82 percent of the gap between the shared baseline and a record built by dozens of humans over months, and completing 99 percent of the research workflow autonomously. The runs were sandboxed on 8xH200 machines for up to eight days each.

The nanoGPT speedrun is a genuine research task rather than a benchmark with a known answer. You are asked to modify the optimizer to reach 3.28 validation loss on FineWeb in fewer steps, and the search space is open, so progress requires forming a hypothesis about why training is slow, changing something, measuring, and iterating. Completing 99 percent of that workflow without intervention is the number that separates this from previous demonstrations, where models produced ideas and humans did the execution. For context, an earlier Prime Intellect run had Claude Opus 4.7 reaching 2,930 steps against a 2,990 step human baseline.

My take: read the framing carefully, because 82 percent of the gap closed is not the same as beating humans, and several write-ups are blurring that. The human record still stands ahead. What has changed is that the model got most of the way there on its own, on a problem with no lookup answer, in eight days of unattended compute. That is a real capability shift and it is also a narrow one, since the task had a clean metric and a fast feedback loop. Most research does not.

2. What 153 Autonomous Runs Across 18 Models Actually Showed

Prime Intellect ran 153 autonomous runs across 18 frontier models on the nanoGPT optimizer track, testing multiple seeds per model and publishing the results in August 2026. It is the largest open experiment to date on how frontier models perform at AI research, and the spread between models was wide.

Multiple seeds per model is the methodological choice that makes this useful. A single autonomous run tells you almost nothing, because these systems have high variance and one lucky trajectory can flatter a weak model. Running several seeds separates models that reliably make progress from models that occasionally stumble into it. Anthropic's Opus 5 closed 52 to 54 percent of the gap in the same experiment, well behind Fable 5's 82 percent, which is a much larger spread than the three-point differences these models show on aggregate intelligence indices.

My take: the gap between how models rank on static benchmarks and how they rank on multi-day autonomous work is the most useful finding here. Fable 5 and Opus 5 sit close together on index scores and far apart on this task, which suggests sustained autonomy is a distinct capability that current leaderboards do not measure. If you are choosing a model for long agentic runs, benchmark scores are a weak proxy. Our AI agent frameworks hub tracks the harness side of that problem.

3. Fable 5 Also Cracked an 87-Year-Old Math Conjecture

A mathematician at Anthropic used Claude Fable 5 to find a counterexample to the Jacobian conjecture, a problem that had resisted proof for decades. The result was reported in early August 2026 and involved the model surfacing a compact formula that overturns the conjecture.

A counterexample is a different kind of contribution from a proof, and in some ways a harder one to produce with search. Proving a statement means constructing an argument, while disproving it means finding a single object with exactly the right properties inside an enormous space. That is closer to what these models are good at, which is generating and filtering large numbers of structured candidates quickly. It sits alongside Anthropic's protein binder result from earlier this month, where Claude designed working binders for 14 of 15 targets at a 22 to 35 percent hit rate against a 10 to 15 percent industry norm.

My take: two genuine scientific contributions from one model family inside a month is a pattern rather than a coincidence, and both share a shape. In each case a human framed the problem and the model searched a space faster than a person could. That is the honest description of where these tools help in research right now, and it is a long way from autonomous science, but it is not nothing. See our August 19 roundup for the protein binder detail.

Cohort program

Claude MasteryCowork & Code

Explore programNo coding needed

4. Fable 5 Is Losing Where It Counts, to Opus 5

Claude Fable 5 has plateaued at 11 percent of Anthropic customer spending two months after launch, according to Financial Times reporting, while costing roughly twice as much per token as Claude Opus 5. Opus 5, priced at $5 per million input tokens and $25 output, has overtaken it in enterprise spend despite being the cheaper and nominally lower tier.

This is the clearest data point yet on how enterprises actually choose models, and it contradicts how labs market them. The top-of-range model wins the benchmarks and the headlines, and then most production traffic routes to the tier below it, because the capability difference does not justify the price difference on the work companies mostly do. Eleven percent of spend for a flagship two months in is a plateau, not a ramp. It also creates an awkward internal problem, since Fable 5 is the model that just closed 82 percent of the research gap and it is the one customers are not buying.

My take: this is the single most useful commercial signal in AI right now and I expect every lab to feel it. The market for absolute frontier capability is small and mostly research, while the market for good-enough-at-half-the-price is enormous. That is the same dynamic driving GPT-5.6 Sol's cut to $4 and $20 and DeepSeek serving 1.6 trillion parameters at $1.32. If you are picking a model, the flagship is probably not your answer.

5. Opus 5 Takes the Top Spot on Artificial Analysis

Claude Opus 5, released July 24, 2026, now ranks first on Artificial Analysis, topping both the Intelligence Index at 63 and the Agentic Index at 55.3. It is priced at $5 per million input tokens and $25 per million output. Anthropic positions it as its everyday model for enterprises, knowledge workers, and developers.

Leading both indices at once is the notable part, because intelligence and agentic performance often diverge. A model can score well on knowledge and reasoning questions and still fall apart across a long tool-using task, which is why the two boards exist separately. Opus 5 holding both while priced below Fable 5 explains the enterprise spend shift directly. For comparison, Kimi K3 sits at 57.11 with published weights at $3 and $15, and GLM-5.3 has been measured tying it around 60, so the open field is roughly three to six points back.

My take: index positions move week to week and this one has moved twice this month, so treat the ranking as a snapshot rather than a settled order. The durable observation is that the gap between the top closed model and the best open-weight model is now small enough that price and deployment constraints decide most real choices. Compare the full field in our best AI models ranking.

6. OX Alpha Fingerprinted as an Unreleased GLM-5.3

The anonymous model listed as stealth/ox-alpha on OpenRouter has been fingerprinted as an unreleased Zhipu GLM-5.3 variant. It carries a 1 million token context window, multimodal input, and reported capacity of 100 trillion tokens per day, and ran free through August 27. Reported scores put it at 80 percent on the DeepSWE coding benchmark against 65 percent for Fable 5 and 52 percent for GPT-5.6 Sol.

The identification rests on technical fingerprints rather than a disclosure: video encoder token consumption patterns identical to GLM-5V-Turbo, and tokenizer alignment with GLM-5.3. Both are hard to fake and hard to produce by coincidence. The 100 trillion tokens per day serving capacity is the detail that says this is a production system rather than a research artefact, since nobody provisions that much throughput for an experiment. The coding numbers still come from limited user testing rather than an audited leaderboard, so hold them loosely.

My take: a free frontier model with no brand attached gets tested harder and more honestly than any launch announcement, which is exactly why labs keep doing this. If the fingerprint is right, Zhipu has now run a public trial of its next flagship, published a leading cyber model, and holds a top-two open coding position, all inside three weeks. The free window closes on August 27, so test it now if you want your own numbers.

7. DeepSeek V4-Flash-Vision-Exp Beats Opus 4.8 on Two Benchmarks

DeepSeek released V4-Flash-Vision-Exp, a multimodal vision model with 284 billion total parameters activating 13 billion per prompt through mixture-of-experts routing. It surpassed the base V4-Flash on six of seven text benchmarks and beat Anthropic's Opus 4.8 on ALE, a suite of over 1,000 multi-step application tasks, and on ZeroBench, a set of 100 hard image analysis problems.

Beating a frontier model on ALE matters more than the image result, because ALE measures multi-step application work rather than single-turn perception. A vision-focused release outperforming a general frontier model on multi-step tasks suggests the vision training did not come at the usual cost to reasoning, which has been the standard trade in multimodal models. Activating 13 billion of 284 billion parameters keeps serving cost near that of a 13 billion parameter dense model, which is what makes an experimental vision model economically deployable at all.

My take: the Exp suffix is doing real work here and should temper expectations, since experimental releases from DeepSeek have previously changed substantially before general availability. Still, this is the second DeepSeek result this month worth attention after V4-Pro-0813 reached general availability at $1.32 and $3.96 with the strongest agent scores in its family. DeepSeek is shipping quietly and pricing aggressively while everyone watches the bigger names.

8. DeepSeek's Compression Cuts Million-Token Cost by 73 Percent

V4-Flash-Vision-Exp uses HCA and CSA compression techniques that reduce the cost of processing a 1 million token context by 73 percent. That is a change to the economics of long-context work rather than to capability.

Long context has been widely advertised and narrowly used, because the price of actually filling a million token window has been prohibitive for most workloads. Attention cost scales badly with sequence length, so a model advertising a million tokens is often quietly unaffordable past a few hundred thousand. Cutting that by 73 percent moves whole-codebase analysis, full-document-set review, and long agent histories from demo territory into budget territory. MiniMax is attacking the same problem from a different angle with sparse attention, claiming per-token compute at roughly a twentieth of its previous generation.

My take: efficiency work on long context is the most underrated line of research in the field right now, precisely because it does not produce a headline benchmark number. Two labs independently shipping large reductions in the same month suggests the constraint is finally being taken seriously. If you shelved a long-context feature on cost grounds in the first half of this year, the arithmetic has changed.

9. MiniMax M3 Pairs 80.5 Percent SWE-bench With Native Multimodality

MiniMax M3 is a 428 billion parameter mixture-of-experts model activating roughly 23 billion parameters, with a 1 million token context window and native text, image, and video input through MiniMax Sparse Attention. It scores above 92 on GPQA and 80.5 percent on SWE-bench Verified, and MiniMax reports more than 9 times faster prefill and more than 15 times faster decoding at 1 million context against its M2 generation, with per-token compute cut to roughly a twentieth. Weights are on Hugging Face under a custom minimax-community licence.

It is described as the first open-weight model to combine frontier agentic coding with native multimodality, and that combination is the point. Most open models pick one: strong code, or strong vision, rarely both in the same weights. A GPQA score above 92 puts it in serious scientific reasoning territory, though still behind GPT-5.4-Pro's 94.4 percent on GPQA Diamond. The custom community licence is the caveat worth checking before commercial deployment, since it is not Apache 2.0 and terms vary.

My take: 428 billion parameters with 23 billion active is a sensible sparsity ratio for people who actually intend to serve the thing, though this is still a cluster-scale model rather than something you run on a workstation. Read the licence properly before building on it. Alongside Qwen3.8-Max at 2.4 trillion and Kimi K3 at 2.8 trillion, the open field now has three credible frontier-adjacent options, which was not true in June.

LLM AGENTSRAG PIPELINESTOOL CALLINGDEPLOYMENT
Let's build

Start building AI agents with Build Fast

Explore Program

10. Qwen3.8-27B Reverse-Engineered a Binary Offline in 30 Minutes

Qwen3.8-27B, Alibaba's 27.8 billion parameter Apache 2.0 model, completed a full reverse-engineering task on a workstation with no network connection in roughly 30 minutes. Working from static analysis, it identified obscured cryptographic material inside a binary and self-corrected a bad key hash mid-task. It ran at about 50 tokens per second using SGLang with NVFP4 quantisation and DFlash2 speculative decoding.

Self-correcting a bad key hash is the detail that matters, because it means the model noticed its own output was wrong and revised, rather than confidently continuing down a broken path. That behaviour is what separates a usable analysis tool from an expensive guess generator. The stack is worth noting too: NVFP4 quantisation plus DFlash2 speculative decoding is what gets a 27.8 billion parameter model to 50 tokens per second on a single workstation, and DFlash2 claims 2.7x to 3.4x throughput over standard autoregressive decoding.

My take: fully offline security analysis at this quality changes who can do this work, and it cuts both directions. A small security team without a cloud budget gains a capable assistant, and so does everyone else. This is also the strongest practical argument yet for local open models: air-gapped reverse engineering is not a workload you can send to an API. Our AI coding tools hub tracks what runs where.

11. GPT-5.4-Pro Still Leads GPQA Diamond at 94.4 Percent

OpenAI's GPT-5.4-Pro holds the lead on GPQA Diamond, the graduate-level science reasoning benchmark, at 94.4 percent. MiniMax M3 posts above 92 on GPQA, placing the leading open-weight model within roughly two points of the closed leader on scientific reasoning.

GPQA Diamond is deliberately constructed so that non-experts with unlimited web access still fail most items, which is what makes a score in the mid-nineties meaningful. It is also close to saturation, and a benchmark above 94 percent has limited room left to distinguish between top models. Worth noting that the leader here is GPT-5.4-Pro rather than a 5.6 tier, which is a reminder that model version numbers do not map cleanly onto capability across every axis, and that specialised or higher-compute variants sometimes hold specific records.

My take: a two point gap between the best closed model and the best open-weight model on graduate science reasoning is the headline nobody wrote. A year ago that gap was wide enough to make the comparison uninteresting. The practical consequence is that for scientific and technical question answering, the choice is now about deployment and price rather than capability.

12. OpenAI's Price Cuts Now Run Across the Whole GPT-5.6 Line

OpenAI has now cut pricing across its GPT-5.6 tiers. Luna fell 80 percent to $0.20 per million input tokens and $1.20 output on July 30, 2026, Terra fell 20 percent to $2 and $12 on the same date, and Sol dropped to $4 and $20 from $5 and $30 in a three month promotional period announced August 21.

Taken together that is a repricing of the entire range inside four weeks, and the shape tells you where the pressure is. An 80 percent cut on the cheapest tier is a response to open models and to Gemini 3.7 Flash at $0.75 and $3.75, since that is the tier competing on volume. The Sol cut is aimed at Grok 4.6 at $2 and $6 and at DeepSeek V4-Pro at $1.32 and $3.96, both of which offer comparable index performance far cheaper. Only Sol's cut is explicitly promotional, which suggests OpenAI is testing whether the volume increase covers the margin loss before committing.

My take: the whole line moving in one direction inside a month is the clearest evidence that token pricing is now set by competition rather than by cost. For anyone with a shelved workload, this is the third repricing this quarter and the numbers are materially different from July. Rerun them. Our GPT-5.6 review covers the tier differences.

13. Gemini 3.7 Flash Nearly Doubled AutomationBench

Google's Gemini 3.7 Flash, released August 13, 2026 just 23 days after Gemini 3.6 Flash, lifted AutomationBench from 17.0 percent to 30.4 percent, DeepSWE v1.1 from 49.0 percent to 65.3 percent, and FrontierCode 1.1 Main from 34.4 percent to 43.6 percent. It ranks first of 186 models on output speed at 340.1 tokens per second, with a 1,048,576 token context window and a 64K output limit.

The AutomationBench jump is the one to sit with, because nearly doubling a score in 23 days on the same tier is unusual and it measures agentic automation rather than knowledge. Combined with the DeepSWE gain, the pattern says Google put this cycle's effort into tool use and multi-step execution rather than raw reasoning, which fits a workhorse tier that mostly serves agents. Introductory pricing of $0.75 and $3.75 runs only through December 31, 2026, after which it doubles to $1.50 and $7.50.

My take: a 23 day gap between workhorse releases is the fastest cadence anyone is running, and it confirms that the middle tier is where the competition actually is. The pricing expiry keeps getting overlooked and it will surprise teams building 2027 budgets. Model Flash spend at the post-January rate.

14. GLM-5.3 Open Weights Are Due This Week

Z.ai plans to publish GLM-5.3 open weights around August 28, 2026 after further safety testing. GLM-5.3 leads the CyberGym cybersecurity benchmark at 84.5 percent, ahead of Claude Mythos 5 and GPT-5.6 Sol, and lifted Terminal-Bench 3.0 from 4.6 percent to 28.3 percent, all from post-training on the same 743 billion parameter base as GLM-5.2.

Publishing weights for a model that leads a cybersecurity benchmark is a genuine dual-use decision, and the delay for safety testing suggests Z.ai knows it. A model at 84.5 percent on CyberGym is useful to defenders analysing their own systems and equally useful to attackers, and unlike an API there is no ability to revoke access after publication. The wider context is that no lab has established a norm here, and this release will become a reference point either way.

My take: this is the most consequential scheduled event of the week and it is getting far less coverage than it deserves. Whatever Z.ai does on August 28 sets an expectation for every subsequent open release of a security-capable model. Watch whether the published weights match the API version or arrive with capability differences, because a quietly reduced open release would be its own kind of precedent.

15. o3 Leaves ChatGPT on August 26

OpenAI retires o3 from ChatGPT on August 26, 2026, completing a 90 day sunset period. GPT-5.4 mini is rolling out to Free and Go users through the Thinking feature and serves as a rate-limit fallback for other tiers. Separately, Google shuts down gemini-robotics-er-1.6-preview on August 31.

Model retirements are the least glamorous and most disruptive events in this industry. Anything pinned to a specific model identifier breaks on the retirement date, and the replacement rarely behaves identically, so prompts tuned against the old model can degrade quietly rather than fail loudly. A 90 day sunset is reasonable notice by current standards, but reasonable notice only helps teams who tracked the announcement, and most did not.

My take: two retirements inside a week is a useful prompt to audit your own pins. Search your codebase for hardcoded model identifiers today, not on August 26. The deeper lesson is that model-agnostic architecture is not an ideological position, it is operational hygiene, because every model you depend on will eventually be switched off.

16. An API Flaw Let Weaker Models Decode Stronger Models' Reasoning

Researchers disclosed an API flaw affecting OpenAI, Anthropic, and Google that allowed weaker AI models to decode the reasoning traces of stronger ones. All three were notified, and the main extraction attack is no longer reproducible as of August 2026.

Reasoning traces are the intermediate steps a model generates before its final answer, and labs treat them as sensitive for two reasons. Commercially, the trace is a large part of what a frontier model's advantage consists of, and exposure enables distillation, where a competitor trains a cheaper model to imitate the expensive one. On safety, traces can reveal information the final output was designed to withhold. That the same class of flaw affected three independent providers suggests a shared design assumption rather than an implementation bug in any one system.

My take: the disclosure being handled properly and patched before publication is how this should work, and it deserves saying given how many AI security stories this month went the other way. The structural issue remains, which is that any API returning structured intermediate output leaks something about the model behind it. Expect labs to expose less reasoning detail over time, which is bad for debuggability and probably unavoidable.

Ready to upskill your team?

Tell us your stack and your goals. We build the programme around them.

Let's makeyour teamAI-native

Book a consultation

17. Where the Token Price War Stands Now

Pricing across the frontier has moved substantially inside four weeks, and the spread between comparable tiers is now large enough to change architecture decisions.

The spread from Luna at $1.20 output to Opus 5 at $25 is roughly twenty-fold, and the models are not twenty times apart on capability. That gap is where most cost optimisation lives: route the simple majority of requests to a cheap tier and escalate only what genuinely needs the expensive one.

18. Industry Moves in Brief

A few non-model items from the same cycle are worth logging.

●       OpenAI's public S-1 remains pending on SEC EDGAR after a confidential draft filed June 8, 2026, with the company at roughly $25 billion annualised revenue and a listing targeted between September 2026 and 2027.

●       Unitree Robotics continues trading after its August 19 Shanghai debut, which opened 629 percent above its 150.8 yuan IPO price, with a Reuters investigation subsequently tracing its quadruped designs to US Army Research Laboratory funded research.

●       Anthropic's Fable 5 adoption problem lands while the company prepares an IPO with Goldman Sachs, JPMorgan Chase, and Morgan Stanley, having reported more than $11.5 billion in second-quarter revenue and its first positive adjusted operating income.

●       Fireworks AI led August funding at $1.505 billion Series D, followed by Together AI at $800 million, with Fractile reaching a $6.5 billion valuation on a roughly $250 million Anthropic inference chip order.

Full industry detail sits in our weekend roundup and the August 21 roundup.

19. Where the Frontier Models Stand Today

Here is the practical state of the model landscape as of August 24, 2026.

The short version for teams choosing today: Qwen3.8-27B if it must run locally or offline, MiniMax M3 if you need open weights with vision and code, DeepSeek V4-Pro or Grok 4.6 if output token cost decides, Gemini 3.7 Flash if latency decides, and Opus 5 if capability decides. Full comparisons live in our best AI models ranking and the Kimi K3 review.

20. What to Watch Next in AI

Four things carry into this week.

●       GLM-5.3 open weights around August 28. Publishing a model that leads a cybersecurity benchmark sets a precedent every subsequent open release will be measured against.

●       The o3 retirement on August 26, and gemini-robotics-er-1.6-preview shutting down August 31. Audit your hardcoded model identifiers now.

●       OX Alpha's free window closing August 27, and whether Zhipu confirms the fingerprint or ships it under a name.

●       Whether GPT-5.6 Sol's $4 and $20 promotional pricing becomes permanent when the three month window closes in November.

The through-line for August 24 is that capability and commercial success have come apart. The model that closed 82 percent of an autonomous research gap is the one enterprises are not buying, while the tier below it leads both indices and takes the spend. That gap between what a model can do at its ceiling and what people actually pay it to do is the most interesting problem in AI right now, and it is a business problem rather than a technical one.

Frequently Asked Questions

Can AI models do AI research on their own?

Partly. In Prime Intellect's experiment of 153 autonomous runs across 18 frontier models, Claude Fable 5 topped the nanoGPT optimizer speedrun at 2,726 steps, closing 82 percent of the gap between the shared baseline and the human record and completing 99 percent of the research workflow without intervention. The human record still stands ahead, and the task had a clean metric and fast feedback, which most research does not.

Which AI model is number one right now?

Claude Opus 5 ranks first on Artificial Analysis as of late August 2026, topping the Intelligence Index at 63 and the Agentic Index at 55.3, priced at $5 per million input tokens and $25 output. GPT-5.4-Pro leads GPQA Diamond separately at 94.4 percent, so the answer depends on which capability you mean.

Who made OX Alpha?

OX Alpha, listed as stealth/ox-alpha on OpenRouter, has been fingerprinted as an unreleased Zhipu GLM-5.3 variant, based on video encoder token consumption patterns identical to GLM-5V-Turbo and tokenizer alignment with GLM-5.3. It has a 1 million token context window, reported capacity of 100 trillion tokens per day, and ran free through August 27, 2026. Zhipu has not formally confirmed it.

Is MiniMax M3 open source?

MiniMax M3 has open weights published on Hugging Face under a custom minimax-community licence, which is not a standard open-source licence such as Apache 2.0, so check the terms before commercial use. It is a 428 billion parameter mixture-of-experts model with roughly 23 billion active parameters, a 1 million token context window, and native text, image, and video input.

Which AI model is best at science questions?

GPT-5.4-Pro leads GPQA Diamond, the graduate-level science reasoning benchmark, at 94.4 percent. MiniMax M3 scores above 92 on GPQA, putting the leading open-weight model within roughly two points of the closed leader on scientific reasoning.

How much does GPT-5.6 Luna cost?

GPT-5.6 Luna costs $0.20 per million input tokens and $1.20 per million output, following an 80 percent price cut on July 30, 2026. GPT-5.6 Terra fell 20 percent the same day to $2 and $12, and GPT-5.6 Sol dropped to $4 and $20 in a three month promotion announced August 21.

Why is Claude Fable 5 struggling with enterprises?

Financial Times reporting puts Fable 5 at 11 percent of Anthropic customer spending two months after launch, plateaued, while costing roughly twice as much per token as Claude Opus 5. Opus 5 has overtaken it on enterprise spend, suggesting the capability difference does not justify the price difference for most production work.

When does o3 get retired from ChatGPT?

OpenAI retires o3 from ChatGPT on August 26, 2026 after a 90 day sunset period. GPT-5.4 mini is rolling out to Free and Go users through the Thinking feature and acts as a rate-limit fallback for other tiers. Google separately shuts down gemini-robotics-er-1.6-preview on August 31.

What is DeepSeek V4-Flash-Vision-Exp?

DeepSeek V4-Flash-Vision-Exp is an experimental multimodal vision model with 284 billion total parameters activating 13 billion per prompt. It beat the base V4-Flash on six of seven text benchmarks and outperformed Anthropic's Opus 4.8 on ALE, covering over 1,000 multi-step application tasks, and ZeroBench, 100 hard image analysis tasks. Its HCA and CSA compression cuts 1 million token processing cost by 73 percent.

Can an open model reverse-engineer software?

Yes. Qwen3.8-27B, a 27.8 billion parameter Apache 2.0 model, completed a reverse-engineering task offline on a workstation in roughly 30 minutes, identifying obscured cryptographic material in a binary through static analysis and self-correcting a bad key hash. It ran at about 50 tokens per second using SGLang with NVFP4 quantisation and DFlash2 speculative decoding.

Recommended Blogs

●       Mystery Model OX Alpha Beats GPT-5.6: AI News August 22-23 2026

●       GLM-5.3 Beats Claude and GPT-5.6 on Cyber: AI News August 21 2026

●       Unitree's Robot IPO Soars 629%: AI News August 20 2026

●       Anthropic Raises Its Own AI Risk Level: AI News August 19 2026

●       Best AI Models July 2026: Ranked by Use Case and Price

●       GPT-5.6 Review: Sol, Terra, Luna Benchmarks and Pricing

●       Kimi K3 Review: Benchmarks, Pricing, and K2 Comparison

Resources & Community

Join our community of 70,000+ AI enthusiasts and learn to build powerful AI applications! Whether you're a beginner or an experienced developer, Build Fast with AI helps you understand and implement AI in your projects.

●       Website: buildfastwithai.com

●       LinkedIn: Build Fast with AI

●       Instagram: @buildfastwithai

●       Founder Twitter: @satvikps

●       Twitter: @BuildFastWithAI

Agentic AI Launchpad 2026

A structured 6-week cohort program that takes you from AI basics to building and deploying real-world agentic AI systems. Includes live sessions, expert mentorship, project reviews, and a builder community network.

Ready to go from learning to building? Join the next cohort: Agentic AI Launchpad 2026

Free AI Resources

Access free tools, workshops, and micro-learning to keep building:

●       AI Workshops: Free resources, upcoming events, and past recordings

●       Unrot: Learn AI in 5 minutes a day (free micro-learning app)

●       Gen AI Experiments: free cookbooks and notebooks on GitHub

GLM-5.3's open weights and the o3 retirement both land this week. Follow Build Fast with AI so each recap reaches you before your standup.

References

●       Autonomous nanoGPT speedrun results (Prime Intellect)

●       nanoGPT speedrun leaderboard (Prime Intellect)

●       Fable 5 enterprise adoption reporting (The News)

●       Fable 5 and the Jacobian conjecture (ScienceDaily)

●       MiniMax M3 architecture and benchmarks (Morph)

●       MiniMax Sparse Attention paper (arXiv)

●       API reasoning extraction flaw (The Hacker News)

●       August model release tracker (Local AI Zone)

●       Model release timeline (LLM Gateway)

●       Independent model evaluations (Artificial Analysis)

●       LLM benchmark leaderboard (BenchLM)

●       Daily AI news roundups (Build Fast with AI)

Enjoyed this article? Share it →
Share:
    You Might Also Like
    Ox Alpha Review: The Mystery AI Model With 1M Context (2026)
    LLMs
    Ox Alpha Review: The Mystery AI Model With 1M Context (2026)

    Ox Alpha is a free stealth AI model on OpenRouter with 1M context, multimodal input and strong early coding results. Here is what is verified, what is speculation and whether it is worth using.

    Mystery Model OX Alpha Beats GPT-5.6: AI News Aug 22-23
    Analysis
    Mystery Model OX Alpha Beats GPT-5.6: AI News Aug 22-23

    Weekend edition: OX Alpha topped coding tests at 80 percent DeepSWE, OpenAI cut GPT-5.6 Sol to $4/$20, and Qwen3.8-Max shipped 2.4 trillion open weights. 21 stories.