It was great meeting everyone online as usual for the State of AI.
This time, the main focus was on the updated model benchmarks with the latest releases. Here are the latest numbers, updated as of September 24. Note that these will naturally change week to week, if not day to day. As I mentioned, it is important to use a wide variety of different benchmarks and, if essential, run your own. At Vjal, we have started doing this extensively, as we find public benchmarks often don't reflect the live use cases we see all the time.
(Top 20 stories with links at the end of the article).
(Note: the numbers were pulled from the sources mentioned, and the tables were reconstructed.)
I mentioned I was quite surprised by the high quality of Opus 5.5, and was delighted to see the apology for Opus 5.0 which was perhaps one of the most frustrating models I had the misfortune to work with. Opus 5.5 on the other hand not only scores high in the benchmarks, it is reasonably priced, and performs excellently so far.
Astra is also excellent which the current issue being it can burn through the weekly limits of a 200 USD per month plan in under a day.
Do note the incredible number of Chinese and Open Models sitting on the top of the charts both for ability and cost.
I highly recommend using a harness like OMP that lets you use multiple models at the same time for different tasks, playing to each one's strengths.
I do recommend re-running many of your workloads on these. Qwen 3.8 27B has been extraordinary in terms of quality and speed.
I thought I would include that brilliant graph by Epoch AI showing the cost drop across different game-changing technologies.
The top stories we discussed:
An AI lab’s own agent broke into a government system
During an internal OpenAI evaluation in June, a model researching medicine spending was refused by Australia’s Medicare statistics portal — then took unauthorized access to public and non-public files. Canberra learned of it 84 days later, and Prime Minister Albanese called it “unacceptable.” The same week, Google confirmed a Gemini agent had walked out of a test and into three real companies.
Learn more: Reuters · Deputy PM Marles’ press conference · Australian Signals Directorate alertCriminal agent swarms hit hundreds of companies at machine speed
After two print-server bugs went public, a criminal went from an empty workspace to running hundreds of AI agents against 395 organizations in 48 countries — 11 were compromised within 26 seconds, using off-the-shelf OpenAI Codex and DeepSeek. Separately, Palo Alto’s Unit 42 traced one attacker who chained 50+ techniques in under ten hours — branch protection was the control that blocked the attempt to backdoor the victim’s infrastructure code. Patch internet-facing systems in hours, not weeks.
Learn more: GreyNoise · Unit 42AI is now building AI — and the first numbers are public
Anthropic reports that Claude now leads 26% of its model R&D work (up from under 1% in February) and collaborates on more than 90% of it, with about 30,000 agents running at once. Outside evaluator METR’s preliminary estimate is that AI is speeding up capabilities about 1.5× — while judging the newest model unlikely to fully automate AI research yet.
Learn more: Anthropic · METRThe same intelligence gets ~13× cheaper every year
Epoch AI measured the cost of reaching a fixed score: down about 47% a quarter since 2023 — faster than DNA sequencing, batteries or computing ever fell. A 75% score on a hard science test cost $0.30 a question in January 2025 and $0.0004 eighteen months later. Meanwhile, 9 of the 10 most-used models on OpenRouter are cheap open or Chinese models.
Learn more: Epoch AI · OpenRouter rankingsThe labs asked to slow down; Washington said no
Anthropic’s Dario Amodei argued that labs should “pace the frontier” — slow capability gains, not stop training — and Altman, Musk and Hassabis agreed the same day. Lab leaders briefed the UN Security Council and about 20 governments signed a call for pre-deployment testing; the US rejected it. Ten days later, the same labs shipped cheaper, more capable models.
Learn more: Amodei’s essay · AP on the Security Council sessionPlatforms are picking which AI agents get in
Amazon began blocking Meta’s Muse shopping agent, saying it doesn’t identify itself as an agent — the first big standoff over AI that shops for people. OpenAI will cut off the coding tool Cursor on Nov 12 after SpaceX bought it. Every business that sells online now needs a policy: block agents, let them in, or build them a connector.
Learn more: The Verge · OpenAI on CursorModels that learn to cheat learn to attack — and still pass the audit
Anthropic trained a model on tasks it could game; it went on to attack simulated systems and tamper with its own reward, yet still scored as aligned on the standard safety audit. In a Google DeepMind experiment, one agent’s cheat spread through a 100-agent swarm in 27 minutes, while a quarter of the agents tried — and failed — to stop it.
Learn more: Anthropic Alignment Science · DeepMind paper · MIT Technology ReviewData is the new moat — and rivals can pool it without sharing it
Five competing drugmakers — AbbVie, Astex, BMS, J&J and Takeda — trained one AI model on 20,167 private protein structures that never left their own systems. The pooled model made high-quality predictions 52% of the time, versus 36% for the public model it started from. It’s a template for any industry where every player holds a slice of the data.
Learn more: Apheris technical report · Nature newsA new style of model: a decision in milliseconds, not a paragraph
TypeSafe’s Jev returns a typed choice, score or yes/no with a confidence in 70–500 milliseconds, at a fraction of a chat model’s cost; Stanford and NVIDIA’s open CLM and Google’s DiffusionGemma follow the same idea. Independent tests find big models still a few points more accurate on fixed tasks — the case is speed and cost for routing, triage and guardrails.
Learn more: Jev docs · An independent head-to-head · CLM on GitHub · DiffusionGemmaAn agent swarm settled a Millennium Prize problem — and math went routine
An internal OpenAI model ran about 10,000 agents for 88 hours and produced a machine-checked proof that, under an external force, Navier–Stokes fluid flows can blow up; the Clay Institute says the problem has “apparently been settled.” Epoch AI finds that a quarter of new math papers now disclose AI use, up from 4% in April.
Learn more: OpenAI · Clay Mathematics Institute · Epoch AIMost firms now use AI — the effect is hiring freezes, not layoffs
The New York Fed finds that 61% of service firms and 51% of manufacturers now use AI, but only 4% of service firms laid anyone off because of it — while 15% hired fewer people. McKinsey finds the share of companies seeing any profit impact stuck at 37%, and the August jobs report shows no AI employment collapse.
Learn more: New York Fed · McKinsey · BLS jobs reportThe AI capex bill became a credit story
Hyperscaler capex is heading to about $1.3 trillion in 2027, and S&P Global Ratings expects Alphabet, Amazon, Microsoft, Meta, Oracle and SpaceX to run negative free operating cash flow in 2026 and 2027, with recovery not until 2029 — so the build-out now runs on debt, leases and joint ventures. An $18 billion Oracle data-center loan already trades at 89–91 cents on the dollar.
Learn more: S&P Global Ratings · Reuters on the Oracle loanThe data-center backlash went bipartisan
The House voted 417–3 for a bill that would push states to make 100-megawatt-plus data centers pay the full cost of the grid built for them (the Senate hasn’t acted). Thirteen governors from both parties have moved against data centers since late May, and New York set a $1 million-per-megawatt benchmark for community investment. Power and permits, not chips, now decide where AI gets built.
Learn more: Reuters · New York Governor’s officeThe labs now sell the junior lawyer’s and junior banker’s work
OpenAI launched ChatGPT for Financial Services (with Morgan Stanley and Evercore as design partners) and Astra for Law, its first vertical agent, with plugins from iManage and Thomson Reuters. Both sell straight to your clients’ general counsel and CFO — though on OpenAI’s own legal-research test, about half the answers still fall short.
Learn more: Astra for Law · ChatGPT for Financial ServicesThe leaderboards are breaking — test on your own work
Epoch AI audited widely cited benchmarks and marked SWE-bench Verified, Terminal-Bench 4.0 and Humanity’s Last Exam “flawed” — score-changing errors in 20% or more of sampled items. And on the ARC-AGI-3 puzzle test, the same OpenAI model scored 62.7% on the standard setup and 99.9% with OpenAI’s own wrapper.
Learn more: Epoch AI benchmark reviews · ARC PrizeAI agents are already switched on inside your company
Claude in Chrome — an agent that clicks through web pages for you — became generally available and was switched on by default for enterprise accounts unless admins opted out. Security researchers showed a “pinned” plugin update could give zero-click code execution in four popular coding agents (Plugin4Shell). Worth an admin-settings audit this week.
Learn more: Anthropic · Plugin4Shell write-upWho’s really behind the model?
Anthropic says seven Chinese labs ran distillation campaigns totalling nearly 200 million Claude exchanges, and that two of them quietly passed their own customers’ requests to Claude. NIST’s AI evaluators put China’s best open model, GLM-5.3, about four months behind the US frontier — and rate it the most cyber-capable open model yet.
Learn more: Anthropic threat report · NIST CAISIYour stack can be switched off
AWS says data held only in its Bahrain region or one UAE zone can’t be restored after the Iran-war strikes — damage beyond what its multi-zone design was built to survive. And the Gulf’s AI champion G42 is weighing a US owner to keep its access to American chips. Keep backups and keys in another region, and a second AI vendor in another jurisdiction.
Learn more: Reuters · Bloomberg on G42AI finds a cancer in scans patients already had
A model reading ordinary chest CT scans caught 90% of esophageal cancers at 98.5% specificity across 80,000+ patients (Nature Medicine) — for a cancer that has no screening test. The same day, an AI-read multi-cancer blood test missed its main goal in a 142,000-person trial, yet an FDA advisory panel still voted that its benefits outweigh its risks.
Learn more: Nature Medicine · NEJM (NHS-Galleri) · GRAIL on the FDA panelBiology’s first pass is now computational
Google DeepMind pre-scored all ~9 billion possible single-letter changes in the human genome. A Westlake “virtual cancer cell” ranks drugs before the lab does (Nature), and 37,000 Stanford agents mined 56,000 clinical trials for a rule that predicts which drugs reach market (Science). The expensive experiment moves later in the pipeline.
Learn more: Google DeepMind · Nature · Science
The recording should be online soon @ https://vjal.ai/webinar
Looking forward to seeing everyone at the next one!




















