SWE-bench Verified moved from 60% to nearly 100% in a single year. 

Wow...

A benchmark built to separate frontier labs from everyone else stopped doing that job in roughly the time it takes most enterprises to renew a cloud contract.

Capability is the part of AI that keeps solving itself.

💡
The eight stats below, pulled from primary research rather than conference keynotes, tell a consistent story: every measure of how good these systems have become is outpacing every measure of how well anyone can trust them.
6 reasons AI engineers can jump into robotics right now
Somewhere between 2,000 and a few thousand engineers in the US can genuinely combine vision-language-action models, sensor fusion, and kinematics. Against that tiny bench, the market is posting more than 65,000 open robotics roles, according to a widely cited analysis from Fruition Group.

Stat 1: The benchmark that stopped meaning much

That SWE-bench jump comes from Stanford HAI's AI Index Report, alongside a blunter finding: industry produced 91% of notable AI models in the tracked period, up sharply from prior years. Capability gains have become the easy part of this story.

A near-perfect SWE-bench score leaves open whether that same model can sustain a coherent multi-turn conversation, how often it still gets things wrong once deployed, and whether anyone can actually demonstrate that reliability beyond the benchmark itself.

The report's own framing lands harder than most vendor decks manage: a field scaling faster than the systems around it can adapt. Somewhere, a marketing team is already turning that sentence into a keynote slide, missing the point on the way to the font choice.


Stat 2: The number that should worry a CISO more than the benchmark score

Veracode's GenAI Code Security Report tested more than 100 models and found the average security pass rate sitting at 56%, essentially flat against the prior year. Java code failed security tests over 70% of the time.

Models write code that compiles at close to a perfect rate. Writing code that resists an attacker turned out to be a separate skill entirely, one that matters most once that code reaches mission-critical systems. The compiler has zero opinion on your firewall.

The same disconnect shows up in agent evaluations, where models keep failing simple selection traps even as headline scores climb.


Stat 3: The transparency score moving in the opposite direction from capability

Stanford's Foundation Model Transparency Index fell from 58 to 40 points this year, and 84% of the most notable recent models shipped with the training code omitted. Google, Anthropic, and OpenAI have all stopped disclosing dataset sizes and training durations for their newest releases.

The labs topping the leaderboards are the same labs publishing the least about how they got there.


Stat 4: The money that says nobody is waiting for the safety data

Two figures put a number on how much confidence the market is placing in capability alone, ahead of the trust question catching up.

  • Global corporate AI investment hit $581.7 billion, up 130% year over year, with generative AI investment alone reaching $170.9 billion, according to Stanford HAI.
  • The top four US hyperscalers pushed combined data center capex toward $600 billion, including Meta, on a path toward $1.7 trillion globally by 2030, per Dell'Oro Group.
💡
Compare either figure to the interstate highway system, and the AI buildout still wins on speed. Eisenhower pulled his off with plain asphalt, long before anyone needed a nuclear power deal just to keep the servers cool.

Increasingly, that spending follows a "train once, infer forever" logic that is reshaping how budgets get allocated.


Stat 5: The language shift that happened faster than most roadmaps updated

GitHub's Octoverse report found TypeScript overtook both Python and JavaScript in August 2025 to become the platform's most used language, the biggest language shift GitHub has recorded in over a decade. Somewhere, a Python maintainer is recalculating a career plan mid-standup.

💡
80% of new GitHub users started using Copilot within their first week on the platform, faster than most of them picked a text editor theme.

A five-year language roadmap became a two-year one, the kind of shift AI architects now plan around faster than most individual engineer checklists get updated.

Bridging the gap from supercomputing to AI factories
A comprehensive industry report on modernizing high-performance computing for production AI, featuring insights from NVIDIA and WEKA leaders.

Stat 6: The skill cluster that outran the org chart

Lightcast's contribution to the Stanford AI Index tracked mentions of the "agentic AI" skill cluster in US job postings and found them growing over 280% in a single year, moving from 0.06% of postings to 0.23% (roughly 90,000 postings).

💡
Overall AI skill mentions climbed to 2.5% of all US postings, up 55% year over year.

Hiring managers are asking for a skill set that barely had a name eighteen months ago, while a lot of resumes still list "proficient in Excel" as the differentiator right above the part where they claim they can wrangle an agent swarm.


Stat 7: The workforce number sitting in survey data before it hits headcount

The same Stanford Index found a third of surveyed organizations expect AI to shrink their workforce in the coming year, concentrated in service operations, supply chain, and software engineering.

Large-scale job losses have yet to appear in aggregate employment data, a gap the report treats as a timing issue rather than a reassurance. That distinction matters more than either half of the sentence read alone, and it is worth revisiting at the next earnings season rather than this one.


Stat 8: The adoption line that crossed mainstream before governance caught up

Organizational AI adoption reached 88%, up from 78%, according to the same Stanford report. Almost every organization worth surveying has adopted something, which in plenty of cases means a pilot license nobody has opened since the kickoff meeting.

Adopting something differs meaningfully from capturing enterprise value from it, and differs further still from building the agent experience layer most of these systems still need. That gap is exactly what shows up in the growing list of agentic deployment mistakes enterprises keep repeating.

💡
What the stats above make clear is that adoption moved fastest of all three curves: capability, trust, and readiness. Only one of those three actually caught up.

So, what should builders and buyers do with the gap?

  • Treat benchmark scores as a single input among several. SWE-bench near 100% and a security pass rate stuck at 56% describe the same model generation. Procurement built on capability scores alone is reading half the page, a lesson every applied machine learning product manager learns early, usually right after a vendor demo that skipped security entirely.
  • Ask vendors what they stopped disclosing, alongside what they shipped. A transparency index falling from 58 to 40 points is a sourcing question every technical buyer should be asking out loud.
  • Budget for the security review the Veracode number implies. A 56% pass rate means roughly half of committed AI-generated code needs the scrutiny a team gives human-written code, at minimum.
  • Hire for the skill cluster before the job posting catches up. A 280% jump in agentic AI mentions is a leading indicator worth acting on before it turns into a talent shortage, especially while the agents themselves keep tripping over basic "why" questions. Interview accordingly.

Capability was always going to move first. The genuinely hard problem, the one these numbers leave open, is closing the distance between how good these systems have become and how much anyone actually knows about them.

That problem lands on engineers first, but it lands just as hard on the chief AI officers steering strategy from Silicon Valley and the tech leaders doing the same from New York. Progress got its keynote. Trust is still waiting for its turn on stage.


Where the infrastructure money actually has to land

The capex figures above describe the top of the funnel. What happens between a hyperscaler writing a check and a production AI system running is a modernization problem most infrastructure teams are still solving in real time.

AIAI's report, Bridging the gap from supercomputing to AI factories, pulls direct insight from NVIDIA and WEKA leadership on that exact gap: where HPC design assumptions break under production AI workloads, what modernization looks like once the check clears, and how teams keep existing infrastructure investment intact while making the shift.

Download the report before the next infrastructure budget gets approved on assumptions worth checking first.