To date, AI industry spending has topped $1.6 trillion, and shows no sign of slowing anytime soon.
So what do we actually have to show for it? Historically, it’s been a whole lot of nothing: as numerous studies have shown us, tools like AI chatbots and autonomous agents have been ineffective at completing real world tasks in a competent way.
The tech industry insists that’s all about to change within the next few years, as AI’s capabilities grow by leaps and bounds, enabling economic growth the likes of which the world has never seen. But is it really?
Not necessarily. A new study out of the University of California Berkeley’s Center for Responsible, Decentralized Intelligence — flagged by the College Fix — shows that frontier AI tools of all makes and models are still incapable of completing the vast majority of workplace tasks at an acceptable level, throwing a major wrench in the tech industry’s assertions that the AI revolution is imminent.
To come to that conclusion, the UC researchers designed a rigorous assessment they call the “Agents’ Last Exam,” developed to test “job-readiness” across numerous state-of-the-art AI models. Basically, the ALE — an impish riff on “Humanity’s Last Exam” — is designed to put an AI system through its paces, covering “more than 1,500 expert-sourced tasks spanning 55 occupations,” the researchers wrote in apress release.
Those test spans the typical line-up of AI-exposed jobs like software engineering and graphic design, but also a substantial number of jobs whose fates remain less certain, such as maritime engineering, agriculture, audio production, and public health operations.
Using the ALE benchmark, researchers took a hard look at advanced “closed” models — proprietary AI systems developed by private companies — like Anthropic’s Fable 5, OpenAI’s GPT-5.5, Cursor’s Composer 2.5, and Google’s Gemini 3.1 Pro. (For good measure, they also looked at two open-source models by Chinese developers.)
As cutting-edge as these AI models are, the research found that they’re far from ready for the complex needs of the modern workplace. Out of all of the models put through the gauntlet, each of them failed spectacularly. OpenAI’s GPT-5.5 came in with the highest score: a passing rate of just 24 percent overall.
“Today’s agents can solve a meaningful fraction of professional tasks,” the researchers wrote. “However, when we look at the hardest tasks that require sustained reasoning, deep domain expertise, and reliable execution over long horizons, they are still far from human-level performance.”
And as tasks became more complicated, even those meager aggregate scores fell off fast.
“On ALE’s hardest tier, every frontier agent we tested, including Fable 5,…
Source link
Disclaimer
We strive to uphold the highest ethical standards in all of our reporting and coverage. We blogs.grocliq.com want to be transparent with our readers about any potential conflicts of interest that may arise in our work. It’s possible that some of the investors we feature may have connections to other businesses, including competitors or companies we write about. However, we want to assure our readers that this will not have any impact on the integrity or impartiality of our reporting. We are committed to delivering accurate, unbiased news and information to our audience, and we will continue to uphold our ethics and principles in all of our work. Thank you for your trust and support.
Website Upgradation is going on for any glitch kindly connect at [email protected]