AA-AnalystAgent Benchmark: Claude Opus 5 Leads at 54%
Artificial Analysis tests AI on real-world spreadsheets and documents. Claude Opus 5 scored 54%; 57% of failures anchored on a wrong early hypothesis.
Hyunkook Kim
Artificial Analysis tests AI on real-world spreadsheets and documents. Claude Opus 5 scored 54%; 57% of failures anchored on a wrong early hypothesis.
Deal with six firms including Goldman Sachs, BlackRock to secure AI infrastructure funding; Bernstein and others warn of depreciation risk
Adobe's interior editing tutorial, introduced via YouTube Shorts, demonstrates generative AI use cases
CEO Sundar Pichai announced on X that Gemini became the 14th Google product to reach the milestone
Lightcap, a longtime associate of Sam Altman, leaves OpenAI as it prepares for an IPO to start something new
OpenAI has released a preview of its ChatGPT and Codex desktop apps for Ubuntu, Debian, and Fedora
OpenAI adds an import feature to its desktop app that pulls projects, conversations, skills, and plugins into one place
Available to Grok Heavy, Cursor Ultra and Team Premium users, bots run real tasks on their own computers
AMIE, previously limited to text conversation, expands to real-time video consultations, demonstrating expert-level performance in simulated consultations
Posted as a short clip with no specific feature details, read as a move amid competition in AI avatars
Higgsfield ran the same 15-second script through seven platforms including HeyGen and Synthesia side by side
The MoE architecture activates only 3B of its total 30B parameters per token, and NVIDIA claims up to 4x faster output than comparable models