Even the best AI models can't score above 60% on pure vision
Moonshot AI's new benchmark ranks GPT-5.6 Sol first at 59.7%, revealing that many "reasoning errors" are actually perception failures
One email each morning — yesterday's AI, sortedGet it in your inbox
Tag
Moonshot AI's new benchmark ranks GPT-5.6 Sol first at 59.7%, revealing that many "reasoning errors" are actually perception failures
Tesla's 2019 vision-only bet has migrated to home robotics — Karpathy's argument, and lidar prices that have fallen to $200
An update combining voice and gesture matches arm direction to a 3D map to clean a specific spot
Offline-capable Gemini and insulin resistance tracking added, price now $399
Instead of typing, deaf users can sign, and Gboard and Live Transcribe turn it into text
Ultra-compact VLM specialized in document and chart understanding, released free on Hugging Face under Apache 2.0
LFM2.5-VL-3B hits 20 tokens/sec on Galaxy S26 Ultra, 228 tokens/sec on M5 Max
An open-weight world model that produces a 10-second clip in 6.8 seconds, optimized for NVIDIA GPUs
Beyond generating reports, a method that measures heart and thoracic width to compute CTR
Multimodal model handling images, video, and text released open-source on Hugging Face
That's the last story.