Gemini 3.7 Flash delivers performance comparable to Fable 5 at a fraction of the cost
We see that many models finish the tasks prematurely, resulting in lower scores and costs
@GoogleDeepMind and @Alibaba_Qwen released new models this week. These are the results:
Gemini 3.7 Flash is the new State-of-the-art model on VGI-Bench.
Qwen 3.8 Max comes in as fourth place in after Gemini 3.6 Flash.
Introducing VGI-Bench: a multimodal, holistic benchmark probing 12 distinct visual and audio-visual skills.
550 human-curated questions, designed to mitigate the common mistakes in today's video benchmarks and expose pragmatic failures of state-of-the-art models.
Best model: 64.73%. Humans: 84.5%.
73K Followers 3K FollowingWe're in a race. It's not USA vs China but humans and AGIs vs ape power centralization.
@deepseek_ai stan #1, 2023–Deep Time
«C’est la guerre.» ®1
12K Followers 804 FollowingI make youtube videos on cool AI research /// AI papers newsletter https://t.co/Xn7GMDbQSd /// paper recap @TheAITimeline /// https://t.co/yigZMs32sO
19K Followers 9K FollowingI push the AI frontier by building tough benchmarks with amazing people. SWE-bench, SWE-agent, SciCode, AlgoTune. Postdoc @Princeton. PhD @nlpnoah @UW.
17K Followers 1K FollowingTraining VLMs at @GoogleDeepMind - ex @Amazon @CarnegieMellon @PoliTOnews @IITalk - I post honest non-AI-generated paper reviews
1.6M Followers 2 FollowingWe're an AI safety and research company that builds reliable, interpretable, and steerable AI systems. Talk to our AI assistant @claudeai on https://t.co/FhDI3KQh0n.