Holy moly. 🤯
Qwen 3.8 27B just scored higher than:
GPT 5.6 Terra
GLM 5.2
DeepSeek V4 Pro
Muse Spark 1.2
Claude Opus 4.8
on the Artificial Analysis Agentic Index.
And you can run this on a single RTX 3090/4090.
Go show your GPU some respect. 🫡
So Anthropic scraped the entire internet to train Claude, yet somehow completely omitted the parts where the CEO's wife ran a porn startup, married a man over 40 years older and pitched her app to Jeffrey Epstein. Peak AI alignment right there.
x.com/amir/status/20…
DeepSeek V4 Pro 0813 scores 53 on the Artificial Analysis Intelligence Index, 8 points above April's DeepSeek V4 Pro - but with a 3.6x price increase and only 1 point above DeepSeek V4 Flash 0731
@deepseek_ai has released DeepSeek V4 Pro 0813, its new flagship model, along with
Alongside the release of GLM-5.3, we are launching the OpenVuln project:
huggingface.co/spaces/zai-org…
If you own or maintain an open-source project and would like additional security support, you can submit its public GitHub repository. OpenVuln will scan the repository for potential
We’ve written an FAQ to answer some of the questions we've received about watermarking.
In summary:
• We’re implementing watermarking to comply with the EU AI Act. Other major model developers have signed the same Code of Practice and will also be implementing watermarking;
• Our watermarking method doesn’t have any practical impact on the quality or content of Claude’s outputs;
• The difference between watermarked and un-watermarked text will not be distinguishable to readers;
• Nothing is added to the text and there are no hidden characters;
• Watermarking doesn’t require extra tokens, and will not be more expensive;
• Watermarks can’t be traced to a specific person, organization, or chat.
Read more: anthropic.com/news/claude-te…
Absolutely insane. This might be the clearest glimpse yet of how AI will transform scientific discovery.
Anthropic asked an unreleased version of Claude to take a real stab at the Riemann Hypothesis, one of the most famous unsolved problems in mathematics.
It failed. But while
🚨 Composer 3 Leaks: It will drop
> It outperform Opus 5 and GPT-5.6 Sol on coding and agentic tasks
> Six different internal variants have reportedly appeared
> Reasoning levels range from Fast → Medium → High → XHigh
> It Will be around 5x cheaper than opus 5 and GPT-5.6
i think future will belong to maintainers cause anyone can make a piece of software today but maintaining it over a period of time suppos 10 year 15 years is nightmare ........ or just simply a agent loop will maintain it ??
🤯Introducing Team Memory, same idea as Agent Memory, except your teammates' agents can read it too
2.0.0 beta out today, and the repo hit #1 on github's typescript trending this week
Highlights:
> Solo builders: one place to manage memory across all your agents and AI tools,
I'm going to cancel Claude. It's just so bad, I can't believe it.
It's just lazy. The most recent example: I have Claude check my inbox for important emails, summarize them, work with them, and send out replies if necessary. I caught Claude again simply not reading the email thread to the end and just ignoring the latest emails.
When I asked him about it, Opus 5 just said: "Valid point. I didn't read it."
I mean, seriously. What the heck? You have to babysit it every time.
DeepSeek V4 Flash 0731 is now open weights!
@deepseek_ai has just released the weights for its new flash tier model, DeepSeek V4 Flash 0731. With a score of 50 on the Artificial Analysis Intelligence Index, it lands among the top 3 open weights models on the leaderboard. The
The tiny Kimi-K3 that you can runs locally on a potato hardware.
- 2.8T to 0.18B.
- 0.10B activated.
- Same architecture.
- Same new attention design scale.
- Same DNA, smaller version
- compressed it into a 0.18B version for testing.
- now fits in 700MB.
You can actually
Open Source AI updates :
> GLM-5.2 ✅ Almost Opus-level
> Kimi-K3 ✅ Almost Fable-level
> Qwen-3.8 🔜 2.4T, Expected to beat Opus
> Deepseek V4 GA 🔜 $0.0028/M Expected to beat Opus
> Minimax-M3-Pro 🔜 3T, Expected to be Fable-level
> GLM-5.5 🔜 Expected to beat Opus
Holy
Which plan we should go with??
1. Kimi k3 is here all plans sold out
2. Qwen 3 8 max is here launching soon.... One of the best and usage heavy literally 18$ feels like $100 of usage according to. Some users
3. Deepseek v4 GA is launching soon.....
4. Or. Claude code 😭😭
50 Followers 132 FollowingMarkets, charts, bad sleep, and occasional bad decisions. Day trading equities & futures. Risk manager by day, risk taker by night. Chasing edges, fighting FOMO
520 Followers 515 FollowingTinterillo legal, politólogo de poca calle y comentarista en whatsup con (armadura de combate + motor de búsqueda + inteligencia artificial) - Ver. 58 Gen X
237 Followers 296 FollowingGeek and coder. Ruby dev and web curious. Tweeting in English and Portuguese. Him/Ele. Dev at @conversationEDU
-
Find me at Mastodon @[email protected]
437 Followers 904 FollowingTechno-adventurer, citizen of the world imprisoned voluntarily in Greece, Web-sucker-surfer, social media observer, music lover.
923 Followers 7K FollowingI'm an innovation bod working on an indie project, exploring possibilities in art/sci and biz/tech. R/L/F≠E. @[email protected] https://t.co/LoTgqWcc1q
70 Followers 1K FollowingA concentração d riqueza produz concentração d poder político. E a concentração do poder político dá origem a 1 legislação q aumenta e acelera o ciclo.
Chomsky
1.3M Followers 790 FollowingFounder/Chair, AMI Labs; Professor, NYU; Partner, 224 Ventures; Ex-Chief AI Scientist, Meta.
Researcher in AI, ML, Robotics, etc.
ACM Turing Award Laureate.
4.0M Followers 2K FollowingFounder and CEO, O'Reilly Media. Watching the alpha geeks, sharing their stories, helping the future unfold. Didn't pay for a blue check, cannot make it go away
9.7M Followers 601 FollowingSix thousand years ago, someone invented the plow, and we all got wealthier. A gentle reminder that all civilizational wealth is driven by invention.