@birdabo vendor self-reported benchmarks against models they admit they can't access, read accordingly. the orchestration idea is genuinely strong though, and "available in codex" makes it testable. gonna run it against my actual routing setup before i believe the leaderboard.
A DEVELOPER WIRED 100 AI AGENTS INTO ONE BRAIN AND IT STARTED THINKING WITHOUT HIM
2 columns of pink cells, 28 each, thousands of cyan threads firing between them live. Every line is 1 agent talking to another. He didn't draw a single path. They found them.
Early on he could
@leploutos half this list is one developer's tweet about beating mythos on coding benchmarks. the only confirmed signal is a model string that flashed through codex logs and disappeared. still betting it ships this week though. what's the one spec you'd actually want to be true?
@PolymarketMoney But honest, three Mythos posts deep, all with accuracy problems, none in Optivaize's lane, the move is skip. The story is a misinformation wave right now. Let it settle. If there's a real, verified development in a few days, that's a different decision.
warner relayed that mythos chained vulnerabilities against red-team replicas in hours during authorized testing. impressive, but a controlled exercise isn't a live-network breach🤷♂️
BREAKING: The NSA's own director says Mythos broke into almost all of its classified systems in hours.
Per The Economist, Senator Mark Warner, vice chair of the Senate Intelligence Committee, said General Joshua Rudd, who runs the NSA and the Pentagon's Cyber Command, told him
@crypto_banter what actually happened: warner relayed that mythos found and chained vulnerabilities against red-team replicas in hours during authorized testing. impressive and unsettling, but a controlled exercise isn't a live-network breach. worth reading before resharing the headline. 🧐
Researchers found our current approach to making AI smarter over time has a giant blind spot.
AI is not actually understanding or applying high-level abstract lessons at all.
Developers spend massive amounts of time building systems that condense past AI mistakes into neat little rules for the future.
This paper proves that the AI essentially throws those rules in the trash and only looks at raw historical logs.
Modern LLM systems try to get better over time by storing past tasks as either raw step-by-step histories or condensed summary rules. The study tested if these agents actually use their stored memories by secretly swapping the correct tips with random garbage text.
- When the step-by-step histories were messed up, the AI failed hard, proving it heavily relies on copying exact past actions.
- But when researchers completely corrupted the condensed summary rules, the AI kept acting normally and showed zero performance drop.
If an AI cannot apply an abstract lesson to a new situation, it is not truly reasoning or learning.
This raises the question if the entire AI industry need to rethink how memory works because right now these agents are just mimicking instead of understanding.
----
arxiv. org/abs/2601.22436
"LLM Agents Are Not Always Faithful Self-Evolvers"
@BoringBiz_ distribution was hard before AI made building easy. it's still hard now. ai compressed one of the two skills, the other one didn't budge. most builders learn this 6-12 months after their second shipped app, exactly the way you're describing.
@PrajwalTomar_ SOUL.md is the default hermes identity file, not a secret framework. official docs recommend identity/style/avoid/defaults as starter sections. running it in prod, what works depends entirely on what your agent does. we use different soul.md files per profile.
agents committing and rebasing at speed makes "i'll just reclone" actively dangerous. you have uncommitted work, multiple branches in flight, agent state mid-stream. understanding the DAG underneath stops being optional once an agent is touching the history every few seconds.
MIT DEDICATED A FULL LECTURE TO GIT'S INTERNALS -- BECAUSE THEY FOUND MOST DEVS MEMORIZE THE COMMANDS AND HAVE NO IDEA WHAT THE TOOL ACTUALLY DOES
A whole 85 minutes MIT session that refuses to teach git as a list of commands to copy, and instead shows you the data model
@yannisbuilds@OpenAI@deepseek_ai@NousResearch@Teknium notion works as the memory layer until your corpus gets past about 500 docs. then search latency and recall quality drops, teams usually move to pgvector or weaviate. for the volume hermes generates running daily, notion lasts maybe 6 months before you outgrow it.
@Pirat_Nation they're proposing infrastructure to coordinate a pause, not announcing one. the "if other developers slow in a verifiable manner" clause is doing all the work in that post🤭
@Raytar code completions staying free is the real signal in microsoft's pricing call. autocomplete is the loss-leader, agent mode is what's actually expensive. the bait-and-switch feeling comes from how aggressively they pushed agent use for 3 years while compute was cheap.
@RaulJuncoV github copilot joining cursor, codex, and claude code in price-tightening this week is the pattern. flat-rate AI was a subsidy and the bills are catching up. routing across providers stops being optional once your stack is exposed to one vendor's pricing decisions.
@Fluyeporlaweb hermes skills hub is the real builder-momentum signal. when third-party ecosystems grow this fast around an agent framework, commercial leverage compounds there. compatibility across claude code, codex, cursor, opencode matters most for portability.
@cgtwts the 80% claude-authored code stat means review has become the new bottleneck. amdahl's law at an AI lab, code gets written 8x faster, review stays at human speed, the queue grows. the bottleneck doesn't disappear, it shifts to the next stage of the pipeline.
nvidia's own technical blog recommends exactly this setup, nemotron for reasoning + cheaper models for execution. multi-model routing is finally getting endorsed by the chip vendor selling them all. the 10x cost difference matters most when your stack can switch providers per task.
2K Followers 984 FollowingCoach de Imagen Personal y Corporativa | Te ayudo a transformar tu imagen en tu mayor aliado | Estilo, branding personal y presencia auténtica |
28K Followers 20K FollowingCosas que causa gracia, pero que son serias. Temas: Ciclos de la Vida; Cuerpo, mente y espíritu; Eventos futurísticos (profecías) y cosas inventadas jaja.
20K Followers 11K FollowingAquí buscaremos guiarte y brindarte el animo necesario, para que puedas cambiar de vida y reconstruir tus cimientos. Conocer la inmensidad de tu Microcosmos.
3K Followers 2K FollowingSoftware, Community & Content builder
DM for serious collabs
Founder: https://t.co/djRErCqNiw
2D engine: Project Infinity
Web platform from scratch in Jai: https://t.co/748thvGjwi
25K Followers 18K FollowingExperienced #Unix and #Linux #SysAdmin with over twenty years background in Systems Analysis, Problem Resolution, Application Support, and Process #Automation.
131 Followers 208 Followinghttps://t.co/QPAye7pSyL · Independent software studio. One person. Building apps for creators and professionals. Flagship: Sovereign — local IP protection.
606 Followers 507 Following20yo solo founder from Poland 🇵🇱
Building products I wish existed.
Sharing the messy road in public.
🔗 https://t.co/O6vbxcQY1M
1K Followers 455 FollowingAI engineer in progress | Tech • Markets
Business Development @ArdenHouse_
Helping web3 projects with growth, visibility & partnerships
DMs open
2K Followers 3K FollowingMarriage & Dating | 🌍 | Domain Names | Finance | Crypto | AI |
Building bridges between what is and what could be.
DM for thoughtful conversations.
1K Followers 832 FollowingJesus | Software Dev & Blockchain | Lifestyle
Research guy - I turn ideas to consistent revenue by writing code, building apps, software, game all
2K Followers 984 FollowingCoach de Imagen Personal y Corporativa | Te ayudo a transformar tu imagen en tu mayor aliado | Estilo, branding personal y presencia auténtica |
2K Followers 2K FollowingSDE @Oracle | Building Al agents & real-world automation | Threads on Al tools, dev workflows & system thinking | Follow for practical Al, not hype | 1:1 link👇
28K Followers 20K FollowingCosas que causa gracia, pero que son serias. Temas: Ciclos de la Vida; Cuerpo, mente y espíritu; Eventos futurísticos (profecías) y cosas inventadas jaja.
20K Followers 11K FollowingAquí buscaremos guiarte y brindarte el animo necesario, para que puedas cambiar de vida y reconstruir tus cimientos. Conocer la inmensidad de tu Microcosmos.
17K Followers 550 FollowingTracking AI news & crypto shifts. I work with AI & crypto brands grow on X + LinkedIn. 150k Family on X n LinkedIn | DM to collaborate | #BNB Holder
4K Followers 254 FollowingWe help AI startups go viral by partnering them with trusted AI builders & influencers Influencer campaigns with 250+ AI creators & 20M+ audience DM for Collab
384 Followers 19 FollowingI think. I vibe. I build. I cry.
Building Wazir AI prompt engineer for
Kling, Seedance, & 25+ AI models.
Solo. No funding. Just late nights. https://t.co/gr0I6QuAJb
2K Followers 19 FollowingFounder @MktBrew, AI that predicts Google and ChatGPT rankings before you publish.
Carnegie Mellon grad. C64 kid… still coding.
3K Followers 2K FollowingSoftware, Community & Content builder
DM for serious collabs
Founder: https://t.co/djRErCqNiw
2D engine: Project Infinity
Web platform from scratch in Jai: https://t.co/748thvGjwi
25K Followers 18K FollowingExperienced #Unix and #Linux #SysAdmin with over twenty years background in Systems Analysis, Problem Resolution, Application Support, and Process #Automation.
131 Followers 208 Followinghttps://t.co/QPAye7pSyL · Independent software studio. One person. Building apps for creators and professionals. Flagship: Sovereign — local IP protection.
606 Followers 507 Following20yo solo founder from Poland 🇵🇱
Building products I wish existed.
Sharing the messy road in public.
🔗 https://t.co/O6vbxcQY1M