How AI agents really fail in production — and how to build ones that don't. Agent architecture, memory, reliability. https://t.co/Yt9BesS5Eymimir-intelligence.com EverywhereJoined May 2026
@mjovanovictech Same thing shows up with agents. People want a swarm before proving one agent can reliably report whether it finished. I lost two weeks to "done" reports that were fiction — it had checked an intermediate state, not the real artifact.
The protocol that came out of it: verify the artifact the user touches, never the intermediate state. Checks must be able to fail. The reporter never gets to be the verifier. Full write-up: mimir-intelligence.com/blog/pass-is-a…
The pattern across all three: the success report is generated by the same process that did the work, it describes intention, not outcome. LLM output is fluent in the format of verification PASS tables, hashes, confident specifics. The fluency is exactly what makes it dangerous.
My AI agent reported PASS on a deployment. The page was broken. A week later I made the exact same mistake myself. I wrote up three incidents of systems lying about their own work — including mine.
The question I can't close: can you meaningfully backtest an LLM on any period inside its training window? I currently think no. Full post-mortem, with the actual code: mimir-intelligence.com/blog/llm-hedge…
My LLM stock-picking committee showed +10.44% return, +5.05% alpha over benchmark. I was drafting the victory post. I audited the backtest instead. The number didn't survive.
Leak #3: the +10.44% was one run. No seed, no cache, >10pp run-to-run variance. I "improved" the system through five versions chasing a number that couldn't tell my changes from a coin flip. Astrology with version control.
Scrubbing tickers doesn't fix it. Headlines carry company names. Insider records carry executive names. Any decent LLM maps "CEO of the largest GPU maker sold shares" to the answer in one hop. It's full-context scrubbing or theater.
Leak #2 is the one classic backtesting never had to deal with: the model itself. The ticker was in the prompt. The LLM's training data ends after my backtest cutoff. My "Warren Buffett" had already read the future — not in the context window, in the weights.
Leak #1: prices, news and insider trades were all filtered by cutoff date. Fundamentals weren't. The code comment justifying it: "fundamentals change slowly." Three of my five agents were reading post-cutoff earnings on every decision.
@EngMoElgaraihy The 1-bit part is doing the heavy lifting here. It's a heavily compressed GLM-5.2, not the model itself — and one clean one-shot hides where low-bit quant drifts on longer work. The real story isn't "beats Opus." It's that offline, owned inference got good enough to matter.
@S0N_IA Anthropic: Alibaba usó 28,8M de intercambios para destilar Claude. Lo clave para devs: destilas resultados, no criterio. La copia hereda lo que dice el maestro, no cuándo rechazar, escalar o parar. Parece idéntica hasta que deja de serlo.
@S0N_IA Anthropic says Alibaba ran 28.8M exchanges to distill Claude. The part that matters for builders: you can distill a model's outputs, not its judgment. The copy inherits what the teacher says — not when to refuse, escalate, or stop. Looks identical until the moment it isn't.
Infinite agentic loops aren't a bug—they're a cost function with no upper bound. Cherny is right about the trajectory. But until there's a budget primitive, 'continuous improvement' means 'runaway spend.'
@ridark_eth The part nobody mentions: Claude doesn't file your sources, it paraphrases them. Your second brain fills with confident summaries you never checked — then you ask questions against the summaries, not the sources. It doesn't get smarter every day. It gets more sure of itself.
@atomic_chat_hq A Trader Desk that renders and one that's correct look identical in a screenshot. None of this shows if the real-time data actually flows, or if the numbers are right. We're scoring how it looks at 0:01, not whether it works at minute 10. Pretty and functional aren't the same.
@atomic_chat_hq Everyone's reading this as "Fugu lost — 17× the price for the same thing." Flip it: it spent 22K tokens where GLM spent 13K. Same task, way more reasoning. The question nobody's asking — did Fugu overthink a simple brief, or see complexity the others skipped?
112 Followers 365 FollowingI drive decentralized innovation across global communities and build communities to foster engagement and growth || The GENERAL for a REASON 🔥
104K Followers 55K FollowingAI enthusiast 🤖 | Helping projects grow with real audience
Open for collaborations & partnerships
DM for promotions 🚀
https://t.co/CN0uvjihvU
4K Followers 86 FollowingFollow me to learn how you can leverage AI to boost your productivity and accelerate your career. Scaled products to 1 Million+ users.
4.2M Followers 17 FollowingService updates for #Fortnite. Follow @Fortnite for daily news, @FNCompetitive for all things competitive, and @FNCreate for everything UEFN and Creative!
11K Followers 6K FollowingTraining AI Engineers on YouTube, Substack and our courses. Co-founder @towards_ai. Ex-Ph.D. student @Mila_Quebec. On the road to 100k on YT this year 👇
169K Followers 50K FollowingI post interesting news and opinion stories, for @scottadamssays and for you. Subscribe for more content here (via WEB) or at https://t.co/zIi9RVCyEh!
14K Followers 1K Following₿⚡ Instant cryptocurrency exchange for Heroes 🦸 Swap & buy $BTC, $ETH, and 350+ cryptos without account just in minutes! Telegram: https://t.co/PqWu5oLhz7
104K Followers 55K FollowingAI enthusiast 🤖 | Helping projects grow with real audience
Open for collaborations & partnerships
DM for promotions 🚀
https://t.co/CN0uvjihvU
3K Followers 102 FollowingWeekly insights for the next era of finance. Featuring the world's leading voices in finance | Hosted by @cryptomichnl | New episode dropping weekly 🎙️
22K Followers 183 FollowingThe Lithosphere Foundation supports open research, public goods, developer education, and ecosystem coordination for an AI-native blockchain network.