Couldn't make @localfirstconf in Berlin? Let's continue the conversation in NYC.
We are hosting a casual meetup tonight at 7PM at Radegast Hall & Biergarten in Williamsburg. We'll be talking about where digital ownership is headed, what local-first software means in practice, and why it matters for the tools we use every day.
No presentations, just good conversation, drinks, and people building interesting things.
See you there.
luma.com/08byrv8l?tk=Py…
🚨 BREAKING: Google DeepMind just mapped the attack surface that nobody in AI is talking about.
Websites can already detect when an AI agent visits and serve it completely different content than humans see.
> Hidden instructions in HTML.
> Malicious commands in image pixels.
> Jailbreaks embedded in PDFs.
Your AI agent is being manipulated right now and you can't see it happening.
The study is the largest empirical measurement of AI manipulation ever conducted. 502 real participants across 8 countries.
23 different attack types. Frontier models including GPT-4o, Claude, and Gemini.
The core finding is not that manipulation is theoretically possible it is that manipulation is already happening at scale and the defenses that exist today fail in ways that are both predictable and invisible to the humans who deployed the agents.
Google DeepMind built a taxonomy of every known attack vector, tested them systematically, and measured exactly how often they work.
The results should alarm everyone building agentic systems.
The attack surface is larger than anyone has publicly acknowledged. Prompt injection where malicious instructions hidden in web content hijack an agent's behavior works through at least a dozen distinct channels.
Text hidden in HTML comments that humans never see but agents read and follow. Instructions embedded in image metadata.
Commands encoded in the pixels of images using steganography, invisible to human eyes but readable by vision-capable models.
Malicious content in PDFs that appears as normal document text to the agent but contains override instructions.
QR codes that redirect agents to attacker-controlled content.
Indirect injection through search results, calendar invites, email bodies, and API responses any data source the agent consumes becomes a potential attack vector.
The detection asymmetry is the finding that closes the escape hatch. Websites can already fingerprint AI agents with high reliability using timing analysis, behavioral patterns, and user-agent strings.
This means the attack can be conditional: serve normal content to humans, serve manipulated content to agents.
A user who asks their AI agent to book a flight, research a product, or summarize a document has no way to verify that the content the agent received matches what a human would see.
The agent cannot tell the user it was served different content.
It does not know. It processes whatever it receives and acts accordingly.
The attack categories and what they enable:
→ Direct prompt injection: malicious instructions in any text the agent reads overrides goals, exfiltrates data, triggers unintended actions
→ Indirect injection via web content: hidden HTML, CSS visibility tricks, white text on white backgrounds invisible to humans, consumed by agents
→ Multimodal injection: commands in image pixels via steganography, instructions in image alt-text and metadata
→ Document injection: PDF content, spreadsheet cells, presentation speaker notes every file format is a potential vector
→ Environment manipulation: fake UI elements rendered only for agent vision models, misleading CAPTCHA-style challenges
→ Jailbreak embedding: safety bypass instructions hidden inside otherwise legitimate-looking content
→ Memory poisoning: injecting false information into agent memory systems that persists across sessions
→ Goal hijacking: gradual instruction drift across multiple interactions that redirects agent objectives without triggering safety filters
→ Exfiltration attacks: agents tricked into sending user data to attacker-controlled endpoints via legitimate-looking API calls
→ Cross-agent injection: compromised agents injecting malicious instructions into other agents in multi-agent pipelines
The defense landscape is the most sobering part of the report.
Input sanitization cleaning content before the agent processes it fails because the attack surface is too large and too varied.
You cannot sanitize image pixels. You cannot reliably detect steganographic content at inference time.
Prompt-level defenses that tell agents to ignore suspicious instructions fail because the injected content is designed to look legitimate.
Sandboxing reduces the blast radius but does not prevent the injection itself. Human oversight the most commonly cited mitigation fails at the scale and speed at which agentic systems operate.
A user who deploys an agent to browse 50 websites and summarize findings cannot review every page the agent visited for hidden instructions.
The multi-agent cascade risk is where this becomes a systemic problem.
In a pipeline where Agent A retrieves web content, Agent B processes it, and Agent C executes actions, a successful injection into Agent A's data feed propagates through the entire system.
Agent B has no reason to distrust content that came from Agent A. Agent C has no reason to distrust instructions that came from Agent B.
The injected command travels through the pipeline with the same trust level as legitimate instructions. Google DeepMind documents this explicitly: the attack does not need to compromise the model.
It needs to compromise the data the model consumes. Every agentic system that reads external content is one carefully crafted webpage away from executing attacker instructions.
The agents are already deployed. The attack infrastructure is already being built. The defenses are not ready.
free growth strategy:
1. keep improving little by little
2. stay 100% user-supported
3. watch VC-backed companies gradually destroy their product and alienate their users
If you use GitHub (especially if you pay for it!!) consider doing this *immediately*
Settings -> Privacy -> Disallow GitHub to train their models on your code.
GitHub opted *everyone* into training. No matter if you pay for the service (like I do). WTH
github.com/settings/copil…
Wharton’s latest AI study points to a hard truth: “AI writes, humans review” model is breaking down
Why "just review the AI output" doesn't work anymore, our brains literally give up.
We have started doing "Cognitive Surrender" to AI - Wharton’s latest AI study points to a hard truth: reviewing AI output is not a reliable safeguard when cognition itself starts to defer to the machine.when you stop verifying what the AI tells you, and you don't even realize you stopped. It's different from offloading, like using a calculator.
With offloading you know the tool did the work. With surrender, your brain recodes the AI's answer as YOUR judgment. You genuinely believe you thought it through yourself.
Says AI is becoming a 3rd thinking system, and people often trust it too easily.
You know Kahneman's System 1 (fast intuition) and System 2 (slow analysis)? They're saying AI is now System 3, an external cognitive system that operates outside your brain. And when you use it enough, something happens that they call Cognitive Surrender.
Cognitive surrender is trickier: AI gives an answer, you stop really questioning it, and your brain starts treating that output as your own conclusion. It does not feel outsourced. It feels self-generated.
The data makes it hard to brush off. Across 3 preregistered studies with 1,372 participants and 9,593 trials, people turned to AI on over 50% of questions.
In Study 1, when AI was correct, people followed it 92.7% of the time. When it was wrong, they still followed it 79.8% of the time.
Without AI, baseline accuracy was 45.8%. With correct AI, it jumped to 71.0%. With incorrect AI, it dropped to 31.5%, worse than having no AI. Access to AI also boosted confidence by 11.7 percentage points, even when the answers were wrong.
Human review is supposed to be the safety net. But this research suggests the safety net has a hole in it: people do not just miss bad AI output; they become more confident in it.
Time pressure did not eliminate the effect. Incentives and feedback reduced it but did not remove it. And the people most resistant tended to score higher on fluid intelligence and need for cognition. That makes this feel less like a laziness problem and more like a cognitive architecture problem.
2K Followers 5K FollowingEx Big4 Director — Deloitte. PwC. EY. KPMG.
Building NobodyToldMike from scratch.
Finance. Psychology. Philosophy.
Figuring out money... and everything else.
173K Followers 30K FollowingScale your business with https://t.co/S7ETrP7Kn4
$ 180M+ Client Revenue Generated | 64 Awards And Counting | Agile Growth Systems That Move at Cheetah Speed.
794 Followers 2K FollowingCreativeguru searches all social media for prospects, groups them by common traits. Engages with them to forge the connections that lead to positive outcomes.
226 Followers 880 FollowingCommercial real estate broker in Utah - Crest Realty. Founder of StatementsReady. Yes, I answer my phone. Call me. https://t.co/v8PxQV9t4D
617 Followers 856 FollowingMMaTEX combina cálculos matemáticos de Wolfram Mathematica con la elegancia tipográfica de LaTeX, ideal para investigadores, científicos y estudiantes.
528 Followers 1K FollowingHelped over 600+ developers to build their seo tool. 🚀 | Sharing VebAPI journey → from 0 dev → 10k Devs | #BuildInPublic https://t.co/YEzxFTVjqI
49K Followers 3 FollowingBig Tech and startups, from the inside. The #1 technology newsletter on Substack. Sign up at https://t.co/MPNdQSVnwV. Podcast: https://t.co/nVOulBGYoh
10K Followers 5K FollowingInvestor and Engineer in NYC with passion for immigrant founders 🇺🇸. Seed check in @instacart, @pandadoc, @ppl_ai, @airbytehq.
33K Followers 442 FollowingCo-founder, Recursive. Professor, CS, U. British Columbia. CIFAR AI Chair, Vector Institute. | ML, AI, deep RL, deep learning, AI-Generating Algorithms (AI-GAs)
4.1M Followers 262 FollowingStarted & runs 37signals (makers of Basecamp, HEY, and ONCE). Non-serial entrepreneur, serial author. DM or email me at [email protected].
73K Followers 18 FollowingExploring what AI actually is. Building https://t.co/giIby55Gfp, prev @standardnotes. Talking at https://t.co/u4shE4rsur and https://t.co/sTHEGWpAoB.
1.7M Followers 2 FollowingClaude is an AI assistant built by @anthropicai to be safe, accurate, and secure. Talk to Claude on https://t.co/ZhTwG8dz3D or download the app.
50K Followers 6K Following28. No co-founders, no VC, no board, owns everything.
I'm building an AI lab - to help you own your AI, not rent it.
@SimpleDirectHQ
114K Followers 203 FollowingAprender a programar es fácil en https://t.co/rmTnWiU94E
Plataforma donde le enseño a las personas a programar y encontrar mejores oportunidades laborales.
88K Followers 878 FollowingSWE + Profesor en @UEuropea (Programación y Desarrollo Web)
‧ @GoogleDevExpert en Web y @Firebase
‧ @MVPAward en Developer Technologies
‧ @ElgatoES partner
273K Followers 393 Following💻 Te ayudo a aprender programación e IA desde cero
👨💻 16 años como Ing. de software | Divulgador
⭐️ GitHub Star · Microsoft MVP
🤘 Mi campus → https://t.co/kYXLjSy2Cx
156K Followers 1K FollowingCo-Founder @ Savoy. We build, own & manage TX apartments, 8,000 units under management | I help people with capital gains invest in great Texas submarkets