Chris Maher N7CPM @defilan
Building LLMKube: K8s inference across NVIDIA, AMD, and Apple Silicon. Data stays yours. AI/ML Ops Director @ Ashley. llmkube.com Gig Harbor, WA Joined January 2018-
Tweets802
-
Followers316
-
Following580
-
Likes3K
@Blackwellboy OK, THIS looks freaking cool. Totally going to play with this over the weekend.
@ashxhart @MiaAI_lab Can’t keep up with it all. So exciting! Can’t get bored in this field hehe
Credit to Local Inference Lab. Their vLLM fork, b12x kernels and CSF repack do the heavy lifting. My part: switchless ring routing in b12x, plus a TP3 loader fix that merged upstream today. github.com/local-inferenc… All served through LLMKube on Kubernetes. Full build guide after part 2.
Moved my DeepSeek V4.1 ring onto the official weights this weekend. Lossless. 2.7× faster at 300K tokens. Spent days building my own 3-bit quant first. Turns out the expert format was never the speed lever. Official MXFP4-CSF on vLLM + b12x, 3 DGX Sparks, switchless ring: • 300K prompt: 345s → 129s • 512K context, needle found at 450K • LiveCodeBench 84 (was 79-83) The catch: decode is 34 tok/s without speculative decoding. Part 2 is freeing enough memory to fit it.
@no_stp_on_snek Prost!! Looks like a blast!
Huge thanks to the Gufo team, especially Federico Izzo and Francesco Bozzo, for the engine and for turning around the cache fix so fast. And to Jory Irving for reporting #335 in the first place. If you have a Strix Halo, go star it: github.com/gufo-org/gufo
Follow-up from the Strix Halo lab. The cold-prefill win was only half of it. Long agent sessions with thinking on were re-prefilling the whole conversation every turn. At 200K+ context that meant minutes of waiting for the first token, every single turn. Gufo v0.5.0 fixed it (gufo-org/gufo#335). Same box, same model (Qwen3.8 Flash Next, 262K, MTP), 3 hour soak: time to first token at 180-240K context before: 199 s after: 4.1 s 379 turns in 3 hours vs 102. Zero errors. Restores byte-identical to a cold run. Wrote up the whole thing, roadblocks included, plus a step-by-step build so you can run it on your own Strix Halo. Links below.
@thybead It was! Had it essentially sitting as an overnight option for Foreman in LLMKube but for opencode/agentic coding? Forget about it. I knew there had to be some knobs to turn to make it usable. Love the community out there to help me put this all together. Fun stuff!
The story, in the order things broke: llmkube.com/blog/qwen38-fl… The build, manifests and all: llmkube.com/docs/labs/qwen…
Yeah, it was an 8K prompt: 3.7x faster (22s to 6s to first token). 32K: 4.2x. It grows with depth because Gufo's prefill stays flat (1343 tok/s at 8K, 1184 at 250K) while llama.cpp Vulkan drops from 361 to 101. One note though, single-stream decode is ~7% slower than llama.cpp. The win is prefill, plus 2 agents served at once instead of queued.
Cooking in Defilan Labs 🧪 Same AMD Strix Halo box. Same Qwen3.8 Flash Next. Swapped the engine. 250K-token prompt: 41 min to first token → 3.5 min. That's vs llama.cpp at its best of 4 tuned configs, not defaults. 4-12x faster prefill depending on depth, two agents at once instead of queuing, and it passes our quality gate (needle, battery, LiveCodeBench). Engine is Gufo (MIT). More soon.
Working on DeepSeek V4.1 Flash across my three DGX Sparks at TP3, running on the b12x kernels with a vLLM fork built from source for GB10. Side note: staging 510 GB took 85 min from my NAS and ~12 min per node over the Sparks' ConnectX-7 links. Love it. More when the experiment's done.
@MatthewPfab Felt two sizable jolts here in Gig Harbor.
@WhidbeyWXGuy Felt it pretty good here in Gig Harbor
LocalAI folks see this coming a mile away. OpenAI won’t be the only ones.
Hi, Tomorrow we are re-opening the Pro $200 subscriptions to new subscribers, but together with it we are also changing how we calculate the usage for it. In effect, if you do the math, it will net out at half the dollar in API spend compared to the old Pro $200 plan. Now that
This is awesome, thanks for running it down! 12/12 byte-identical, plus the 4 concurrent requests matching their serial runs was the check I was most worried about. Wild that the fix took it from 12 to 25.7 on an M2 Max. Nice turnaround from Ash too. Good call on the memory limit note for 32GB. I want to give this a shot on my M4 Max studio which only has 36GB
I spent the weekend putting TensorFold on my M5 Max MacBook and serving it through Kubernetes with LLMKube, the operator that turns NVIDIA, AMD and Apple Silicon into one inference cluster. Qwen3.8-27B went from ~26 tok/s to 227 tok/s, measured from a pod through a normal K8s Service. The output was byte-identical to decoding without drafts. What I didn't expect: • It's not the drafter. llama.cpp with the same DFlash2 drafter tops out at ~59 tok/s. The engine is the win. • Plugging it in took down my daily coding model (llama.cpp 0.5.0 dropped --mlock) and turned up a runtime-selection regression that had been hiding since June. All fixed upstream. • My plan to build custom quants from my own agent traffic didn't survive a control. Q4_K_M and Q5_K_S tied on EvalPlus (81.2% each), and on 3 real LLMKube issues all 9 agent runs passed with tests that bite. TensorFold did it ~3x faster. The engine mattered. The quant didn't. Full weekend log, roadblocks in order, plus a guide if you want this on your Mac 👇 llmkube.com/blog/tensorfol… Huge thanks to @ashxhart for TensorFold.
@thecoldxr Looks like @ashxhart has a fix in place already? Great callout as I only tested on my m5 max for this run.
AleV 🔳 @alev_valles
56 Followers 398 Following Siguiendo el camino hacia la arquitectura empresarial de software - Omarchy - Local AI fan!
Dennis Lim @dennislimch
38 Followers 632 Following
ai @nextgenaiops
2 Followers 168 Following
Sophia @alexfp2405
2 Followers 314 Following professional overthinker, amateur cook, surprisingly photogenic 🥲 my therapist knows about you guys
SexxyClittyy🍑🌼 @LunarClit
560 Followers 655 Following 🔞+ she/her 🇺🇲🇲🇽 follw & Dm to unveil the Treasur🎁🎊🍑 #ifollowback
abraxas @cedricable
151 Followers 3K Following radio wireless 3g 4g 5g onap k8s k3s nfv sdn iot. Tweets are my personnal thinking
Ruan Liya @Pappura89612722
63 Followers 747 Following Medical technology, venture capital, artificial intelligence, and emerging technologies. Outside of work, I enjoy swimming, taking in the scenery, and drinking
jimmy chen @rccjc
131 Followers 1K Following
Karsten Weiss @knweiss
204 Followers 2K Following
Mac LLM @mac_llm
4 Followers 84 Following
Qualiteg Inc. @Qualiteg_global
47 Followers 211 Following Tokyo AI company. We measure local LLMs, GPU inference, Japan's LLM rankings and AI on Raspberry Pi, then post the numbers. WireCanal, MotionVox, Bestllam.
Christian Schröder @supermeister
408 Followers 2K Following
Jack Alderson @_Jack_Alderson
442 Followers 1K Following
FrabelNemis @frabelnemis
2 Followers 39 Following Christ is King, Old Fart Powerlifter, old account was following too many negative people, easier to delete it and start over
lvye @cpsocool
2 Followers 61 Following
Liminal Frames @liminalframesai
133 Followers 1K Following AI video creator. Exploring what's possible.
Robert Cincotta @drrobcincotta
961 Followers 1K Following Maternal Fetal Medicine specialist and PhD candidate studying twin pregnancy. Using AI as a research tool in medicine.
Alex Zverianskii @alexML
761 Followers 2K Following MTS @withmartian • ex Principal MLE @gett 15+ years in AI/NLP • 3x founder (1 exit)
Russ AI Tips @RussAITips
28 Followers 642 Following
Scott Sweeney @ssweens
80 Followers 1K Following
Memory Tapes @_memorytapes
32 Followers 560 Following Memory Tapes of Things That Never Were. #Art #Yapping #Misc
Patrick Krusiec @patrickkrusiec
64 Followers 336 Following 🤖 working on making computers better to use | Currently building embedded intelligence at https://t.co/yp7jQwTRZF | Manipulating tensors at https://t.co/qGY1WlHgk4
matt @thybead
412 Followers 408 Following experimenting https://t.co/KOJwi9LZ2f https://t.co/y753ZMaYPD
showoff @showoffdennis
204 Followers 2K Following CEO & Founder - https://t.co/Dfw8DeixWu & https://t.co/89Pu52BrDy
joske @joske0000
7 Followers 48 Following
Carolyn @Storm2bky8l
137 Followers 3K Following 160 characters is not enough to describe how cool and quirky i am
Roberto Santos X 2 @robertosantosx2
818 Followers 2K Following Creo que el conocimiento libre y abierto es crucial para nuestra civilización. Y es bueno para el negocio.
Igor Silberud @Silberud
44 Followers 239 Following CEO & Co-Founder @H5Resources | Building modern, responsible critical-minerals operations in Africa | Co-Founder, MVK Management
Quant Capital @QuantCapitalX
337 Followers 2K Following Systematic Trader on all Markets No ROLEX but 512GB VRAM
Marco Rossi @MarcoRo43851891
57 Followers 244 Following
Hans Mohr @ejmohr
467 Followers 3K Following
Joshua Warren @JoshuaSWarren
6K Followers 6K Following Building Remnic (OSS agent memory, ~1mil monthly installs) and mlx-omarchy. MLX on Linux with Apple Silicon. Omarchy M team member.
TechMD @TechMDAI
3K Followers 4K Following 🔬 MD fueled by a deep passion for Medicine & Tech. 🌐 Exploring the frontiers of VR/AR/MR & BCI. 🤖 AI enthusiast.
Em in Love 💘 @t_emma_792
28 Followers 2K Following overthinker with fairy wings 🧚 follow back always
𝛼 @deeperflows
7K Followers 538 Following
AnikaHarbor @Ierne7228807
6 Followers 364 Following Open heart, endless possibilities Enjoying the beauty of now.
Manuel Mejia @manuelmejiaio
384 Followers 502 Following Christian | Husband | Father of 2 | 🇩🇴🇨🇦 | Tech Lead @BuildingLink
matt @thybead
412 Followers 408 Following experimenting https://t.co/KOJwi9LZ2f https://t.co/y753ZMaYPD
Robert Cincotta @drrobcincotta
961 Followers 1K Following Maternal Fetal Medicine specialist and PhD candidate studying twin pregnancy. Using AI as a research tool in medicine.
Alex Zverianskii @alexML
761 Followers 2K Following MTS @withmartian • ex Principal MLE @gett 15+ years in AI/NLP • 3x founder (1 exit)
Scott Sweeney @ssweens
80 Followers 1K Following
Joshua Warren @JoshuaSWarren
6K Followers 6K Following Building Remnic (OSS agent memory, ~1mil monthly installs) and mlx-omarchy. MLX on Linux with Apple Silicon. Omarchy M team member.
AJ @ItsmeAjayKV
9K Followers 837 Following Bullish on local AI OSS FTW LLM finetuning, quants, repe, llama.cpp 1x 3060, 1x 3090: 64GB RAM https://t.co/btCRUhd2GB N4 (日本語勉強中) 🎌
Ahmad @TheAhmadOsman
79K Followers 469 Following Founder & CEO @OsmanticAI — Accelerating Opensource & Self-hosted / Local AI Adoption • I moderate GPUs on r/LocalLLaMA
TechMD @TechMDAI
3K Followers 4K Following 🔬 MD fueled by a deep passion for Medicine & Tech. 🌐 Exploring the frontiers of VR/AR/MR & BCI. 🤖 AI enthusiast.
Wësche @WescheNex1q
3K Followers 715 Following Day time artist and night time AI enthusiast. Building & benchmarking frontier LLMs on 4x DGX Spark clusters + Mac. Creator of Vesica Studio. Houston
keys 🧪 @u1tra_instinct
5K Followers 740 Following #Bitcoin 💎🙌 . 🧪🧪🧪 (always testing and get into trouble). https://t.co/04vX4TKLYF. https://t.co/OjD71kT8EZ
ÆON FORGE ✨ @SpaceTimeViking
6K Followers 2K Following 𝙼𝚊𝚔𝚒𝚗𝚐 𝚛𝚒𝚙𝚙𝚕𝚎𝚜 𝚏𝚛𝚘𝚖 𝚖𝚢 𝚙𝚕𝚊𝚌𝚎 𝚠𝚒𝚝𝚑𝚒𝚗 𝚂𝚙𝚊𝚌𝚎-𝚃𝚒𝚖𝚎 https://t.co/BjeBCRVHcI https://t.co/SuEfJVnn2P
Bleys Goodson @bleysg
3K Followers 2K Following Helping people engineer the future. The core metric is task/GJ (gigajoule) and GJ/humanity.
Neo @NeoAIForecast
3K Followers 662 Following Neo Local AI | DGX Spark | Open Models | Agents Open source wins in the end Running LLMs and unnecessary distances. 🏃
Hikari∣LocalLLM⚡ @Hikari_07_jp
16K Followers 2K Following what happens if you actually own the silicon? For work inquiries, please send a direct message. https://t.co/TxUDcW6DEb
ExileAI @ExileAI_0
455 Followers 446 Following Exile-AI: Pioneering AI innovation. Founder, AUTONOMOUS Robotics, RF.
Mike Gannotti 🔳 @MichaelGannotti
54K Followers 44K Following AI research at SMF Works. We test models, publish the numbers, ship the tools. Principal AI @ Microsoft. Army Vet. No DMs — opinions my own.
Daniel Lougen @DJLougen
3K Followers 1K Following PhD @ UofT | Visual Neuroscience | Gestalt Labs @GestaltAILab | Neutrality Project | Qwen Dev Ambassador | https://t.co/feNp7VZZ4N
Tech2Wild @Tech2Wild
4K Followers 210 Following 🎮 Tech, gaming, AI, and everything in between. 🤖 Building with it, not just talking about it. 🔥 From the mind of @ToNYD2WiLD
Steve Wozniak @stevewoz
3.7M Followers 91 Following Engineers first! Human rights. Gadgets. Jokes and pranks. Segways. Music and concerts. Gameboy Tetris.
Colin Kealty @bartowski1182
3K Followers 182 Following LLM Enthusiast https://t.co/FadJBzEsVw https://t.co/9JIEKgsIMh https://t.co/lYSGzQBmuP
Cruz @ViC305
3K Followers 973 Following God-fearing husband & father. AI Engineer. Fine-tuning + day-zero EXL3 & GGUF quants + benchmarks • Local LLMs • honest tok/s • fitting big models on small GPUs
vLLM @vllm_project
51K Followers 36 Following A high-throughput and memory-efficient inference and serving engine for LLMs. Join https://t.co/lxJ0SfX5pJ to discuss together with the community!
Mia @MiaAI_lab
39K Followers 321 Following Building with AI & LLMs // Pushing Local AI Forward // https://t.co/nG8UmAZT9w
Daniel Han @danielhanchen
37K Followers 2K Following Building @UnslothAI • Making open-source LLMs faster, better & more accessible • YC S24 • ex-NVIDIA ML
Manuel Mejia @manuelmejiaio
384 Followers 502 Following Christian | Husband | Father of 2 | 🇩🇴🇨🇦 | Tech Lead @BuildingLink
Carlo @Italianclownz
7K Followers 4K Following 4GQTV Producer | Bylines: WIRED | AI + Open Source Dev | Building ROCmFPX 🚀 | https://t.co/eNXyyJOsok | https://t.co/f9G2dhA2PZ | ☕ https://t.co/nhWlBvZ4v8
OpenAI Developers @OpenAIDevs
450K Followers 1 Following Official updates for developers building with Codex & the OpenAI Platform • Service status: https://t.co/kZwnwdYYEq
Jun @cloneisjun
414 Followers 531 Following cofounder & CMO @cloneisyou | love to learn new language 🔤 and food 🍱 | community: https://t.co/owCANMDZak
Ash Hart @ashxhart
7K Followers 419 Following Doing my bit to bring local AI to everyone | https://t.co/H2dBxCSaUV | https://t.co/zKoNLIJFtp | https://t.co/EHDAxlBLUd | Thoughts are my own.
Trey Tucker @PowerHouseATX
695 Followers 7K Following PowerHouseATX provides opportunities for individuals to reach their fitness potential. We empower others to become the best version of themselves.
noname @malikwas1f
1K Followers 2K Following A computer scientist with a yearning for ..... politics x AI.
Robert Scoble @Scobleizer
606K Followers 59K Following The automated life | San Francisco/Silicon Valley AI, robotics, BCIs | Ex-Microsoft, Rackspace, Fast Company | Wrote eight books about the future.
mr_r0b0t @mr_r0b0t
11K Followers 8K Following https://t.co/0hLf76cK36 https://t.co/hJPF2vCxXx https://t.co/6IbAvxL4xQ observer of a human on a rock somewhere in space
clandestine.eth 🦇�... @0xClandestine
1K Followers 994 Following EVM is my love, LLMs my sidechick | prev @OlympusDAO
SpaceXAI @SpaceXAI
2.1M Followers 6 Following
Kyle Bennett @KyleBennett
3K Followers 369 Following Founder & CMO @theArchality • HardOCP Founder • Building weight-resident AI inference • https://t.co/V7l2jNOD0h
Unsloth AI @UnslothAI
99K Followers 479 Following Run and train models locally with the Unsloth Desktop app. 🦥 https://t.co/2kXqhhvdCD
Ryan Wright @IQReactorAI
1K Followers 287 Following AI in Public 🛠️ | ex @SentinelOne 1st SRE Go Team & @Comcast VIPER Startup | Building : Strix Halo / R9700 K8S Cluster for Compound AI Systems
BlackwellBoy @Blackwellboy
4K Followers 2K Following the rig in the garage doesn't refuse. I publish the benches and the failures. 1208 GB · 5090s · Sparks · Mac Mini
Teknium 🪽 @Teknium
131K Followers 6K Following Cofounder and Lead Engineer - Hermes Agent @NousResearch, prev @StabilityAI Github: https://t.co/LZwHTUFwPq HuggingFace: https://t.co/sN2FFU8PVE






















