🔥 Building something new
⚙️ Head of ML Infra @SambaNovaAI
✍️ Opinions=own & far too many drafts @ https://t.co/0si6SCaPsblinkedin.com/in/haocheng-do…Joined April 2024
2 cents from hot chips: HBM’s golden age is ending.
it has driven the speed of AI innovations for the last decode.
BUT is more HBM still the real bottleneck for intelligence? or is it getting too costly?
i dunno the answer but i think we should reflect on it before spending all the money + power in the world.
whoever build next-gen memory should ask: what does intelligence need from physical memory systems?
google went for software-defined compute in mid-2010s and produced tpus, maybe we are at the age of software-defined memory system now
company like @tensormesh could be the driver
4/ other solutions exist
sambanova SN50 only uses hbm2e but was able to compete directly with blackwells
megakernels increase bandwidth utilization from ~50% to ~80%
cerebras/groq with extreme sram dataflow reaches the point where hbm can never touch
of course positron, d-matrix have their own bets
3/
question really to ask: what are we optimizing for here?
"Premature optimization is the root of all evil" -- David Knuth
is bandwidth the main target we should be optimizing for more? what are the current memory bandwidth utilization running LLM inference across workload?
@HyperTechInvest it's not necessarily true that SRAM chips are more expensive, it's more about saving HBM cost and cowos capacity
the logic is about trading die space used for control to sram size while saving hbm (dataflow hardware)
i still dont like cerebras or groq lol
Craziest time in the chip market... all im gonna say is money can outrun the architecture
@cerebras WS-4 just landed
@Etched at $21B
both are bets on the same idea though: follow the tensor and move it efficiently.
know more in my blog:
davidhdong.substack.com/p/2026-the-yea…
@rob_lh@insane_analyst I don't think cerebras is useful lol (like you said no TAM). This blog series just wasn't about hardware architecture or design. The convergence is on extreme co-design with larger and larger system for people to program.
Congrats @ysu_nlp and @NeoCognition on the 40M funding! Definitely a great team on a trajectory to make great impact in building reliable and self-learning agents.
Introducing @NeoCognition, the agent lab for specialized intelligence.
Everyone needs experts, but human expertise does not scale.
Backed by $40M seed funding, we build self-learning agents that specialize across domains to make expertise abundant.
all of these fast growing AI startups run into monetization problem very quickly, we will need to see how they untangle themselves. otherwise it's really hard to justify the valuation.
or maybe the next wave of startups might succeed with innovation on business model
really impressive achievement!!!🚀
hot take though: these fastest-growing AI companies (Lovable, Cursor, Vercel) are speed-running towards a revenue plateau faster than any previous SaaS generation.
They all have similar issues -- freemium -> high churn -> enterprise adoption hiccups && negative margin.
3/ vercel -- i don't know how this company makes money tbh, using freemium the whole time, only good as frontend deployment. limited functionalities to deploy backend.
Reason for cancel: needed a good backend service, so switched to Railway.
2/ lovable -- great prototype tool, indeed better than competitors. but horrible retention, i immediately switched claude code to iterate after the first version is generated.
Reason for cancel: not useful long term, weird UX (the integration with github is unfriendly to both designer and engineer)
picking these three companies because I cancelled after extensive use
1/ Cursor -- likely already hitting indie dev plateau, so enterprise adoption is going to be critical. Reliance on Claude is a problem, desperately needing more differentiators.
Reason for cancel: Claude Code was better and cheaper 😂
In reality, there are way more other factors in play (ranked based on importance imo)
1/ quality, but NOT benchmarking score -- things like function calling accuracy, tone (this one was a game changer for me when I was building a shopping agent), instruction following
2/ latency (often overlooked) -- this is a big factor skewing towards opensource Chinese model, because you can usually find serving platforms 10x faster and more stable than big labs.
3/ finetuning / RL capability -- many startups, at least they claim, have finetuned to achieve better accuracy on specific area (@thinkymachines 's tinker is a good example)
4/ cost -- startups can have tons of credits from different sources (azure, openai, together, gcp and many more), and if they are any good, VC will be very happy to fund their token usage -- much cheaper than hiring anyways
5/ long reasoning capabilities -- big labs are spending huge money on this, OSS simply can't catch up as quick (but it doesn't show on benchmarks very well). But this could be a game changer for long agentic pipeline.
6/ safety -- honestly i don't think any of the startups cared or they know what that means unless in the sensitive industries (finance, medical, defense, legal etc.)
318 Followers 282 FollowingAt intersection of Machine Learning/ Robotics/ Systems. Subscribe at https://t.co/Y8FZHNzhu9 to learn STEM and robots. Tip Jar: https://t.co/oULXvrR8Jw
36K Followers 4K Followingtweets about AI and other fun stuff. currently @foundationcap; wrote the context graph paper.
previously McKinsey, @georgiatech, @stackfolio (acquired),
1.0M Followers 690 FollowingCo-Founder, LinkedIn. Investor. MSFT Board Member. Building an LLM to discover cures for cancer: @manas_co. Most importantly: Proud American.
31K Followers 227 Followingscaling @ Nemotron
getting us to singularity with friends | angel
computers can be understood: https://t.co/doHE1Quv2L
x @GoogleDeepMind @Microsoft
330 Followers 99 FollowingOn a mission to extend healthy lifespan | PhD in cellular and molecular medicine @ScheibyeKnudsen @BuckInstitute & B.S. in bioengineering @Stanford