training models from scratch gives you so much control on core competencies of the final models. Say you wanna deploy a fast model for cybersecurity applications, you can verifiably steer the model during “pre”-training towards fundamental behavior that during “post”-training” allows for the strongest guardrails reliability & resiliency.
This you cannot do with a checkpoint you don’t know its pre-training data, and initial conditions.
Over the summer, we have been sprinting on safety priorities; it's more important than ever for capabilities and safeguards to advance together. We have more to do but have made a lot of progress. We are also going to be launching our next model soon.
There is an obvious tension here: on one hand, Astra is very good and we are excited to see what people will build with it. We are proud of our work.
On the other hand, we are clearly in a phase of development where we believe caution is warranted, and we are pacing our progress to ensure that we can meet the safety standards required by new capability levels.
Astra has been done training for a while now and is a significant step forward in both capabilities and alignment. For the models after that, we have been slowing things as needed to ensure that we can do sufficient work on safety and alignment.
AI is getting extremely capable; no one fully understands the consequences of this. Managing the transition to a world with abundant and powerful AI to optimize for safety and benefits to people should be one of the highest priorities in the world. It is our highest priority at OpenAI.
We have been living with the tension between being excited and anxious about progress for some time, and it is still discordant for us. We know it is much more discordant for other people. And yet, we believe strongly that the world needs to understand where AI is going and how models perform in the real world. More importantly, we believe the world will need aligned AI to manage the future phases of this transition.
An iterative loop where society and this technology evolve together is what will lead to the highest chance of getting this right.
So we hope you enjoy our new model, and we hope the world continues to take what’s happening in AI extremely seriously.
Gavin, spot on.
AI is bringing manufacturing back to America and reindustrializing the nation after decades of offshoring.
AI is creating demand that drives investment in our aging power grid and sustainable energy, powered by market forces, not subsidies.
AI is creating construction and manufacturing jobs across energy plants, chip fabs and data centers.
AI is creating new companies and industries. $400 billion has been invested in AI startups in the past six months alone.
Builders must partner with communities to build in their hometowns, earn trust and create local benefits.
We have an opportunity to create lasting benefits for communities across America and help America lead the next industrial revolution.
Regret the tone of my post on data centers yesterday.
What I should have said:
There were reasonable concerns about data centers 18ish months ago: water, taxes, jobs, electricity prices, the environment and what they would do to small towns. Well-structured data center projects
Proud to see my co-founder, phd co-advisor and postdoc advisor Daniela Rus, on the TIME's 100 most influential people in AI.
Daniela is one of the defining figures of modern robotics and AI. it is a humbling experience to share a 9 years professional journey with her as a mentor, advisor, and business partner.
As she describes it herself about her work @MIT_CSAIL and @liquidai, it is "all about the science & engineering of intelligence, & especially embodied intelligence” ✨
time.com/collection/tim…
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
This one was long overdue in the local AI community! A benchmarking platform for on-device intelligence that is 1) independently validated, 2) community driven evaluation components, and 3) taking into account all aspects of local intelligence that do not show up on the cloud!
It was a pleasure working with the amazing @ArtificialAnlys team on this project.
✨ enjoy
Today we release Pipette, a model evaluation suite for on-device intelligence, in partnership with @ArtificialAnlys
Most benchmarking platforms are optimized to measure core capabilities and speed profile of foundation models served in the cloud. Pipette gives the field a
Which model should you run on the iPhone 17 Pro?
Announcing our new intelligence and inference testing for small models on mobile devices: independent measurement of how capable small models are in typical on-device tasks, and how they perform on popular phones - in partnership
Which model should you run on the iPhone 17 Pro?
Announcing our new intelligence and inference testing for small models on mobile devices: independent measurement of how capable small models are in typical on-device tasks, and how they perform on popular phones - in partnership with @liquidai
We have partnered with @liquidai to deliver mobile device inference benchmarking, covering a range of models in 4-bit or lower precision on the iPhone 17 Pro and Galaxy S26 Ultra
We’re publishing our phone-scale intelligence evaluation results in combination with @liquidai's inference performance benchmarks to give users and developers a holistic view of how small models are performing on phones
Inference benchmarking is conducted in a controlled environment using Liquid AI’s inference benchmarking software, which Artificial Analysis has examined and is open-sourced on Liquid AI’s GitHub. Results cover end-to-end generation time, output speed, peak memory usage and other metrics. The inference benchmarking app, ‘Pipette’, is available to download for free on iOS and Android, allowing users to test a variety of models on their own devices
We are ranking phone-scale model intelligence based on each model’s average score in five evaluations chosen for the task-based work these models do in practice: BFCL, IFBench, AA-Omniscience, GPQA Diamond and MATH-500. These evaluations are run by Artificial Analysis using our independent methodology. By default, we limit models to 16K context on each of these evaluations, representing the lack of memory space for significant KV cache on mobile devices. This leads to some intelligent but verbose models dipping in relative score - they were not designed for the constraints that phone memory imposes on token use
We are defining our portable device category as including models that fit within 8 GB of memory after quantization, including KV cache, at 8K context
We expect both the intelligence and inference benchmarks to evolve over time, as new models, devices, inference frameworks and quantization techniques are released. Our pages will remain up to date with these new additions
Initial results:
➤ Nanbeige4.2-3B and LFM2.5-2.6B share the top average evaluation score at 63 (with a 16K context limit), ahead of Ornith-1.0-9B at 62 and Qwen3.5 9B (Reasoning) at 61. LFM2.5-2.6B achieves its score more efficiently: on an iPhone 17 Pro it answers a standard 1,024-token prompt in 8.0s using 2.3 GB of memory, against 21.4s and 4.0 GB for Nanbeige4.2-3B and 25+ seconds and 6.9 GB for the two 9B models
➤ The 16K context limit shapes the leaderboard: as an example, Qwen3.5 9B (Reasoning) spends 74.5M output tokens across one pass of the benchmark set, hitting the 16K limit on 29% of its generations and landing in fourth place overall. With the limit raised to 64K, Ling 3.0 Tiny takes first place with a score of 66, ahead of Nanbeige4.2-3B (65), Qwen3.5 9B (Reasoning, 64), and LFM2.5 2.6B (64). But a 64K window does not fit in mobile phone memory, and at 55 output tokens/s on an iPhone, generating 64K tokens could mean a 20+ minute wait and a lot of battery use. This is why our primary results are capped at 16K, but we're also publishing a set of results capped at 64K, and another capped at one minute of generation time
➤ The speed-intelligence Pareto frontier is short: six models are unbeaten on both intelligence and speed on an iPhone 17 Pro: LFM2.5-230M (27 at 0.9s), MiniCPM5-1B (45 at 2.9s), LFM2.5-8B-A1B (58 at 5.7s), Ling 3.0 Tiny (59 at 5.7s), LFM2.5-2.6B (63 at 8.0s) and Nanbeige4.2-3B (63 at 21.4s). LFM2.5-8B-A1B and Ling 3.0 Tiny are mixture-of-experts models that activate ~1B parameters per token, which is how they answer in under 6s with 8B-class weights
➤ Leading models have opposite strengths: Nanbeige4.2-3B is the most balanced (76% on BFCL, 96% on MATH-500, 67% on GPQA Diamond); Qwen3.5 9B (Non-reasoning) is the strongest tool caller (77% on BFCL) and scientific reasoner (79% on GPQA Diamond); LFM2.5-2.6B follows instructions best of any model measured on the iPhone (59% on IFBench), clears 90% on MATH-500 and hallucinates far less on AA-Omniscience (79% non-hallucination, against 33% for Nanbeige4.2-3B and 24% for Qwen3.5 9B (Reasoning))
More details below in thread ⬇️
Today we release Pipette, a model evaluation suite for on-device intelligence, in partnership with @ArtificialAnlys
Most benchmarking platforms are optimized to measure core capabilities and speed profile of foundation models served in the cloud. Pipette gives the field a common, reproducible way to measure the quality, speed, latency, and memory use of AI models on devices such as phones, laptops, PCs, AI boxes, and embedded hardware.
> Pipette is open source
> In Pipette, models get compated as model + quantization + runtime + device from one interface.
> It comes with a warehouse of verified benchmark results, currently with 10k+ results across 35 model classes, 7 quants, llama.cpp runtimes, and 4 devices.
> Pipette is a dynamic platform, allowing new contributions from day one, adding new devices, runtimes, model families, and quantization levels.
🧵
the three personalities of digital writing
me in emails:
> Please review and let me know what questions you have
me in slack:
> ready for review. wdyt?
me to my agent:
> chkec the lofficla ocumentation
Today, we release updated 4-bit checkpoints for LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B trained with Quantization-Aware Distillation (QAD). These checkpoints recover accuracy lost to 4-bit quantization while retaining the low memory footprint and high decode throughput of the Q4_0 format.
All four checkpoints reach roughly 97% of their BF16 averages.
🧵
So you’re telling me I can swap my local LFM2.5-2.6B from F16 to QAD Q4_0 and go from:
5.4 GB → 1.6 GB
21 → 64 tok/s
3.0s → 1.2s tool-call latency
while keeping ~97% of BF16 performance?
@liquidai what did you just do 😭
yesterday we made them more compressed! today we make them faster than ever with speculative decoding!
up to 4x decode speed up on device for our 1.2B, 2.6B and 8B moe.
You gotta try these LFMs for function calling applications on device or latency critical load on the cloud! work of art by our very own @tugot17
enjoy 🚢🚀
Today, we release DSpark draft models for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B. These add a speculative decoding path that trades a minimal memory increase for a large decoding speedup without changing output quality.
A lightweight draft model proposes a block of
We gave LFM2.5-VL-3B from @liquidai a generated skin-like image and let it choose tools.
It mapped 6 regions, drew the contours, measured L04 at 8.8 × 8.1 mm, then picked what to review first.
Local on a Mac Studio. Visual review, not diagnosis. 🧵
@liquidai's LFM2.5-VL-3B, now native on Apple Silicon
The INT4/INT8 weights map to MLX bit-exactly (max diff 0.0) — then we chased the remaining bytes.
- Decode 12.16 → 62–68 tok/s (5.1–5.6×) · text TTFT 167 → 62 ms
- Checkpoint 2.79 → 1.63 GB (−42%) · peak RSS ~4.8 → ~1.9 GB (−60%)
- M3 decode is bandwidth-bound (~0.94 GB weights/token @ ~70 GB/s) — every win was a bandwidth cut, not a kernel trick
Go download it.
huggingface.co/konic-labs/LFM…
Sharing confidential information with AIs is a norm these days! We can do better. Our latest encoder models find and strip 40 types of PII across 16 languages, all in one simultaneous pass for a secure information sharing with AIs.
LFM2.5-Encoders identify all entities in the
141 Followers 110 FollowingThere is a horse in my heart, longing for freedom and the open wilderness. Here, I am simply my truest self. Thank you for passing by.
605K Followers 58K FollowingFan of new things and the people who build them | San Francisco/Silicon Valley AI | Ex-Microsoft, Rackspace, Fast Company | Wrote eight books about the future.
72K Followers 24K FollowingCustomize dashboard with multiple charts of different intervals and sentiment analysis with the help of twitter data.
#SmartViewAi #AI #Crypto
https://t.co/VYBrLM9Ji6
48K Followers 16 FollowingI've been in the industry for O(40) years and have written O(1M) LOC. I don't think I'll ever write O(another) line again, but I'll be launching more than ever.
11K Followers 173 FollowingChatGPT Consumer Product Lead
Former Head of Product @ https://t.co/h3TIy4Sbjx, product at https://t.co/AO9Z6nGQsT. New Yorker in SF.
3K Followers 2K FollowingBorn in the late 1900s, survived the early 2000s ✶ Christian ✶ Southerner ✶ EMS ✶ AI Enthusiast ✶ Independent Researcher ▶︎ •၊၊||၊|။||||။၊|• 0:10
1.1M Followers 63 FollowingIt's time to build.
https://t.co/A9eTFq6Xbx
Posts are not investment advice or an advertisement for investment services. See https://t.co/nX2FtaLE06.
473K Followers 9K FollowingFounder/CEO of Henry Intelligent Machines PBC and Creator Buddy ($300,000 ARR). Building a 100 trillion dollar economic engine
79K Followers 141 FollowingHave questions, or building something cool with Cloudflare's Developer products? We're here to help. For help with your account please try @CloudflareHelp
42K Followers 187 FollowingA North Star for open AGI. Co-founders: @fchollet @mikeknoop. President: @gregkamradt. We're hiring mission-driven builders: https://t.co/GswTSnyCoJ