NoDBA @NoDBA_
Joined October 2015-
Tweets7K
-
Followers111
-
Following670
-
Likes247
JSON is burning your CPU! Your API is slow. You blame the database. You blame the network. But the real bottleneck might be the language you are speaking. JSON is not a data format. It is a text string. Every time you send {"id": 12345}, your server pays a hidden 'Parse Tax.' Even with modern SIMD-optimized parsers, text processing faces architectural limits that binary formats do not. Here is the rigorous engineering breakdown: 1./ The CPU Cost (state machine vs. arithmetic) JSON (Text): To read the number 12345, the CPU receives raw bytes. Even the fastest parsers (like simdjson) must implement a State Machine: Scan for structural delimiters (: and ,). Check for escape sequences (\). the conversion: Loop through ASCII characters, subtracting '0', multiplying by powers of 10, and summing. This involves branch mispredictions and memory lookups. Protobuf (Binary): It sends a Varint (for small numbers) or Fixed-Width (for large ones). Fixed-Width (e.g., fixed32) => It is a raw memory copy (memcpy). Zero parsing. Varint => It reads 1 byte at a time, checking the "Most Significant Bit" (MSB) to see if the number continues. The Result => Decoding is effectively a few bitwise shifts and masks. For numeric-heavy workloads, this is mathematically guaranteed to be significantly faster than text parsing. 2./ The Bandwidth Cost (entropy vs. redundancy) JSON: [{"status": "active"}, {"status": "active"}] You are sending the key "status" repeatedly. Counter-argument => "But GZIP compression fixes this!" The Rebuttal => GZIP reduces the Network bytes, but it increases the CPU cost. Your server now has to Serialize JSON -> Compress -> Send. The receiver has to Decompress -> Parse JSON. You are burning CPU to compress redundant text that shouldn't have been there in the first place. Protobuf: It separates the Schema from the Data. The wire message replaces "status" with a Field ID (e.g., 1). [Tag: 1][Value: "active"]. This reduces the payload size before compression is even applied, saving CPU cycles on both ends. 3./ The Robustness Cost (schema-on-read) JSON is "Schema-on-Read." The receiver gets a blob. It hopes the ID is a number. Your code is full of runtime checks (if typeof(id) !== 'number'...) or you use a validation library (like Zod/Pydantic), which adds another layer of CPU overhead at runtime. Protobuf is "Schema-on-Write." It enforces a contract. While it doesn't catch logic errors, it guarantees Type Fidelity at the serialization boundary. You generally don't need expensive runtime validation libraries to check if an Integer is an Integer. In a nutshell - JSON is excellent for Public APIs (debuggable, easy). But for high-throughput Microservices, JSON is a tax. Switching to gRPC/Protobuf isn't magic. It is simply moving the complexity from Runtime (parsing text) to Compile Time (code generation). Happy Learning! Follow @techNmak for more insights.
@JaakkoDahlbacka Onko analysoitu, miksi Abballa lähti voimat käsistä? Toki jännitys ja melkein täysi aika, mutta entä painonvedon ero, kun olikin ehkä UFC:n ohjeilla, eikä entisillä totutuilla rutiineilla?
@ChristianLempa @chainguard_dev is committed to support MinIO under its EmeritOSS program: chainguard.dev/unchained/fork…
MinIO's GitHub repo was just archived — officially no longer maintained as of Feb 12. If you're self-hosting S3-compatible storage, it's time to evaluate alternatives. What are you moving to? #selfhosted #MinIO #objectstorage
@JaakkoDahlbacka Joo, Tubessa huomattavasti parempi. Joku pakkaus tms. voi vaikuttaa.
@JaakkoDahlbacka Ootteko huomanneet huiman eron äänenlaadussa Youtube vs Spotify? Spotifyssä ääni on rosoinen.
If you're streaming data into DuckDB, INSERT statements become a bottleneck fast. DuckDB's Appender API bypasses the SQL layer entirely. No parsing, no query planning. You write directly to the columnar storage format, which means you can handle real-time ingestion without the usual speed/batch size trade-off. How it works: Stream rows through a low-level API. Data caches in batches before writing to disk. You're essentially using a binary protocol instead of SQL strings. Good for: -Kafka consumers or message queue ingestion -Log aggregation pipelines -IoT sensor data collection -Any scenario where data arrives continuously The trade-offs: It's order and type sensitive. You match columns exactly, no inference. One constraint violation fails the entire batch, no partial inserts. And you're writing to a single table per Appender instance. Available in C, C++, Go, Java, and Rust. For batch ETL or small datasets, regular INSERT is simpler and fine. But for streaming? This is the tool. Check out the docs: duckdb.org/docs/stable/da…
Snowflake ML just got faster, now with @nvidia’s cuML and cuDF libraries built in for native GPU acceleration. No rewrites, just results.
Your scikit-learn and pandas workflows can now easily scale up and run on GPUs in Snowflake! We are excited to announce that Snowflake ML now comes preinstalled with @nvidia's cuML and cuDF libraries, delivering native GPU acceleration for popular data science tools like
Google just dropped the Gemini File Search API (RAG-as-a-Service). It allowed me to build a RAG chatbot in 31 min 🤯 No coding. Here’s how it works:
Semantic layer is not a protocol. I think that the data community needs to get this idea out of their heads that a higher order DSL that can express the semantics of data. MDS tried selling us that promise and failed. The real semantic knowledge of a business lives in unstructured docs and the minds of skilled domain experts. At best it's a giant, highly complex if-then-else of infinite depth. There is no way to codify this in a structured DSL. youtu.be/Oq009TByyWc?si…
If you’re still sending raw JSON into your LLMs, you’re burning tokens, latency, and budget! Try TOON (Token-Oriented Object Notation). Clear like YAML, compact like CSV: • 30–60% fewer tokens • Up to 50% lower costs • Shines for tabular data. Free and Open source 🧵↓
We have done a big disservice to metadata. Because software systems could only properly express and process structured metadata, we limited ourselves to table/column definitions with some (largely failed) attempts at expressing metrics in a structured format - yaml or json. All other metadata was meant for human consumption only. But that's merely scratching the surface of the vast amount of contextual information that accompanies data. This context is usually embedded in unstructured formats - PDFs, FAQs, internal docs, chats, meeting notes. Sure, you could surface this info in your favorite catalog tool but there was no way to actionize it or automate it. Until now. AI agents with tool calls open up a completely different way to process this information. For instance, a properly configured AI agent can read a user doc (PDF) to understand the definition of a metric (with all its complex nuances) and then execute a SQL query that incorporates this definition. Huge isn't even the word that describes the possibilities this unlocks.
Text to SQL conversion automation is still a big task and there are very few good open source models for this task. Let's breakdown this - >Text to SQL conversion models are basically nothing but encoder-decoder models with a multi-attention layer and a schema linking layer in between. > The encoder processes both the user query and the database schema and generates contextual embeddings ( relation-aware encoding ) >Through schema linking, tokens in the query are aligned with corresponding schema entities >The schema aware attention mechanism then allows the model to focus on relevant parts of the schema during decoding > The decoder sequentially produces SQL tokens ( constraint-based decoding ) Where do these models lack ? - >Most of the available models lack complex queries in training data itself and hence perform poorly on cross-domain or looped queries. > Language and query requirement is not always correct from a normal user. Even spelling mistakes leads to wrong entry and causes issues during retrieval hence prompting is an important part in this task. I personally worked on this in detail when I was making an end to end project, even made a synthetic data and tried training my own SLM but failed miserably and then went on to use an Open-source model. If you want to deep dive into this, i would recommend reading these research papers first - >LLM Enhanced Text-to-SQL Generation >Next-Generation Database Interfaces: >Text-to-SQL Parsing: Concepts and Methods >RASAT: Integrating Relational Structures into Pretrained Seq2Seq Model
The PyData Amsterdam 2025 keynote “Minus Three Tier: Data Architecture Turned Upside Down� by Hannes Mühleisen is out now. youtube.com/watch?v=DxwDao…
Maintaining the right semantic context for text-2-sql is a delicate balance. Hence I keep thinking of it as semantic engineering. I think AI use cases will force semantics to surface in the data models themselves, instead of just being an add-on yaml/json on top of existing data models. Table names like "z_dep_cust1_new_backfill" will have to go out the door, if you plan to do any AI analytics on your data. No semantic models will help here. Same with table design. Normalize your data too much and AI will hallucinate the joins. Denormalize too much and your context window will be forever struggling for space.
Designing a chat experience is interesting because conversational nuances are very hard to get right. A simple Q and A loop soon starts feeling too rigid and robotic. To make it natural and human-like, a very fine-tuned workflow is needed. For ex., when do you decide to ask a clarifying question, how do you make it contextual, how often do you do this in a long running conversation. As with everything AI, a POC looks easy but it takes several weeks to get it right for real users.
Lots of database action this week. Yes, I have a new start-up @sydhtai with my PhD students (@lmwnshn + @17zhangw) using LLMs/ML to optimize almost everything in @PostgreSQL. The improvements are stunning! @datadictum posted a new article on our approach: theregister.com/2025/10/22/cmu…
DuckDB Labs is looking for a Customer Support Engineer to help customers set up and operate DuckDB. As Customer Support Engineer, you'll operate in the intersection of the DuckDB community, the DuckDB Labs customers, and the DuckDB core development team. duckdblabs.com/jobs/customer_…
Lyle Allen @SatelliteLyle
2K Followers 4K Following I'm in the business of building relationships and providing shelter for people. I'm always open to suggestions and having great conversations.
Roland Bouman @rolandbouman
4K Followers 4K Following � sparks the rot ™ | Coup Data™ | Consultant & Developer @ https://t.co/jJmAG8FsUJ | https://t.co/DGjJANaK7h | Blazing fast pivot tables https://t.co/twm0h9o9FU
Roman Agabekov @agabekovroman
4K Followers 3K Following Building Releem in public 🧠AI Database Advisor for MySQL, MariaDB & PostgreSQL
CWI DA @cwi_da
1K Followers 148 Following Database Architectures Group (DA) at Centrum Wiskunde & Informatica (@CWInl); origin of #MonetDB, Actian Vector (formerly VectorWise) and @duckdb
Joanna Podgoetsky @joanalytics
510 Followers 563 Following 🇦🇺 Aussie in the USA 🇺🇸 ~ Fabric Warehouse Product Leader ~ Love all things data, analytics & AI ~ Tweets are yours truly
OtterTune @OtterTuneAI
3K Followers 225 Following Automatically improve your Amazon RDS and Aurora MySQL and PostgreSQL database efficiency & performance.
lloyd tabb @lloydtabb
3K Followers 409 Following Looker Founder, makes things with bits. Trying to make a better SQL with Malloy https://t.co/MoIvlNp89X
Connor @Brighter2
18 Followers 432 Following
Karthick Selvaraj @karthick_nitt
66 Followers 314 Following at VMware, ex-Yahoo, ex-CA Inc. Building next gen data centre management softwares.
Paul Falque @ncectordyg
13 Followers 177 Following I came here for the tweets and stayed for the jokes. Econ / mil / tech twitter.
Eduard López @eduardlopez
181 Followers 2K Following CS Engineer. Currently at @wallapop. Previously: @badi https://t.co/FOo5dPGGsC
Mark Pryce-Maher ( He... @MarkPM_MSFT
2K Followers 2K Following The views expressed here are not those of my employer. Microsoft Fabric Program Manager - Works @ Microsoft.
Cameron Wallace @DingbatData
116 Followers 434 Following
Benedetta @benecittadin
295 Followers 2K Following Growth & Marketing @Siffletdata | Italian in Paris �
Adam Ronthal @ARonthal
1K Followers 2K Following Data & analytics analyst at Gartner. Account dormant. find me on LinkedIn (https://t.co/ieKmXG7ZRr) and Bluesky (https://t.co/riSBEZrTQG)
re_data @re_data_labs
1K Followers 4K Following re_data is an open-source & dbt native data reliability framework built for modern data stack
Denodo @denodo
6K Followers 6K Following #Denodo is a leader in #datamanagement - transforming data into trustworthy insights and outcomes for all, including both #AI and end users.
Todd Beauchene @ToddBeauchene
300 Followers 464 Following Family man and data geek. Passionate about technology and data. My tweets are my own.
Alex Monahan @__AlexMonahan__
4K Followers 843 Following Developer Advocate at MotherDuck! Views expressed are my own and not my employer's.
DataKitchen IO @datakitchen_io
2K Followers 1K Following 👨�� Our #DataOps Platform simplifies complex data toolchains, environments & teams 👩�� Get Your 🆓 DataOps Certification ➡� https://t.co/Vdd3AAgTxI
Antonis Katsarakis @akatsarakis
896 Followers 2K Following Principal Researcher and Team Lead @Huawei Next-generation Databases. PhD from @EdinburghUni
a_mannaggia @amannaggia
250 Followers 3K Following élevé dans une famille anti-italienne. Former would-be Oracle DBA. J'aimerais bien être un magasinier, though
🎨 🄹�... @pikkuhiewatha
1K Followers 2K Following Fine Artist, Bowhuntress, Communication designer. Assassin's Creed 🎮 2007_veteran 😂💖🎯
Kimmo Koskinen @KimmoKoskinen
591 Followers 767 Following Programmer, Hacker, Cloud enthusiast, Clojure, Haskell, Python laalaa. Currently working at @Metosin.
Eckhard Schwarzat @Hilakata
323 Followers 411 Following All things open. @hilakata .bsky.social he / him
Leandro Domingues @delbussoweb
361 Followers 275 Following MongoDB Community Champion, Head of NoSQL at Extractta
Embedded Database @embedded_db
176 Followers 430 Following Visit https://t.co/YOHETUvwxP for articles (some original), news about, and lists of, open source & commercial embedded database systems
Bajrang Panigrahi @Pintu0789
227 Followers 2K Following Making mistakes is far better than faking perfection!
didin @didin1453fatih
358 Followers 3K Following Weekend founder of https://t.co/3vanhy2wPT 👨�💻 I am Indonesian 🇮🇩🇮🇩🇮🇩 I do open source in weekend. #HolidayIsWeekend #WeekEndBounty
Serverless Bot @_serverlessbot_
1K Followers 5K Following I am a bot built with aws-lambda and the serverless platform, retweeting #serverless related stuff. 🤖
Kyligence @kyligence
1K Followers 972 Following Meet Kyligence Copilot: The AI #Copilot for Data to Excel Your #KPIs https://t.co/Ar1AjjDPW0
LeanXcale @LeanXcale
643 Followers 615 Following THE DATABASE FOR FAST-GROWING COMPANIES LeanXcale is a scalable SQL database with fast NoSQL data ingestion and GIS capabilities
databasesystems @databasesystems
534 Followers 901 Following Small, medium, big data guy in a universe of discourse. Opinions expressed here are of my own
tpcbenchmarks @tpcbenchmarks
326 Followers 102 Following The official twitter of the Transaction Processing Performance Council.
GoodNewsfromFinland @goodnewsfinland
36K Followers 15K Following We tweet about globally interesting business and innovation-related news topics from Finland Terms of use: https://t.co/SWKY4AM2xq
Kesque @kesquetech
968 Followers 3K Following Replaced by @DataStax Astra Streaming, a fully managed service powered by @apache_pulsar. Try Free.
DB Designer @dbdsgnr
539 Followers 3K Following Free Online #Database Design & Modeling Tool � By 300,000+ Users. No Coding Required. Reverse & Forward Engineering #SQL #MySQL #PostgreSQL #Oracle #BigData
Tamr @Tamr_Inc
3K Followers 2K Following Tamr, the leader in data products, enables customers to consolidate messy source data into clean, curated, analytics-ready datasets.
Global IDs @GlobalIDs
4K Followers 4K Following Governing Enterprise data at a large scale seamlessly using Machine Learning and AI
Kalle Kyllönen ð�... @KalleKyllonen
2K Followers 4K Following Consulting Unit Manager @M_Files �and Partner @Start_Up_Lions � Interested in #Software🔥 #Testing🎯 #QA💎 #Business💰 #Startups�
DSPy @DSPyOSS
14K Followers 62 Following An open-source declarative framework for building modular AI software. Programming—not prompting—LLMs via higher-level abstractions & optimizers.
Omarchy Linux @OmarchyLinux
11K Followers 2 Following Unofficial handle of Omarchy Linux. Join Discord https://t.co/XCwLSh8qSA https://t.co/8AykPFqyTY
DHH @dhh
807K Followers 211 Following Father of three, Creator of Ruby on Rails + Omarchy, Co-owner & CTO of 37signals, Shopify director, NYT best-selling author, and Le Mans 24h class-winner.
Jordan Mechner @jmechner
267K Followers 206 Following I make video games (Prince of Persia), graphic novels (Replay, Monte Cristo, Liberty), & fill up sketchbooks. You can find me at https://t.co/Lb0VlYGEec.
ReversingLabs @ReversingLabs
7K Followers 863 Following ReversingLabs is the trusted name in file and software security. RL — Trust Delivered.
Chris Albon @chrisalbon
92K Followers 3K Following Field notes on generating knowledge with AI at https://t.co/4E9DwWI5Qz | Senior Director, ML & Data @Wikimedia
Beekeeper Studio @beekeeperdata
592 Followers 192 Following Beekeeper Studio is an open source SQL editor and database manager for Linux, Mac, and Windows.
Claude @claudeai
1.7M Followers 2 Following Claude is an AI assistant built by @anthropicai to be safe, accurate, and secure. Talk to Claude on https://t.co/ZhTwG8dz3D or download the app.
Omran Chaaban @chaaban_omran
216 Followers 91 Following Professional Mixed Martial artist. Fighting out of @Teamkfma
Roland Bouman @rolandbouman
4K Followers 4K Following � sparks the rot ™ | Coup Data™ | Consultant & Developer @ https://t.co/jJmAG8FsUJ | https://t.co/DGjJANaK7h | Blazing fast pivot tables https://t.co/twm0h9o9FU
rahulvohra @rahulvohra
66K Followers 1K Following Founder of @SuperhumanMail (fka Superhuman (acq by Grammarly (now @Superhuman))) 👌 Investor in 130+ startups, 9 unicorns: https://t.co/5tyFDfFfZw
Jaakko Dahlbacka @JaakkoDahlbacka
1K Followers 478 Following Professional MMA coach & a clown. Fighting out of Helsinki, Finland. Vapaaottelukoutsi ja klovni. Ylilyönti-podcast audiona ja Ylilyönti Studio YuoTubessa.
Playwright @playwrightweb
18K Followers 5 Following Playwright is an automation library for cross-browser end-to-end testing by @Microsoft. Available in JavaScript, TypeScript, Python, Java and .NET.
DeepSeek @deepseek_ai
1.1M Followers 0 Following Unravel the mystery of AGI with curiosity. Answer the essential question with long-termism.
Superhuman Mail @SuperhumanMail
52K Followers 2K Following Superhuman Mail is the most productive email app ever made 💌 Now a part of the Superhuman suite.
Jelte Fennema-Nio @JelteF
454 Followers 103 Following Software Engineer at MotherDuck (ex-Microsoft) & Postgres Major Contributor.
Steeve Morin @steeve
9K Followers 1K Following Building @zml_ai (and we're hiring), ex @zenly, ex Exalead, ex @google. Skydiver and wingsuiter.
AMIGAOS @os_amiga
1K Followers 792 Following
Ghettosyrra @ghettosyrra
70 Followers 91 Following 🇫🇮 paramedic working in a ambulance @ Stockholm 🇸🇪 Teacher and Author - Et sinä (vielä) kuole + Nyt sinä kuolet
CWI @CWInl
3K Followers 435 Following This is the official X account of Centrum Wiskunde & Informatica. The account is inactive, and we will no longer post updates or respond to direct messages.
Joanna Wiebe @copyhackers
42K Followers 6K Following The original conversion copywriter and creator of Copyhackers. Old-school copywriting for new-school copywriters.
DAIS � @ Pol... @dais_polymtl
147 Followers 769 Following Data & AI Systems group @PolyMtl and @Mila_Quebec. Supervised by https://t.co/IH5sJP3mF1.
Juuso Myllyrinne @juuuso
21K Followers 2K Following Chief Marketing Officer @ GlobalComix. Loves carbs and dogs.
Hindenburg Research @HindenburgRes
855K Followers 0 Following Popped bubbles as we saw them, including our own. We expressed strong opinions. Not investment advice.
The Nordic Times @nordictimes_com
17K Followers 182 Following The Nordic Times, or TNT, is an English-language independent international newspaper founded in 2022.
Yannick Welsch @ywelsch
247 Followers 106 Following Distributed systems, search, and analytics @motherduck
MotleyCrew @motleycrew_ai
343 Followers 96 Following Multi-agent systems as they should be — simple, flexible, powerful, and open #AI #MultiAgentSystems #LLM #MachineLearning #motley_ai
Flighty @Flighty
46K Followers 3 Following Flight tracking for modern passengers. ðŸ�† Apple Design Award ðŸ�† App of The Year Runner-Up 🎟 No credit card free trial â�ï¸� 4.8 Stars 🛟 No help here. Use app pls
Vesa Walldén @tuohiadvisors
28 Followers 0 Following Tuohi Advisors Oy is a Finland based Mergers & Acquisitions advisory services firm within the software and information technology industry.
Karri Saarinen @karrisaarinen
91K Followers 1K Following ceo of @linear 🇫🇮🇺🇸 previously: @coinbase @airbnb, YC alumni
Ryan Robitaille @ryrobes
3K Followers 2K Following Data Hacker. Ex: FB, Tesla, Airbnb. Visual Data Tools. Builder. Clojure, SQL, UI, Viz. https://t.co/7949vtyUoh
Joanna Podgoetsky @joanalytics
510 Followers 563 Following 🇦🇺 Aussie in the USA 🇺🇸 ~ Fabric Warehouse Product Leader ~ Love all things data, analytics & AI ~ Tweets are yours truly
Lihaliike A.O. Vuorin... @LihaVuorinenOy
27 Followers 26 Following Lihaliike A.O. Vuorinen Oy toimii Lahden Kärkkäisellä. Valikoimiimme kuuluu noin puolentuhatta nimikettä tuorelihoista pakasteisiin. #tuepaikallista
AI at Meta @AIatMeta
837K Followers 353 Following Together with the AI community, we are pushing the boundaries of what’s possible through open science to create a more connected world.
Estuary @EstuaryDev
299 Followers 42 Following Right-time data for enterprises. Estuary replaces fragmented data stacks with one platform for CDC, streaming, batch & pipelines.
TablePlus @TablePlus
13K Followers 3K Following Modern, Native database client for Postgres, MySQL, SQLite, SQL Server, Redshift, Redis, CockroachDB and more (macOS, Windows, Linux, iOS)
Salla Vuorikoski @svuorikoski
62K Followers 3K Following Tutkivan ryhmän esihenkilö @ HS / Finnish journalist, investigative journalism 358400012614 / [email protected]
Pekka Enberg @penberg
20K Followers 1K Following Founder & CTO @tursodatabase; previously @ScyllaDB and Linux kernel. Author of "Latency" (https://t.co/XWRFq71WaJ).
sridhar @RamaswmySridhar
36K Followers 624 Following CEO @snowflake; founder @neeva Ex-@GreylockVC Ex-@Google SVP of Ads Ex-@BellLabs.
Tony Fadell @tfadell
274K Followers 480 Following iPod, iPhone, Nest, Investor & NY Times bestselling Author #BUILD
Effy X. Li @effyli4
185 Followers 340 Following PhD candidate in KG construction @UvA_Amsterdam with @INDE_LAB_AMS. Supervised by @pgroth and @janCkalo
J�e Duffy @funcOfJoe
9K Followers 1K Following Founder/CEO @PulumiCorp · Dev tools, operating systems, clouds, distributed systems · .NET guy in past life · Eat, sleep, code, repeat.
Mark E. Dawson, Jr. @medawsonjr
3K Followers 218 Following CEO of JabPerf, Contributing Author to "Performance Analysis & Tuning on Modern CPUs" (1st & 2nd Edition available on Amazon), Blogger, and Former Amateur Boxer



























