cornelistools @cornelistools
Building advanced language technologies and robust AI systems for the Dutch language 🇳🇱 Research & development in #nlp #nlu #ai #rl #conversationalAI cornelistools.nl Gooise Meren (NL) Joined October 2020-
Tweets172
-
Followers49
-
Following395
-
Likes568
Effectiever trainen van Nederlandstalige taalmodellen door onderzoek in niet-opknipbare zelfstandig naamwoorden en het verminderen van ambiguïteit in tokenizers. cornelistools.nl/blog/2025/10/e… #LLMs #AI #tokenizers #linguistics #NLP
Again an amazing @MLStreetTalk with the great @LauraRuis about reasoning and agency in AI Models and her paper "Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models". How Do AI Models Actually Think? youtube.com/watch?v=14DXtv…
What's really going on in machine learning? Just finished a deep dive using (new) minimal models. Seems like ML is basically about fitting together lumps of computational irreducibility ... with important potential implications for science of ML, and future tech... writings.stephenwolfram.com/2024/08/whats-…
# RLHF is just barely RL Reinforcement Learning from Human Feedback (RLHF) is the third (and last) major stage of training an LLM, after pretraining and supervised finetuning (SFT). My rant on RLHF is that it is just barely RL, in a way that I think is not too widely appreciated. RL is powerful. RLHF is not. Let's take a look at the example of AlphaGo. AlphaGo was trained with actual RL. The computer played games of Go and trained on rollouts that maximized the reward function (winning the game), eventually surpassing the best human players at Go. AlphaGo was not trained with RLHF. If it were, it would not have worked nearly as well. What would it look like to train AlphaGo with RLHF? Well first, you'd give human labelers two board states from Go, and ask them which one they like better: Then you'd collect say 100,000 comparisons like this, and you'd train a "Reward Model" (RM) neural network to imitate this human "vibe check" of the board state. You'd train it to agree with the human judgement on average. Once we have a Reward Model vibe check, you run RL with respect to it, learning to play the moves that lead to good vibes. Clearly, this would not have led anywhere too interesting in Go. There are two fundamental, separate reasons for this: 1. The vibes could be misleading - this is not the actual reward (winning the game). This is a crappy proxy objective. But much worse, 2. You'd find that your RL optimization goes off rails as it quickly discovers board states that are adversarial examples to the Reward Model. Remember the RM is a massive neural net with billions of parameters imitating the vibe. There are board states are "out of distribution" to its training data, which are not actually good states, yet by chance they get a very high reward from the RM. For the exact same reasons, sometimes I'm a bit surprised RLHF works for LLMs at all. The RM we train for LLMs is just a vibe check in the exact same way. It gives high scores to the kinds of assistant responses that human raters statistically seem to like. It's not the "actual" objective of correctly solving problems, it's a proxy objective of what looks good to humans. Second, you can't even run RLHF for too long because your model quickly learns to respond in ways that game the reward model. These predictions can look really weird, e.g. you'll see that your LLM Assistant starts to respond with something non-sensical like "The the the the the the" to many prompts. Which looks ridiculous to you but then you look at the RM vibe check and see that for some reason the RM thinks these look excellent. Your LLM found an adversarial example. It's out of domain w.r.t. the RM's training data, in an undefined territory. Yes you can mitigate this by repeatedly adding these specific examples into the training set, but you'll find other adversarial examples next time around. For this reason, you can't even run RLHF for too many steps of optimization. You do a few hundred/thousand steps and then you have to call it because your optimization will start to game the RM. This is not RL like AlphaGo was. And yet, RLHF is a net helpful step of building an LLM Assistant. I think there's a few subtle reasons but my favorite one to point to is that through it, the LLM Assistant benefits from the generator-discriminator gap. That is, for many problem types, it is a significantly easier task for a human labeler to select the best of few candidate answers, instead of writing the ideal answer from scratch. A good example is a prompt like "Generate a poem about paperclips" or something like that. An average human labeler will struggle to write a good poem from scratch as an SFT example, but they could select a good looking poem given a few candidates. So RLHF is a kind of way to benefit from this gap of "easiness" of human supervision. There's a few other reasons, e.g. RLHF is also helpful in mitigating hallucinations because if the RM is a strong enough model to catch the LLM making stuff up during training, it can learn to penalize this with a low reward, teaching the model an aversion to risking factual knowledge when it's not sure. But a satisfying treatment of hallucinations and their mitigations is a whole different post so I digress. All to say that RLHF *is* net useful, but it's not RL. No production-grade *actual* RL on an LLM has so far been convincingly achieved and demonstrated in an open domain, at scale. And intuitively, this is because getting actual rewards (i.e. the equivalent of win the game) is really difficult in the open-ended problem solving tasks. It's all fun and games in a closed, game-like environment like Go where the dynamics are constrained and the reward function is cheap to evaluate and impossible to game. But how do you give an objective reward for summarizing an article? Or answering a slightly ambiguous question about some pip install issue? Or telling a joke? Or re-writing some Java code to Python? Going towards this is not in principle impossible but it's also not trivial and it requires some creative thinking. But whoever convincingly cracks this problem will be able to run actual RL. The kind of RL that led to AlphaGo beating humans in Go. Except this LLM would have a real shot of beating humans in open-domain problem solving.
🥁 Llama3 is out 🥁 8B and 70B models available today. 8k context length. Trained with 15 trillion tokens on a custom-built 24k GPU cluster. Great performance on various benchmarks, with Llam3-8B doing better than Llama2-70B in some cases. More versions are coming over the next few months. llama.meta.com/llama3/
At least once a year I come across the argument that "scale is all you need: the more neurons a species has, the more intelligent it is; humans have over 2x more neurons than gorillas and that makes all the difference; future AIs will have even more neurons than us! If we give them 1,000x more neurons they will be 1,000x more intelligent!" As a reminder, the species with the most neurons are whales and African elephants (a whopping 3x more than us). And for any particular neuron count, you will find species with wildly different levels of cognitive ability.
@fchollet I agree. The strength of the Stack Overflow is still a community with diversity, vision and experiences of many programmers around the world. And one answer is often not an answer. In coding the why is at least important as the how. LLMs can't provide that vision yet.
We’re releasing Gemma 2B and 7B, which achieve best-in-class performance for their sizes compared to other models, and can run on a developer laptop or computer. They also surpass much larger models on key benchmarks while meeting our standards for safe and responsible outputs.
@pkuhar @GaryMarcus Indeed, with RAG you limit the allowed "knowlegde" of the model to "talk" about only the retrieved results provided in a prompt. As the retrieval system is basically search (or "semantic" or whatever search), it has many limitations on reasoning, nor sense of the total dataset.
@GaryMarcus I mean, if the expectation is a talking search engine, it works "ok", but mainly defined by the quality of the external retrieval system. If the expectation is reasoning about the question and results, the RAG method is useless. LLMs do nothing very smart here & even can be SLMs.
@karpathy I tried managing my schedule with an LLM and had the same result 😉
@GaryMarcus As you need to insert with RAG your company's knowlegde externally into the LLM, most of the preliminary work is done in the R-mechanism (retrieval) outside the LLM. Using a LLM to generate pre-presented data is actually a waste of the knowlegde it has. SLMs work fine here too.
@christineliebr @KvW Ik gebruik deze ook en denk dat deze aan de goede kant zit. Deze kristalliseert (voornamelijk eigenschap van kwalitatieve honing volgens volkstuin.woltersweb.nl/waarom-echte-h…), komt uit bloemen en echte honing is idd niet geschikt voor heel jonge kinderen.
Today, we’re launching Aya, a new open-source, massively multilingual LLM & dataset to help support under-represented languages. Aya outperforms existing open-source models and covers 101 different languages – more than double covered by previous models. cohere.com/research/aya
@ohmypy And for many scripting languages you don't really need a configuration language at all. Having a config .py or .php file is not very wrong. However, to separate code and configuration language completely or having a compiled language, JSON would be fine and a well-known standard.
@Justin_Halford_ @fchollet If the task you want to solve is more beneficial than the energy consumption, it may not. However, knowing spending a 500ml bottle of water when I have a conversation with ChatGPT, I would like to see more earth friendly solutions. There is a lot that needs to be improved there.
@fchollet Also: Human: "Let's eat 2 slices of bread with peanut butter for the morning" Tesla bot: "The production version of the Optimus bot will be equipped with a 2.3-kilowatt battery pack"
@fchollet 12 watts is about 2x Raspberry Pi 4B's at full load 💪 I love to see people already experimenting with some 7B GGUF models on PIs, and I hope to see way more focus on energy efficient models.
Mijn gevoel dat Chain-of-Thought prompting vooral werkt is omdat het deels "Chain-of-Completion" is. LLMs zijn getraind voor completion (next token prediction) en door een redenatie vraagstuk stap voor stap uit te schrijven, word redenatie een behapbaarder completion vraagstuk.
AIMagic @Real_AImagic
9 Followers 243 Following AIMagic is a Dutch agency - blending AI, strategy and creativity.
Simme volkers @simonsayscya
64 Followers 178 Following seo copywriter plus kiteboarder on sea #seocopywriter
Corki @Corkiis9p
47 Followers 5K Following
Thescrs @Thescrsmp7y
48 Followers 4K Following
Peter Kuhar @pkuhar
2K Followers 2K Following Entrepreneur, coder, maker. Mobile Health and Machine Learning. https://t.co/hjYo5bICZU
anjadec @AnjadeCastro
33 Followers 196 Following
Clienten Plein @TEL1NL
16 Followers 142 Following https://t.co/PO3ZAac63G webio user interest view unity.
Frans Olsthoorn @FransOlst
41 Followers 270 Following CEO at Scriptix: making content accessible to everybody
𝚑𝚎𝚗𝚔 𝚟... @henkvaness
55K Followers 9K Following Cutting through #AI for sharper investigations. Workshops worldwide. Trusted by Pulitzer winners, law makers and NGOs. Mediadetector https://t.co/A7pkSs4VDY
Erik Meijerink 🌍 @erikotn
656 Followers 437 Following concert photographer / pr @ podium https://t.co/biCUkV1Glo | landscape photographer @ https://t.co/wgWa2EQrSL | creative strategist @ https://t.co/ob2LjF0oJM
Pieter van de End @Pieter_vd_Ende
127 Followers 3K Following studies MSc. Health policy, Law and Innovation - defend the free West 🇳🇱🇺🇲🇮🇱🇺🇦
SkyWalker - customer ... @SkyWalker_nl
55 Followers 2K Following SkyWalker - customer service management is dé specialist in recruitment, executive search, consultancy en interim management voor customer service
Peter Cornelissen @PajCornelissen
0 Followers 2 Following
De Brominocs - Voorle... @joinprettypinky
3K Followers 4K Following Volg ons op YouTube voor spannende avonturen! Nederlandstalige voorlees- / luisterverhalen van de Brominocs.
Nathan Benaich @nathanbenaich
71K Followers 36K Following solo member of superinvestment staff @airstreet @airstreetpress @stateofai @raais
Jane Cotterell @jane_cotterell
2K Followers 5K Following
Laurens Vreekamp @campodipace
1K Followers 845 Following 📖 Boek 'The Art of AI' ✺ Future Journalism Today 🎓 Mentor @SVDJnieuws 📨 Super Vision, nieuwsbrief over AI voor creatieven 🔍 Ex-Googler
Loy @loybeek
1K Followers 3K Following Dad, Robotics at https://t.co/1rs6ck77aG & @TechUnited RoboCup@Home, scout @ScoutingBoxtel. Interested layman in anything. Mastodon: @[email protected]
Paul Molenaar @pmolenaar
836 Followers 308 Following
gert koot @gertkoot
2K Followers 2K Following Co-Founder HumanTales. Consultant. Advertising. Marketing. Reclame. New Media. Teacher. Speaker. Storyteller.
Tech Today @tech_vandaag
505 Followers 1K Following Dutch and English. 🇳🇱 🇬🇧 Volg ons voor je dagelijkse portie tech. Van Apple tot Zigbee en van Google tot Tiktok.
Jesse Wienholts @JesseWienholts
491 Followers 988 Following [email protected] | digital accessibility and innovation
Wluper @wluper_
330 Followers 674 Following Advanced Conversational AI to provide powerful voice-first experiences for any industry.
ARxVision @ARxVision
419 Followers 1K Following ARxVision makes the ARxAI, the wearable AI device compatible with smartphones apps and makes the world more accessible through audio generative AI .
European Language Tec... @EuroLangTech
779 Followers 928 Following The channel of the EU projects European Language Equality and European Language Grid, fostering the LT community towards digital language equality by 2030!
Martijn Eindhoven @MEindhoven
393 Followers 163 Following
Barend Jungerius 🤖 @BarendJun
1K Followers 2K Following Freelance Consultant - #ConversationalAI #GenerativeAI. Founder of @Botrepreneurs networking events 🤖
M. Lens-FitzGerald @Dutchcowboy
6K Followers 5K Following Innovation Executive focused on government, emergence, people, teams and change. I instigate movements that shape the future.
Lee Boonstra 🏳️�... @leeboonstra
2K Followers 789 Following Software Engineer #SWE for the Office of the CTO @ #Google, #Innovation #ConversationalAI, O'Reilly & Apress book author, public speaker & LGBTQ mom.
Charlotte van Hooijdo... @vhooijd
230 Followers 319 Following Ass. prof. Language & Communication Utrecht University| Webcare & Chatbots | Health Communication | Visual Metaphor | Advisor & Speaker|
Botrepreneurs @botrepreneurs
94 Followers 136 Following Community & Online+IRL Networking Events about #ConversationalAI, #Chatbots & #Voice. Tweets: @BarendJun
bart stroeken @bartstroeken
181 Followers 395 Following developer | singer | guitarist | culture-addict
VUIchallenge @VUIchallenge
291 Followers 1K Following Daily and free challenges for you to improve your VUI skills. Subscribe: https://t.co/XN5ri5RphC #vuichallenge
Incentro @Incentro_
1K Followers 787 Following As an digital agency we create digital happiness. We believe in happy people; our main KPI | Best Workplace 2018
pc @pc78412925
2 Followers 9 Following
Gerard van Nieuwenhui... @GerarDOSKVK
628 Followers 360 Following Ooit hoofdredacteur van https://t.co/osaT5uXwpp. 100% succesvolle fleets. Co-host van @animegesprek
EU Chatbot & Conversa... @EuropeanChatbot
614 Followers 782 Following Connecting you with customers & Accelerating the adoption of CAI in the EU market. #conversationalAI #voice #chatbots #virtualassitant #generativeai #euchatbot
ReadSpeaker @ReadSpeaker
2K Followers 1K Following Empowering Accessibility and Inclusion with Text-To-Speech Solutions. Revolutionizing the way we listen, engage and learn 🎧
Tim Coronel @TimCoronel
922K Followers 113K Following #Dakar2026 #DakarInSaudi #DakarRally 20th time 🐫🐪🐫🐪🐫 #rally CoronelDakar is a marketing team that happens 2 race. Fun is our motto
Novum @novum_nu
183 Followers 552 Following Innovatielab van de Sociale Verzekeringsbank, gericht op het verbeteren van de dienstverlening. Snelheid is het credo en de burger staat centraal. #innovatie
Jeroen Vonk @jeroenvonk_
2K Followers 3K Following AI Enthusiast | Innovation Designer | Creative Problemsolver | Wine Writer at https://t.co/NwDkAdQyRE
Danny Uyterlinde @dannyadamse
235 Followers 557 Following Papa van Jet & Teun | man van Roos | MSc (UvA) | Senior Developer @ Endeavour | Founder @ DSJ Digital, Store24, NetworkingDay
PC @Peter20107473
2 Followers 14 Following
Kimi.ai @Kimi_Moonshot
346K Followers 137 Following Built by Moonshot AI to empower everyone to be superhuman. @Kimidevs built for developers ⚡️API: https://t.co/XCrgjXAqMw DC: https://t.co/wBsBTn6Ncc
Dave W Plummer @davepl1968
105K Followers 86 Following Hi! I'm Dave Plummer. You might remember me from such Windows components as Task Manager, Windows Pinball, Calc, ZIPFolders, Product Activation, etc. Cheers!
Thinking Machines @thinkymachines
179K Followers 1 Following Thinking, beeping, and booping. @tinkerapi
Shane Legg @ShaneLegg
81K Followers 67 Following Chief AGI Scientist & Co-Founder, Google DeepMind Work website: https://t.co/E4SyeGVYXk Personal blog: https://t.co/LL9JNdNpW1
Jeff Geerling @geerlingguy
94K Followers 5K Following Father, author, developer, maker. Sometimes called "an inflammatory enigma". #stl #ansible #k8s #raspberrypi #crohns #ostomy
Dwarkesh Patel @dwarkesh_sp
246K Followers 1K Following Host of @dwarkeshpodcast https://t.co/3SXlu7fy6N https://t.co/4DPAxODFYi https://t.co/hQfIWdM1Un
DHH @dhh
773K Followers 203 Following Father of three, Creator of Ruby on Rails + Omarchy, Co-owner & CTO of 37signals, Shopify director, NYT best-selling author, and Le Mans 24h class-winner.
Godot Engine @godotengine
143K Followers 3 Following Your free, open-source game engine 🎮🛠️ Develop your 2D & 3D games, cross-platform projects, or XR ideas! https://t.co/RlOWFgRKhP
LÖVE @obey_love
4K Followers 5 Following A 2D game engine that allows for rapid development using Lua.
LaurieWired @lauriewired
158K Followers 294 Following researcher @google; serial complexity unpacker; https://t.co/Vl1seeNgYK ex @ msft & aerospace
Heroes Dutch Comic Co... @dutchcomiccon
7K Followers 543 Following November 21 & 22 | 2026 | The Netherlands' biggest pop culture event 🧚
Simon Willison @simonw
201K Followers 6K Following Creator @datasetteproj, co-creator Django. PSF board. Hangs out with @natbat. He/Him. Mastodon: https://t.co/t0MrmnJW0K Bsky: https://t.co/OnWIyhX4CH
DuckDB @duckdb
25K Followers 62 Following DuckDB is an analytical SQL database management system. "DuckDB" and the DuckDB logo are registered trademarks of the DuckDB Foundation.
Michael Levin @drmichaellevin
79K Followers 3K Following Scientist at Tufts University; my lab studies anatomical and behavioral decision-making at multiple scales of biological, artificial, and hybrid systems.
Harrison Kinsley @Sentdex
109K Followers 451 Following gpus and tractors. Director of AI and Engineering @ https://t.co/H4St8dd1ip Neural networks from Scratch book: https://t.co/hyMkWyUP7R https://t.co/8WGZRkUGsn
DeepSeek @deepseek_ai
1.1M Followers 0 Following Unravel the mystery of AGI with curiosity. Answer the essential question with long-termism.
ClickHouse @ClickHouseDB
19K Followers 63 Following ClickHouse is the fastest open-source OLAP database ⚡ Download: https://t.co/3JKlDJbkcH GitHub: https://t.co/bjCe9qIetg Slack: https://t.co/d95c6jVeJm
Verna Dankers @vernadankers
1K Followers 337 Following Postdoc @Mila_Quebec @McGill_NLP 🇨🇦 PhD from @Edin_CDT_NLP 🏴 memorization vs generalization x (non-)compositionality x interpretability. she/her
Renate🇳🇴 @Renate_FE
12K Followers 350 Following Interest: UAP / UFO enigma. President of the org. UFO-Norway. Engineer. Project Director technical entrepreneur and 🪖 Be kind 🙂 https://t.co/Tf3FDbrYtn
Richard Sutton @RichardSSutton
69K Followers 59 Following Student of mind and nature, libertarian, chess player, cancer survivor. @ https://t.co/pioQVVjSXz, UAlberta, amii, https://t.co/vPRUv44glx, The Royal Society, Turing Award
Neel Nanda @NeelNanda5
42K Followers 122 Following Mechanistic Interpretability lead DeepMind. Formerly @AnthropicAI, independent. In this to reduce AI X-risk. Neural networks can be understood, let's go do it!
Jürgen Schmidhuber @SchmidhuberAI
209K Followers 0 Following OG of: P and T in ChatGPT, 100x deeper learning, meta learning and RSI, neural distillation, GAN/World Model... Co-authored most-cited AI paper of 20th century
Interconnects @interconnectsai
9K Followers 2 Following What you need to know about AI research trends, from @natolambert Wednesday mornings weekly, sometimes extra posts.
Fei-Fei Li @drfeifei
915K Followers 1K Following Cofounder/CEO @theworldlabs, Prof (CS @Stanford), Co-Director @StanfordHAI, #AI #SpatialIntelligence #GenAI #computervision #robotics #AI-healthcare
Stephen Wolfram @stephen_wolfram
182K Followers 4 Following Creating ideas, technology, science, companies, books, ... #WolfLang #WolframPhysics #WolframAlpha #Mathematica @WolframResearch
Sabine Hossenfelder @skdh
225K Followers 809 Following German Physicist. Author of "Lost in Math" & "Existential Physics". There is no strength in numbers, have no such misconception. rt's are not endorsements
kyutai @kyutai_labs
27K Followers 14 Following
Sepp Hochreiter @HochreiterSepp
15K Followers 372 Following Pioneer of Deep Learning and known for vanishing gradient and the LSTM.
Isomorphic Labs @IsomorphicLabs
59K Followers 82 Following Solve all disease. Developing and applying frontier AI to unlock deeper scientific insights, faster breakthroughs, and life-changing medicines.
Cleo Abram @cleoabram
80K Followers 982 Following Video journalist making optimistic science and tech explainers. HUGE* If True. Watch: https://t.co/EI32Qgtigc
Mira Murati @miramurati
841K Followers 638 Following Now building @thinkymachines. Previously CTO @OpenAI
Phillip Lippe @phillip_lippe
3K Followers 535 Following Research Scientist @GoogleDeepMind, RL/Multimodal/🍌 Gemini | Prev @UvA_Amsterdam - Opinions my own
Jeff Dean @JeffDean
451K Followers 6K Following Chief Scientist, Google DeepMind & Google Research. Gemini Lead. Opinions stated here are my own, not those of Google. TensorFlow, MapReduce, Bigtable, ...
Cohere Labs @Cohere_Labs
27K Followers 268 Following @Cohere's research lab and open science initiative that seeks to solve complex machine learning problems. Join us in exploring the unknown, together.
Lamini @LaminiAI
6K Followers 9 Following The LLM tuning & inference platform for enterprises. Factual LLMs. Deployed anywhere.
the tiny corp @__tinygrad__
79K Followers 199 Following We make tinygrad; sell tinybox for the GPU middle class. Our mission is to commoditize the petaflop.
LlamaIndex 🦙 @llama_index
118K Followers 33 Following The most accurate agentic OCR platform for production AI. LlamaParse: https://t.co/yQGTiRSNvj Docs: https://t.co/us6GCS1Clb
BlinkDL @BlinkDL_AI
10K Followers 167 Following RWKV = 100% RNN with GPT-level performance. https://t.co/TkdxOJSFWX and https://t.co/86DzS6arA0
François Fleuret @francoisfleuret
53K Followers 476 Following Research Scientist @meta (FAIR), Prof. @Unige_en, co-founder @neural_concept_. I like reality.
Lilian Weng @lilianweng
279K Followers 190 Following Co-founder of Thinking Machines Lab @thinkymachines; Ex-VP, AI Safety & robotics, applied research @OpenAI; Author of Lil'Log
Phillip Haeusler @PhillipHaeusler
453 Followers 720 Following builder. Staff ML @ Canva. thinking about video. innovation, adventure, community

















