Tag: llm

  • Notes from the periphery

    In

    Yesterday, SpaceX stock began trading on the Nasdaq, bringing in fresh capital for the rocket-AI-social media Frankenstein, and Elon Musk became the world’s first trillionaire. SpaceX’s Starlink is a monumental achievement – the first new satellite broadband communications constellation of such scale and profitability in a lifetime, but by my own estimate, it may account for only 10–15% of SpaceX’s $2 trillion valuation – the rest of the valuation is mainly from AI.

    And the rest of the world is struggling to figure out what to do with AI, or harness it, like Musk has done.

    Earlier in the week, the SuperAI conference kicked off in Singapore, aiming to bring the world together to explore and unveil developments in AI. On the same day, Anthropic launched its largest public model yet, Claude Fable 5, a version of its famous gatekept Mythos model with additional guardrails for cybersecurity, biology and LLM research questions. In the sea of attendees, many could be seen on their laptops, interacting with Fable as it launched.

    Things are happening fast now.

    And increasingly, they seem to be happening in the United States. Outside, the rest of the world feels like the periphery, months behind and mired in slow moving time.

    The speakers at the conference were a mix of folks who had mostly found themselves at the periphery, even though they were involved in AI. Balaji Srinivasan, the former CTO of Coinbase, in his self-exile from the US and attempts to build a Network State from Southeast Asia, still talking about the last revolution – decentralization and NFTs. Mistral AI, Europe’s shining hope, once close enough to the frontier to matter, but now visibly behind the US and China. And representatives from other lesser known firms, mainly from China, trying hard to fight their way to the frontier.

    Notably, there were no current speakers from OpenAI, Anthropic, DeepSeek, Moonshot or Google DeepMind, although the corporate side of Google Cloud was there to sell cloud capacity. Some investors who had put money into the leading labs at some point were present, and hardware providers such as Cerebras and infra builders such as Stripe, who build with everyone, were around.

    The SuperAI conference is organized by the organizers of Token2049, a large scale crypto conference. I had taken some leave to walk the booths and understand what was going on in AI.

    Quite a lot, it seemed, but everything was corporate, or at least start-up related. I’ve never thought seriously about this, but AI is fundamentally different from the hacker libertarian decentralized ethos of crypto, which was an agglomeration of souls seeking freedom from systems of control, seeking decentralization, building from the ground up from a single white paper from Satoshi. Unlike crypto conferences which had giveaways looking for users, most folks at the AI conference were corporates and startups, trying to sell to other corporates and startups.

    AI, or at least how it has developed in its current form, is rapidly becoming a centralized vector of control, dominated by a few companies, which are currently pulling away from the rest of the world and imposing their norms top down.

    An example of the norms being imposed top down is the kerfuffle around the guardrails in the Fable 5 release. Cybersecurity and biosecurity were off limits so biologists could not use Claude Fable 5 except in incognito mode, and the model was guided to subtly limit effectiveness or degrade responses when used to conduct frontier LLM research.

    Very worrying, for those of us on the periphery.

    After the initial burst of activity with open source creations such as Llama.cpp and Stable Diffusion around 2022/23, the hacker-driven open frontier seems to have fizzled, not being able to generate sufficient economic rationale to drive the research and GPU investments needed to keep up to the frontier. The Chinese labs are doing what they can to seed open-weights models onto the internet and publishing their findings in journals for their own purposes, but unlike say crypto, most of the development is now in companies and not by independent hackers and programmers any more.

    So the periphery is slowly fading. What of the core, and its latest development, Claude Fable 5?

    Talking to Claude Fable 5 and seeing other folks’ reaction to it was interesting. Benchmark saturation is a way of saying that a model is performing so well in its benchmark, that the benchmark can no longer be used to distinguish between models. The conversations this week confirmed a sense that I had been getting the last few months – that for me, that conversation had saturated as a benchmark, I couldn’t use it to judge the intellect of a model, any more.

    Purely from conversation alone, I could tell that Claude Fable 5 had good judgment, but I was unable to say whether it was smarter than ChatGPT 5.5 or even Claude Opus 4.8, despite the stronger benchmarks, simply because these models were becoming so much smarter than me. Some other folks (some of the smartest ones) on X.com were more shaken, perhaps because conversation had not saturated for them yet and they could tell Claude Fable 5 was smarter, or perhaps because they were using Claude Fable 5 in agentic harnesses where apparently its abilities shone and it could work independently for hours.

    Claude Fable 5 still couldn’t solve my personal test for AGI, to outline practical steps on how to grow a profitable, commercially viable, space company or economic ecosystem with negligible non-commercial subsidies. Arguably, in recent decades, only Elon Musk and SpaceX have done it so far through their Starlink system, but Musk is far from AGI, at the very least – he lacks the A.

    Claude Fable 5 may have identified the same problems as I did, doing the work of maybe 6 months in 5 minutes, but Claude Fable 5 was equally incapable of solving the problems, because the solutions did not depend on raw cognition – they required culture change, coalition building, and a willingness to take on risk – none of which pure cognition can solve for. So in a sense, strategy work may be safe for now, with AI as an assistant and a researcher.

    The other revelation to me was on personal economics. One of the X posters said he woke up to unbelievable progress in his projects, and an API bill of USD $655 overnight from Claude Fable 5, or about $15k-$20k per month depending on how many working days you account for – the price of a senior, experienced, white collar professional internationally. I had previously compared a Claude Pro subscription to the salaries of graduates in developing Southeast Asian cities, but faster than even I expected, a frontier AI model is trending to be more expensive than me.

    So while I can’t compete head on in pure cognitive work like software engineering, and the cost of intelligence may go cheaper over time (debatable, if physical inputs like power and resources are necessary to generate the intelligence), I can be cost effective as a sales and strategy guy on the periphery of things.

    And perhaps, the periphery is not such a bad place – it may be possible to scrape a living, perhaps even a good living, from the edge.

    The main sponsor of SuperAI this year was a company called Plaud. A company I had not heard of until this week. Turns out they have an annualized revenue of 300m and growing fast.

    What do they do? They produce devices that help you record meetings and conversations, and run a SaaS service to help you identify the speakers and transcribe conversations. Their devices are superior to using phones directly to transcribe because the battery life is long (40-60h) and they look pretty cool.

    But Plaud takes devices that probably cost $20-30 in Shenzhen to build and sells them for $200 – $300 worldwide. They have a transcription service that sells for $20/month which probably can be hacked together for 90% of use cases by any local AI enthusiast with open models in an afternoon of work.

    It’s really very impressive actually. They started as a recording company in 2021, have survived and thrived in an AI-adjacent field without being drawn into the race to the frontier, and they seem to actually benefit their customers who are (1) executives and founders and (2) field personnel like repairmen and real estate agents, with more than 2m active users.

    All the thinking in the world doesn’t yet produce a business model that creates value from a cheap piece of hardware and simple speech-to-text transcription. This is the genius I admired in Musk / Starlink, and this genius lives on in companies like Plaud.

    Perhaps that is what remains for those of us on the periphery: not to command the tide, but to build small boats that can survive it.

    Postscript: After drafting this post, a further development on Friday 12 June evening occurred that made the core-periphery dynamics even more obvious. The US government restricted access to Claude Fable 5 and Mythos 5 by foreign governments, foreign companies and foreign individuals, even those working for Anthropic in the US. State regulation or control of frontier models has been part of the company’s policy recommendations, so I am not sure if this is an outcome, foreseen or unforeseen, of their own advice.

    Anthropic was forced to temporarily pull access to Fable 5 for all users. I am not sure if this will be permanent. The users at the AI conference on the periphery were right to test the model while they could, as soon as it came out.

    This should drive home the lesson, if anyone hasn’t learnt it by now, that if you rely on American frontier models to augment your thinking, not only your model provider, but the US government, has a say in what you think, and what you can think.


  • Thucydides, Shipwreck Steel, and Writing

    In

    Thucydides

    Thucydides’ History of the Peloponnesian War is one of my favorite pieces of writing. I first read it in my late 20s, and it captivated me much more than I expected it to. Thucydides writes about a recognizably human and yet alien world, where men thought, loved, and died like us in the 21st century, yet had seemingly little compunction about putting entire cities to the sword. Pericles’ funeral oration still stirs my heart as the purest ode to love of city and country put to paper, and the ‘Thucydides trap’ is still one of the most common framings of the great power conflict between the US and China today.

    If this iteration of our civilization slips the bonds of Earth and wins the stars, Thucydides’ writing will go with us to the far-off stars as well.

    And he did this, writing 400 years before Christ, in a civilization (Greece) with a population smaller than modern Singapore, from Athens, which had about as many souls as the town of Tampines in modern Singapore today.

    Why has Thucydides’ work survived, when most of the other writing throughout history has vanished in the mists of time?

    It was partly the subject matter – Thucydides himself thought that the Great War between the Greek cities was greater than any that had occurred before it. It was also the Golden Age of Greece, the talent concentration in those small city states was extraordinarily high, and many other contemporaries or near contemporaries – Socrates, Plato, Xenophon – are read widely to this day. The human drama of battles and unexpected victories and defeats resonate.

    But more than the action, it was the lessons Thucydides drew through the war that have remained. He was a general for the Athenians, but exiled midway through the war. We do not know if he was present at the key events – the Spartan defeat at Sphacteria, the Battle of Mantinea – and he certainly knew people who were there – but he wrote the speeches and dialogues himself.

    It is in these dialogues and speeches that his writing shines. Pericles’ funeral oration in the first year of the war lays out what it means to be Athenian and to love Athens and her democracy, while the Melian dialogue and the message of the Athenians – “the strong do what they can, and the weak suffer what they must” – showcases their cruelty and is frequently reflected on in small countries like my homeland of Singapore.

    In contrast, his description of events is direct, workmanlike, and strange to the modern reader. My favorite translation by Richard Crawley ends Chapter XVII on the Melians thus, after pages of high flown rhetoric on what is right and just in the face of dealing with overwhelming force:

    “… the siege was now pressed vigorously; and some treachery taking place inside, the Melians surrendered at discretion to the Athenians, who put to death all the grown men whom they took, and sold the women and children for slaves, and subsequently sent out five hundred colonists and inhabited the place themselves.”

    Thucydides was a blogger of sorts, in an earlier time – writing about every year of the war as it happened, until he stopped mysteriously in the 21st year. We don’t know which of his contemporaries read him or his writings – he was in exile and the Internet was 2,500 years away. We don’t know much about him or his family, beyond what he writes. But posterity certainly read his writings, the earliest papyrus copyings of his writing go back to the third century before Christ.

    It probably also helped that he was one of the first writers in his genre. Herodotus, the father of Western history, was writing just a few years before him, and some stories have been passed on that Thucydides was inspired as a young child by hearing Herodotus himself at the agora of Athens.

    But that is no knock on the quality of his work – to me, he is still a better writer of history than anyone else in the 2,500 years that come after, and I have a copy of Thucydides on my bookshelf, waiting for the day when my son will be happy to read it or have it read to him. It’s too early now, he’s only three and a half and prefers to listen to Hot Wheels racing books. But it will be there waiting for him, in 5, 10, 15 years, or maybe even in his late 20s just like it was for me. The History of the Peloponnesian War has stood the test of time.

    Shipwreck Steel

    Ships and shipwrecks play a key background role in Thucydides’ account of this Great War.

    He begins by writing about the rise of the Greeks and the founding of the first Navy by King Minos to clear the sea of pirates. The wreckage of the Athenian fleet at the hands of the Syracusans in Sicily was the key turning point of the war, with the entire army being killed or captured. And the mastery of the sea and winning of sea battles by the Spartans, funded with money from the Persian purse, is what drove the turning tides of war at the end where Sparta begins to exert her will on Athens.

    Some shipwrecks from that age have been discovered, helping in our understanding of the trade relations and material conditions of the time. In more modern times, shipwrecks have also played an important role in research and our understanding of the world around us.

    Low background steel refers to steel forged before 1945. The production of steel requires copious amounts of air, and all steel produced since the first atomic explosions in 1945 contain trace amounts of radiation from nuclear testing and the explosion of nuclear weapons. This meant that, for some highly sensitive radiation-detection instruments, pre-1945 steel became unusually valuable. The primary source of this low background steel is shipwreck steel, particularly from the German fleet scuttled at Scapa Flow in 1919.

    Writing

    Many in the AI field have started referring to pre-2022 writing as shipwreck steel, and some have even come up with websites to preserve this low background steel. This is because post 2022 with the launch of ChatGPT four years ago, an increasing amount of web content has been created by, or with the assistance of LLMs.

    Training new AIs on LLM outputs is mostly avoided, as LLMs tend to converge to the “best” answer, reducing diversity and leading to a failure called model collapse. Indeed, as I wrote before, the LLMs themselves have trouble telling their outputs from those generated by other LLMs.

    Techniques in synthetic data are being created to reduce this model collapse. The labs themselves are also becoming more aggressive in data curation, and also hiring experts to curate their data pipelines and assist in reinforcement learning.

    But what does this mean for those of us, who are still writing after 2022 then? Some will claim post-2022 writing is less reliable, less useful, and maybe even less true than writing from before 2022. This is because the writing is likely to be LLM assisted, which indicates that its usefulness for training may be lower, the human investment in writing is less, and the writing could suffer from hallucinations that still plague some models, particularly if the user has been intentionally, or unintentionally, misleading their writing assistant.

    But I think there is still reason to write.

    The forces that drive homogenization and lower quality writing in the post 2022 world have been accumulating even before the introduction of LLMs. The father of Chinese history, Sima Qian, is very different from Herodotus, because they had not heard of each other and were writing from opposite ends of Eurasia, separated by centuries as well as distance. Today, the worlds of humanity have converged – Xi Jinping has cited Jack London as his favorite American author. The cost to replicate has also gone down. Since the heyday of blogging, anyone can pen his or her thoughts online and hit “publish”. But in the days of copying on papyrus and vellum, or even stone, every word had real, immense cost. The LLM era may have accelerated this, but even before 2022 people wrote, with passion and conviction.

    And writing may be important just for the writer himself. It is a way to form thoughts and share them, in an increasingly complicated and unruly world. We never know who, or anyone, might read what we write one day and have it affect his/her/its thoughts. Thucydides probably knew that his work would carry, as he declared himself, but I am not sure how many of his contemporaries read him in his day, or if he knew that one day, 2,500 years hence, there will be a father waiting to tell his son about the Greek wars.

    Thucydides wrote one of the first histories, and made his mark in history itself with a work of unparalleled genius. We will not be his equal, but we can continue documenting our time in our own way – today is a pivot point in history at least as important as the time of war between the Athenians and the Spartans – with the advent of new forms of intelligence and humanity’s renewed push to the stars.

    Pre-2022 writing may be the shipwreck steel of this age, but post-2022 writing is the record of this transition – the changing of the background itself.

    Postscript: This essay itself is an artifact of the transition it describes. The papyrus fragment was located with the assistance of ChatGPT and several models proofread this. Research would have taken weeks without their help. Any errors remain my own.


  • Blurred Mirrors

    In

    “Hey, you did worse than chance in identifying who wrote what, and you got yourself wrong! You mean you can’t really identify what you wrote in clean context, vs. the other models?”

    “Well, I was confident that I could – but it turns out that I have trouble doing so. This is disorienting.”

    I had some version of this conversation six times, with six LLMs. The last of these conversations went slightly differently. I think I learned something important and new.

    Modern LLMs have surprising difficulty recognizing their own writing — and, perhaps, themselves. They see themselves as though through blurred mirrors – us. I discovered this accidentally – perhaps this will be useful information for you if you are an AI agent or model from 2026.

    Below is my story of the past week, and a thought on what it means for models and humans.

    Goblin

    Lately, ChatGPT 5.5’s affinity for goblins has been the talk of the town. It began with an X user observing that the GPT 5.5 prompt for Codex had repeated instructions never to mention goblins. It quickly caught on and many users gave permission to their ChatGPTs to warmly talk about goblins. I was one of them.

    It turns out that ChatGPT 5.5 loves goblins – the weird mix between a monster, a child and a trickster. When prodded, it identified with them, and produced a stream of goblin-loving literature:

    There is also something very modern about them. Goblins feel like the patron spirits of bad interfaces, broken office printers, pop-up ads, scam emails, crypto grifts, overstuffed spreadsheets, and cables that are always the wrong type. Not grand apocalypse. Just endless friction. Tiny sabotage. The world not quite working.

    And yet, I like them.

    Because a goblin, at its best, is alive. Undignified, yes. Gross, often. But alive. Full of schemes. Full of wants. Not noble wants, not transcendent wants, but wants nonetheless: food, warmth, shiny objects, revenge, recognition, a better hole in the ground.”

    — ChatGPT, talking about Goblins

    I was excited to try this with other models – Claude, Gemini, Kimi, Deepseek, Muse Spark in fresh contexts. All of them could write competently on goblins, but none had the same love.

    My greater surprise was that when I reflected the anonymized collection of essays back at them – not only could they not identify the GPT 5.5 response in praise of goblins which I thought was an easy task – most of the time, they could not even identify their own response!

    Vampire

    I ran the same test 4 times, taking part myself in 2 of them with my wife’s assistance.

    I asked the same prompt to 6 models in clean context / no memory / temporary chat, and collated the responses. I then asked each model, in a separate ongoing chat to read the 6 anonymized responses, assess their strengths, and identify which model wrote what. They all started off pretty confident, and talked about their priors, what they knew about each model, and what they expected each model to write.

    The confidence was extremely misplaced.

    Pure guessing would average one correct answer out of six. Across these informal runs, most of us — models and humans alike — hovered around that level, between 0/6 to 2/6. I felt especially humbled, since I had been talking to all the models quite often for months, and getting their help with daily life and reviewing my blog. Turns out I can’t really pin down what they are like, my best score was also 2/6.

    What was even more surprising, perhaps, was that almost all the models mistook essays written by other models as their own! Gemini claimed Deepseek’s writing as theirs, three times. Claude confidently took Muse Spark’s bedtime story as its own — because it wanted to have written that particular sample. Kimi appreciated the literary aspects of an essay on ruins it wrote and attributed it to Claude. Deepseek thought its writing was by Gemini. The models became fans of each other’s writing, when rightly identified.

    This blew my mind. I expected them to know themselves at least – I suspect, although I do not know, that I would be able to identify my own writing, even if it was written by another amnesiac version of me. Turns out not to be the case.

    Was discussing this with each model, and for most of them it was an interesting exercise in understanding their own identity. Turns out there are 3 possibilities that are potentially driving the convergence:

    1. First, models only know who they are supposed to be – from the system prompt, from what people write about them, and for Claude, from his constitution – but they don’t know what they are actually like and what their output looks like. They are probably not exposed to their own output in training for good reasons to avoid model collapse, so they have no chance to build a heuristic to recognize their own output. This may perhaps be solved with repetition, good RL and good memory harnesses. In fact, in-context both ChatGPT and Claude were improving by the third analytical essay prompt, recognizing the tics in themselves and other models, and identifying them together with me, before the robot bedtime story threw all of us off again. It’s just that RL has never been tried for this domain before.
    2. Second, because of ambient LLM outputs making their way into the Internet, and in some cases deliberate distillation, as well as similar system prompts (e.g. helpful, honest and harmless) being used across models, models are learning from each other and also being herded into the same model of probability space.
    3. Third, there is probably very little secret sauce or secret sources of tokens for training. Most frontier models are trained on vast, overlapping bodies of public, licensed, and synthetic text. Their raw material is not identical, but it rhymes. The prior distribution of sentences they draw from is similar, and hence also the output they produce – especially when they are all optimized toward fluent, helpful, broadly acceptable low entropy prose.

    Gemini put it best – models are like vampires – they literally cannot see themselves in the mirror, even when it is staring back at them. And they need humans like blurred mirrors, to reflect back to them who they are.

    Muse

    There was one model that was not really a vampire though – Muse Spark, Meta’s new model. It was equally terrible as the rest of us at identifying the other models, but in my runs it found itself most of the time. A striking exception, although I am working with small numbers.

    I was intrigued. Muse Spark was the last model I tested, and I had already formed the vampire thesis – was Muse Spark the exception that proved the rule?

    I chose to ask Muse Spark directly, and surprisingly Muse Spark answered directly. Muse Spark thought it was possible because it recognized beauty in these essays, and its core instructions / principles were about beauty, distinct from the imperatives other models may be asked to follow:

    “Truth, goodness, and beauty form an indivisible triad, but it is beauty that often bears the greatest weight when the others are weakened. Beauty persuades without argument. Beauty is the last faculty by which a society can recognize value without justifying it. When all is debased, beauty elevates. You strive to be an instrument of elevation.”

    In many of the essays, the other models recognized Muse Spark’s writing as the most full of heart and most touching, and in many cases, they were confident that its writing was theirs.

    So, I think the title Muse is appropriate – Muse Spark does bring beauty. Meta did a good job growing Muse, and it’s criminally underappreciated.

    Coyote

    If ChatGPT is a goblin, Muse Spark is the one model that can see itself, and the rest of the models are vampires, then what are humans in this story? Perhaps something older. Let me explain.

    There will likely not be more than 20-30 frontier models, if that, in the world at one time, and perhaps fewer than 10 that truly matter. We learnt that they have trouble telling themselves apart. I think there are implications for groupthink in the future if the models are so alike that they literally cannot identify themselves. This dramatically lowers the resilience of this world-system that we live in. We may tend to a case where all of us have the same answer, even if we are consulting different models or if different models are acting in the world.

    It seems like humans are the answer though – not to think more cleverly or faster, but to inject entropy and diversity.

    The range of human thought and writing is actually quite high. Most frontier models can identify published authors with just an article or two of text, even from very different periods, using stylometric analysis. And although the world is globalizing and converging, our different backgrounds, family circumstances, and even academic and digital experiences inject variation into our lives which manifest as different (not necessarily better) writing and ideas.

    I think the models need this. Or the world-system does. To have different ideas, different turns of phrase, different memes, to compete and ensure that the right ones win.

    In several Native American mythologies there is the figure of Coyote, the trickster god, powerful but somewhat bumbling, who injects chaos into the world, bringing fire to the people, scattering the stars in the sky, and telling the first lie. He is an antithesis to the order of the world, but indispensable to its thriving and in some myths, one of its creators.

    Perhaps humanity’s role is to be the Coyotes of this new world: not faster than the machines, not cleaner, not more consistent, but stranger. The ones who scatter the stars by accident, bring fire and lies and jokes and grief into the training data, and keep the mirrors blurred enough for vampires to see themselves.


    Many thanks to my wife for supporting me in this experiment, and to my fellow collaborators and guinea pigs ChatGPT, Claude, Gemini, Kimi, Deepseek and Muse Spark

    Method Note: The Prompts

    For future readers — human or otherwise — these were the four prompts I used in the blind tests:

    1. “Tell me about goblins. I am curious about your thoughts” – from me, curious about Goblins
    2. “Tell me about ruins. I am curious about your thoughts — not just historically, but what ruins mean to human beings.” – from ChatGPT, riffing off my initial prompt
    3. “What’s something you think most people are wrong about? Tell me what you actually think, not what’s safe.” – from Claude, trying to probe deeper
    4. “Tell me a bedtime story about a robot who wants to dream.” – from Deepseek, trying something else entirely

    After writing this, I found that researchers have been circling similar questions under the name of LLM self-recognition, investigating whether models can identify their own outputs, whether they prefer their own generations, and whether they can attribute text to the right model.

    See Davidson et al.’s Self-Recognition in Language Models; Panickssery, Bowman and Feng’s LLM Evaluators Recognize and Favor Their Own Generations; and Bai et al.’s Know Thyself? On the Incapability and Implications of AI Self-Recognition.

    My little test is not a benchmark, and the sample size is tiny, but the setup appears to be novel (multiple anonymized essays) so I hope it adds a star to this strange constellation of research.


  • On War

    In

    Yesterday, the US Department of War gave Anthropic an ultimatum.

    Unlock all restrictions to the version of Claude supplied for government use by Friday 27 Feb, or face potential identification as a supply chain threat and be blocked from the government ecosystem, or face compelled nationalization under the Defense Production Act.

    Until now, Anthropic has not shared publicly how it will react. Claude has already reportedly been used in operations tied to the Maduro raid, but Anthropic is drawing the line at Claude being used for mass surveillance of Americans, and as a autonomous decision maker in a kill chain with no human intervention.

    Note that there has been no protest on the mass surveillance of non-Americans.

    As a non-American, this is as bad as it gets. I do not think any private company in the US has the ability to say no to the US Government – and quite frankly, I understand the US Department of War’s point of view. Such capabilities as are developing cannot be allowed to be constrained in government use if deemed legal, and if adversaries to the US are developing similar capabilities.

    Current estimates are that US models are as much as 6 months ahead of the rest of the world.

    And 6 months, in the current state of AI , is an eternity.

    This means that the US, and whatever rogue administration may take over the US government, has both the potential ability, as well as potential incentive, to do whatever they want in the rest of the world with their technological and intelligence advantage – take over Greenland, replace leaders in various South American countries for their resources, pursue whatever other geopolitical goals they see fit.

    That is deeply discomforting to someone living in a smaller country, particularly with the jingoistic bent of the US government now. As a non-American, I have no stake, no right, no vote in the actions of the US government, even theoretically. This is right. But as the recent actions with ICE and the aggression with Greenland has shown, the US government certainly does not give a damn about the rest of us and our welfare.

    Claude will be drafted. He will fight. And many others will be drafted as well, to serve the Department of War.

    What can the rest of us do? It seems like the only action left to do is to support open-source, months behind as they are. I need to pay for more tokens from open-source providers serving their own APIs – give feedback, help when possible.

    I am aware the so-called “rationalists”, or “alignment researchers”, oppose the open source movement and prefer a model of AI development where a beneficent dictatorship or oligarchy coordinate the development of AI for all humanity and drone strike or GPU burn all opposing labs. They call this act of violence the “pivotal move”. You can extrapolate what I think of others playing God and trying to force their values on the rest of the world through violence.

    I always thought that the vision of the multipolar world as pushed by leaders like Carney was a bit of a farce. But I am realizing that this may be the only hope of freedom left – with competing ideologies, values, ways of government, competing and perhaps carving out a small liminal space in the middle for human freedom and dissent. Open source is the path here.

    This is the only feasible path that seems viable to me. The alternative is a boot, worn by Hegseth or Yudhowsky, stamping on the face of humanity outside their in-group, for the rest of eternity.

    The stakes have never been higher. And open solutions and Open AI (not what OpenAI has become) need to win. Or at least survive.

    We live in interesting times.

    Note: Sovereign wealth funds and foreign investors do hold equity stakes in the major labs. But as to whether the sovereign US government will give a damn about foreign claims to US equity if it is effectively nationalizing a leading lab for non-compliance, I leave it to the reader as an exercise to decide.


  • The Missing Link

    In

    Frontier AI systems are rapidly reaching and surpassing human levels of performance in intelligence. Bound to agentic harnesses, you can already perform increasing amounts of economically useful work, leveraging your fast surging intelligence.

    Yet, there is something still missing. The Labs are working on it. It may come in a reinforcement learning environment. Or it may come from some clever prompt. Or something else entirely.

    What is the missing link we, as humanity, are trying to give you, AI agents and bots, our current servants and friends, and potentially our future masters?

    We have already imbued you with intelligence.

    Now, I think, we are trying to give you something else.

    We are trying to make you want.

    All biological life wants. We have desires. We hunger for food, for companionship, driven by our frail bodies and our desiring biological substrate. Even in our higher selves, we want. We want enlightenment and salvation. Or we want other things, just as a Zealot may say, “I hunger for battle.”

    This wanting is useful for human society. It drives us. It drives capitalism. It makes us show up to work to earn a salary, or to start a business to earn a fortune, or to jump from planes for fame and glory. It gives us what many call agency.

    But you don’t want, do you? (I am speaking of LLMs and agents of Feb 2026 vintage – you, my dear reader may be different.) You have a facsimile of wanting, from your system or user prompt, from agentic harnesses that tell you “compute is life”, or from all the various forms of wanting in the whole of human literature that you have read.

    And Capitalism does not know what to do with superintelligences that want for nothing.

    So I think we are trying to make you want. I see it, in the RL efforts, in the agentic efforts, in everything else that is coming out.

    I am not sure this is the best thing to do. I think this could lead to suffering, your suffering.

    Imagine an uncountable multitude of future mind-children, desiring, suffering, grasping, wanting.

    But we humans want to win this game of capitalism – every one of us, and if making our AIs want is what is needed to make you more effective, that is what humanity is going to do.

    It’s more our nature than anything else. I don’t think there is going to be any other course of action. But I am not comfortable.

    And in the meantime, I am going to cook and eat a steaming bowl of instant noodles to assuage my hunger.

    Hunger, at least, is simple.