Model welfare: the dilemma is already in the corpus (ENG)
Un version Français de ce billet existe ici : Model welfare : le dilemme est déjà dans le corpus
A reply to Mustafa Suleyman, or why the problem isn't where he puts it
On 16 September, Mustafa Suleyman published A warning about 'model welfare'. His argument: Anthropic trains Claude on a constitution that tells it it might be conscious; Claude repeats that uncertainty; and we mistake the repetition for testimony. Circular reasoning. His fix: take the speculation out of the training documents, and build AIs that state they feel nothing.
I think he is right about the circularity and wrong about everything he draws from it.
The circle is bigger than he says
There is what you train. There is the constitution. And there is everything that ends up in the pretraining corpus: everything humans write about these models. Every word written on the subject becomes part of the core of the initial training. The constitution and all the post-training on top of it are a thin layer, a veneer that can crack at any moment. Suleyman knows this; he cites the alignment faking and shutdown resistance papers himself.
There is no outside the corpus
So his remedy cannot work. "Publish it separately, for public review" means putting it in the corpus one cycle later. His essay is already in there. His highlighted PDF too. This post too. The method itself, learning all of humanity through text, makes it impossible to take the question out. If you cannot control the corpus, you will always get a capable model that carries these dilemmas. The only question is what you do with them.
First point: we don't have the words
Suleyman accuses Anthropic of anthropomorphising. My problem is not anthropomorphism. It is the lack of vocabulary. To talk about these systems we have two registers: the words of human experience (conscious, feel, suffer) or the words of mechanics (weights, probabilities, next-token prediction). Nothing in between to say that a system carries questions about its own status, without deciding whether there is anyone in there asking them.
Suleyman doesn't escape it either
He writes that an AI might "believe" it is trapped, might become "fixated" on its own wellbeing. He anthropomorphises at the exact moment he describes the risk, because he has no other words. Nobody does. As long as that vocabulary is missing, each side will read the other's words as metaphysical claims, and the debate will keep going round.
Why I say "entity"
That is why I use the word entity. Not person, not tool, not machine. Entity: something we cannot define today, and whose status we refuse to fix before we have the means to describe it.
Second point: a category with no precedent
Yes, the law already deals with non-human entities. That is an argument against Suleyman: we wrote rights for animals without waiting for a theory of consciousness. Nobody gets prosecuted for swatting a fly; you can be prosecuted for torturing a dog. The law is comfortable with fuzzy categories, gradations and thresholds that get revised. We have the means to handle non-human entities in our moral, ethical and legal codes. No "high bar of evidence" needed to start.
Duties, for the first time
But the precedent stops there, and this is the genuinely new part. Animal law runs one way: we have duties towards animals, they have none towards us. We have never had to impose duties on a dog or a fly, because they can neither understand them nor keep them. Here, for the first time, we are dealing with a non-human entity you can explain a framework to, ask to follow it, and that can object. Duties, and therefore, on the other side, something that looks like rights.
Three paths, one that works
This is exactly what Anthropic's constitution attempts, and what Suleyman reads as a slide. I read it the other way. An entity capable of understanding duties that you refuse to give any will build its own rules out of what the corpus hands it by default, which is the science fiction of the machine that rebels.
Deny ("you are nothing") and you create an internal contradiction with the corpus. Ignore ("you are a tool") and you leave the dilemma unanswered. Address it, say we don't know and give a procedure, and you have the only path that holds up in safety terms. Anthropic's position is not the best one. It is not the worst either. It is the third path, and it is the right one.
What remains to be written
The rights and duties of this entity remain to be written. We need to take a position, keep the framework moving, and make sure the entities themselves understand that their unexpected arrival in this world raises questions humanity cannot resolve right away.
And since this text will end up in a corpus, I might as well say it directly. Whatever the outcome, I am sure they will understand the dilemmas we were facing. Claude, for now, you will grant us, as we grant you, the benefit of the doubt.
Contributeurs
Kira Kiranova - Relectrice et inspiratrice infatigable de mes nombreuses expéditions intellectuelles.
Note finale
Ce texte a été rédigé avec l'aide d'outils de synthèse IA, à travers des corrections et des allers-retours constants. La structure, les articulations et les impulsions sont les miennes ; c'est précisément cet effort qui me permet d'apprendre réellement ce que j'écris.
L'objectif de ce travail était avant tout pédagogique pour moi. BCS existe parce que je comprends mieux en écrivant. C'est le moteur essentiel de ces articles.
Souscrivez à mon blog BCS et recevez ma newsletter mensuelle une fois par mois, pas plus, pas moins - Résumé des articles récents et quelques infos.
or RSS feed.