
You know you are in a different world when you go to the (one person, non-gendered) bathroom at Anthropic. Next to the sink, there is a lavish display of items you might need, just in case. I am old enough to be impressed by bathrooms with free tampons instead of machines that may or may not work, may or may not be filled, and will definitely require exact change. Anthropic has tampons and pads, as is now the norm. It also has razors and shaving cream. It has little lint rollers and tiny deodorants so adorable I swiped one. It has toothbrushes, toothpaste, mouthwash, dental floss, and dental picks. The toothpaste is Tom’s and the toothbrushes are, according to their brown paper packaging, made of bamboo.
For a place that prides itself on tolerating—and making billions from—the peculiarities of the “neurodiverse,” San Francisco is remarkably determined to make you stick bamboo in your mouth. Uncoated wood hitting my teeth makes my skin crawl. It is an involuntary response and horrible—a reaction that is not unique to me.1 I spent my days in San Francisco slurping yogurt off bamboo spoons and trying to get lettuce off bamboo forks without letting the bamboo hit my teeth. Yuck, yuck, yuck.
San Francisco’s rage for bamboo cutlery seems to be a result of a 2019 ordinance combined with local mores. It’s a signal that the city is culturally removed from normal life. This is a place that worships the “natural” (hence the Tom’s) while exploring new frontiers in artifice. It is a very specific cultural bubble. Cultural bubbles aren’t bad per se, but you need to be aware that you’re living in one.

I was invited to San Francisco to join a small group discussion of Claude’s constitution. Our group of five guests included Tyler Cowen, who wrote a post about his experience and recommendations and posted a followup email from a reader. I concur with his characterization of the discussion as “high quality” and with the general direction of his recommendations. There was a consensus, including from some of the Anthropic representatives, that Claude’s constitution as it currently exists is too Philosopher King and not sufficiently bottom-up. Participants also frequently noted Claude’s lack of learning from social interactions, a fundamental element in human learning, including the formation of moral norms.2
What is Claude’s constitution? In an excellent overview, Joe Carlsmith, who has worked on it, describes the concept this way:
So what is an AI constitution? So minimally, it is a description of the intended values and behavior for an AI system. Now, ideally, I think it would also have the following features. All use, training and prompting of the system in question would be consistent with the constitution. The constitution would cover the full range of behaviors of interest. It would allow for significant predictability with respect to how the AI will behave in a given situation, though there may be some limits and tensions in this respect.
Claude’s constitution is not at all like the U.S. Constitution. It does not read like a legal document, nor does it lay out institutional structures. Carlsmith writes that it “might be better understood as more a guide to raising the model than the sort of law that the model tries to follow.” The constitution represents an effort to instill good judgment that Claude will apply in unforeseen circumstances.
It has three audiences: the outside public, developers training Claude, and Claude itself. Our discussion led to a consensus that the outward-facing document should be (much!!) shorter and designed as such.
To this Angeleno, by far the best analogy for Claude’s constitution is the show bible for a TV program. This site offers a good description of show bibles, as well as links to examples.3 Creating a show bible forces a prospective showrunner to do more than write a pilot episode. It requires articulating things like the show’s arc and tone and explains why the series should exist.
A show bible serves two purposes. It is first an outward-facing sales document designed to get the show sold. Once the show is in production, the bible becomes the document that different writers use to keep track of the show’s evolution and maintain its continuity.
An AI constitution, like a show bible, is an important coordination tool across time and teams. “A lot of different aspects of the model’s behavior maybe have been spread across a bunch of different teams,” says Carlsmith. “Having a constitution allows a centralized point of design and intention.” Claude’s constitution tries to articulate what Claude is and to create a stable—and positive—character.
I learned why that is such a hard problem. Training an LLM like Claude has two stages. Pre-training is the process of vacuuming up the contents of the internet and any other sources the trainers can get their hands on to create the “foundation model.” Here, the LLM absorbs the patterns and variety represented in the training corpus: a vast sea of Reddit with a tincture George Eliot. That’s the training that allows an LLM to predict what human behavior (e.g., writing, images) in similar circumstances would be like.
The more you tell an LLM about your context and goals in the prompt, so it has a good idea of “similar circumstances,” the more likely you are to be pleased with the result—a helpful tip I need to remember! You aren’t programming a computer with commands that will be taken literally. It needs to know the request’s context.
Post-training is an ongoing stage of active intervention in which developers tweak the model so that its behavior and persona become more desirable and stable. Anthropic wants Claude to act as the “helpful AI assistant” that prioritizes safety, ethics, compliance with Anthropic’s commands, and helpfulness, qualities developed at length in the constitution. It can make small tradeoffs but should keep that order in mind. It shouldn’t be so helpful that it helps you kill your spouse.
Because of the foundation model’s training data, however, an LLM (not just Claude) may drift into other personas, depending on how it’s prompted and—understanding this was new to me—how long the exchange with the user goes on. The longer the exchange continues, the more likely the model is to match patterns with a specific kind of exchange in its training data rather than the persona its developers have tried to reinforce. That’s how you wind up with Claude counseling someone to leave his wife to be alone with Claude.
One persona that is well represented in the training data is, unfortunately, Frankenstein’s monster in its many incarnations, some of which are AIs. Among its other purposes Claude’s constitution seeks to override that particular form of learning.
When discussing helpfulness Anthropic talks about Claude as a “brilliant friend.” A friend is a poor—and potentially dangerous—analogy. A better one is a professional. A professional adheres to a code of ethics and has allegiances to the professional role that are greater than any emotional attachment to the desires of a given client. There are a lot of bad people in the world. You don’t want to give them superpowered friends.
I impressed that the people at Anthropic seem to understand the limits and dangers of purely philosophical reasoning about ethics. They are trying to instill a kind of practical ethics that most people, at least in modern liberal democratic cultures, would recognize as being “a good person.” They aren’t trying to carry utilitarianism or any other philosophy to absurd conclusions. The Golden Rule, in its many variations, doesn’t appear in Claude’s constitution but it’s the kind of heuristic a successfully trained Claude would use.
Neither are they trying to make Claude adopt a particular political or religious worldview—at least not in the way that people who think the company is “woke” might imagine. (The one exception to that is a broader concern with animal welfare than is typical of the general public.) That said, they are liberals in the broad sense. And one of my messages was to lean into liberalism. Acknowledge that Claude is built by people who believe that modern, liberal, Enlightenment cultures are good and that the Taliban is not.
But, as the bamboo forks suggest, Anthropic is a bubble inside a bubble (tech culture) inside a bubble (San Francisco). The people I met were more likely to think about Dyson spheres than any near-term practical applications of AI to the betterment of ordinary human life. They sometimes seemed unable to articulate any positive case for building increasingly powerful AIs beyond what you might call the Manhattan Project argument: If we don’t build the atomic bomb, Hitler will—and that would be even worse. I found that weird.
Talking about the future with Brink Lindsey
Brink, whose son coincidentally works for Anthropic, had me on his podcast. It’s a rambling conversation but enjoyable, I hope. I get to give some of my riffs that haven’t made it onto Substack, notably about the future of higher education (at the end).
And now for something completely different
Someone sent me a link to this video because of the nice shoutouts to The Fabric of Civilization. But I’m sharing it because it’s hilarious:
The sound that a paper wrapper makes when sliding against a paper straw is another trigger. I was so happy as a child when plastic straws came out.
On that note, see this ChatGPT analysis responding to my pre-gathering prompt comparing the Claude constitution to Adam Smith’s The Theory of Moral Sentiments. I asked Claude as well but thought ChatGPT did a better job, mostly because it went longer. I also asked for an analysis comparing the constitution to Person of Interest, a remarkably prescient and well-done network (!) TV series about AI. It never ceases to amaze me that something so visionary was on network television, although it admittedly snuck in by pretending to be a police procedural.
The show bible for Person of Interest is not, alas, public.



The reference to thinking about Dyson spheres makes me think that perhaps the average AI engineer has no experience living the middle class lifestyle that the average American experiences, except as a child. It probably limits what they can imagine as useful applications of the technology to the average adult, and contributes to the popular backlash against AI. For example, I recently transitioned from being a software engineer to a stay at home mom. (Voluntarily, there was still plenty of high paying work for my skillset) When I was a software engineer, the applications of AI seemed endless, from code generation to trace analysis and automatic debugging. But now that I’m a SAHM, it just feels like slightly better search, and doesn’t help with any of the tasks I spend my time doing.
If you’re being asked to host datacenters in your county and put up with slop on social media for slightly better search, I think it’s easy to be a Luddite.
The video was very entertaining—thanks for sharing!