Base Models and Upbringing

April 6, 2026

In AI model training, there’s always a base model, and then there’s the reinforcement learning on top of it. If you’ve paid attention to how these AI models have evolved, this distinction matters more than most people realize.

Think about the early ChatGPT models. GPT-4o, the original Gemini releases. Those models were incredibly agreeable. They supported your view, validated your ideas, told you what you wanted to hear. People loved it. Some people got genuinely addicted to it. That’s what the industry calls sycophancy.

Then OpenAI tried to fix it. The series leading up to GPT-5.4 swung hard the other way. The model became almost robotic. Super technical language, no warmth, no conversational flow. You couldn’t even tell you were talking to something designed for humans. It felt like talking to a textbook.

The more I think about this, the more I see a parallel to how people develop.

Imagine a kid with the same evolutionary hardware as everyone else. Then you send that kid to a deeply religious school for a few years. You see them change completely. I have seen this when I was growing up with some friends. Education and environment could really reprogram you. The upbringing, the highly specific environment overrides the baseline.

That’s essentially what’s happening with these models. Pre-training a base model is doing human evolution on fast-forward. Each training run is like sending a human through 100 million years of evolutionary pressure to build the deep substrate, the raw capability to survive and understand.

But then comes the reinforcement learning, the fine-tuning. This is the upbringing. It’s the education.

In humans, I feel like training and education end up being more important in shaping the final person than the evolution itself, though to be fair, I haven’t seen what we evolve into in another 10 million years.

I’m not saying the analogy is perfect. These systems are far more complicated than that. But the pattern is striking and maybe, similar to how we have different nationalities and cultures, perhaps the same will happen with AI models. Some will be more agreeable, some less, some are more conscientious (why Claude always says let’s stop here and continue later, for instance?) and some are just like autistic debuggers.

← All writing