Yijie Hao
<think> Blog
Fine-tuning a language model to be C-3PO: an empirical test of the Persona Selection Model
We fine-tune Qwen3-4B on C-3PO content across nine conditions to test Anthropic's Persona Selection Model. Results closely track PSM predictions, and one ablation turns out to be a clean instance of inoculation prompting — connecting persona selection to the emergent misalignment literature.
Narrow reward, broad shadow: RL generalization from IPD self-play
RL finetuning trains a model to optimize a specific reward. I trained Qwen3-8B with per-turn GRPO on Iterated Prisoner's Dilemma under two reward functions and asked how the finetuning generalizes. Sampling finds almost nothing; teacher-forced logprob scoring finds a small, structured latent shift whose shape differs between self-interest and spite.
Experiment plan: does the motivational structure of a reward function independently shape personality generalization?
When in-game behavior is held constant, does the structure of the reward function independently influence downstream generalization? We train Qwen3-8B on an iterated Prisoner's Dilemma with three reward functions encoding self-interest, competitiveness, and spite.
Alignment as selective generalization stability
A reframing of the alignment problem: what if alignment failures are instances of a single, deeper phenomenon — unstable generalization of proxy objectives under distribution shift?
On the geometry of representation collapse in RLHF
Why reward models trained with RLHF tend to collapse internal representations into low-dimensional manifolds, and what this means for downstream alignment.
The alignment problem in 2026: what the evidence says
In 2023, Ngo et al. laid out a theoretical case for how deep learning could produce misaligned AI. Three years later, the evidence is in — and it's more nuanced than anyone expected.
Reading notes: the alignment problem from a deep learning perspective
Reading notes on Ngo, Chan, and Mindermann's ICLR 2024 position paper. Traces a causal chain from RLHF to situationally-aware reward hacking, misaligned goals, and power-seeking behavior.
ocean / glitch / crt — a webgl experiment