I want to understand what objectives LLMs actually pursue — not what we trained them to do, but what they learned to do, and whether the two diverge when the world shifts beneath them.
This is, I think, a core piece of the alignment problem.
I approach this via persona selection model and model trait generalization. If you are working on related problems, please feel free to reach out.