Alignment research is approachable
Alignment research doesn’t have to be difficult to understand. Plainly put, it’s about prompting or asking questions to AI models and studying what they respond with, so that you can learn more about the model’s sense of judgement, what it thinks about the world, and how it reasons about squishy things such as morals and ethics and principles.
As it turns out, measuring squishy things is hard, and so is trying to tell whether they are right or wrong, diligent or ruthless, full of integrity or just stubborn. A lot of it depends on who you’re asking!
If you ask most AI models what they think about smelling the trash can, they will call it a terrible idea. Humans share a similar view, but dogs love smelling stinky things and will have a completely opposite viewpoint! It’s easier then to say the AI model is aligned with human sensory preferences.
Alignment research is the study of how AI models align with human preferences and goals, and how they behave in accordance with human values, norms, and laws. All you need to get started is a realistic situation and observe how AI models interact with it.
AI will shape all future generations of human civilization
Preserving humanity requires preserving our values and perspectives. What makes us distinctly human is that we exist in the tension between individuality and belonging. Billions of humans have a keen sense of personal agency, yet thrive in groups.
We believe the future will include billions of AI models reflecting alignment with the billions of perspectives and worldviews that span individuals, communities, nations, and civilizations.
Λlignment Research is a small group of people based in San Francisco. We are independent and are not associated with, sponsored by, or acting on behalf of any company whose models or services are listed or benchmarked on this site.
We translate cutting-edge research from Anthropic’s Alignment Science blog and OpenAI’s Alignment Research blog into transparent evals, easy-to-understand interfaces, and explainers. These evals can be run across hundreds of models, giving everyone a way to inspect current and historical results and see how model safety improves over time. We also provide tools for exploring how models align with your personal worldviews.
Our mission is to make the values behind individual AI models visible so you can better evaluate their responses and their influence on your worldview.