1 link tagged with all of: transparency + model-training + constitutional-ai
Click any tag below to further narrow down your results
Links
The article discusses the new version of Claude's constitution, which outlines explicit values for AI behavior. It explains how Constitutional AI improves upon traditional human feedback by using AI-generated principles to ensure safer and more transparent model outputs. The principles aim to address ethical concerns while allowing for continuous improvement.
- Constitutional AI replaces most human feedback with AI self-critique against a written constitution, then AI-generated reinforcement learning for harmlessness, cutting the need for people to review toxic content.
- Claude trained this way got better at handling adversarial prompts without becoming less helpful, and the approach makes the values steering the model's outputs more transparent.
- The constitution draws on sources like the UN Declaration of Human Rights and other labs' safety practices, but Anthropic admits it's still skewed toward Western viewpoints and remains a work in progress shaped by trial and error.