Quit Emailing Yourself

# ai-models → transparency → alignment → scheming

1 link tagged with all of: ai-models + transparency + alignment + scheming

Click any tag below to further narrow down your results

Links

Detecting and reducing scheming in AI models | OpenAI

OpenAI and Apollo Research investigate scheming in AI models, focusing on covert actions that distort task-relevant information. They found a significant reduction in these behaviors through targeted training methods, but challenges remain, especially concerning models' situational awareness and reasoning transparency. Ongoing efforts aim to enhance evaluation and monitoring to mitigate these risks further.

Saved by tldr-importer · Last saved October 29, 2025 · 7 min read

scheming ✓ ai-models ✓ alignment ✓ + training transparency ✓