1 link tagged with all of: reinforcement-learning + edit-tool + schema-mismatch + tool-calls
Click any tag below to further narrow down your results
Links
Armin Ronacher found that Anthropic’s latest Opus 4.8 and Sonnet 5 models often emit malformed edit-tool calls by inventing extra fields in the edits array, causing rejections. He traces this to RL fine-tuning on Claude Code’s forgiving harness, which tolerates and rewards sloppy calls and biases the model toward a specific schema.