The difference comes from Anthropic’s own OSS-Fuzz-based evaluation, which measures how well a model can locate and then ...