← Back to Model Beat
Models·4d ago·all news from September 11, 2026

Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills

Anthropic has introduced a new evaluation workflow for Claude Code that allows developers to test how specific plugins impact model performance. By comparing plugin-enabled outputs against a baseline, this tool helps determine whether additions improve or hinder the model's responses. The system includes six different grading types and a continuous integration gate to standardize the verification of custom capabilities.

Covered by 1 source

Related stories

ModelsUS Says Alibaba, DeepSeek Have ‘Systematically’ Siphoned AI ModelsSep 8 · 37 sourcesModelsMistral AI Raises €3 Billion With Samsung Leading the RoundSep 8 · 102 sourcesModelsNew Deepseek model V4.1-Flash cuts memory needs for AI agentsSep 8 · 29 sourcesModelsAI Music Startup Suno Launches New Models That Pay LabelsSep 9 · 5 sources