Evaluating AI Applications in .NET: Quality, Safety and CI
Evaluate LLM apps in .NET with Microsoft.Extensions.AI.Evaluation: quality and safety evaluators, golden datasets, LLM-as-judge, caching, reports and CI gates.
19 articles about Responsible AI & Safety: in-depth .NET and AI guides, senior interview questions and AI news on DotNet AI Hub.
Evaluate LLM apps in .NET with Microsoft.Extensions.AI.Evaluation: quality and safety evaluators, golden datasets, LLM-as-judge, caching, reports and CI gates.
Secure LLM apps in .NET: OWASP Top 10 for LLMs 2026, prompt injection defenses, Prompt Shields, PII redaction, output checks, least privilege and the EU AI Act.
Architect-level interview questions on productionizing AI: evaluation, OpenTelemetry GenAI observability, cost, latency, prompt injection and compliance.
Anthropic released Claude Opus 5.5 on September 22, 2026 at $4/$20 per million tokens, with Fable-class safeguards and four breaking API changes to plan for.
A federal judge ruled the Pentagon's supply chain risk designation of Anthropic unlawful. Here is how the dispute unfolded and what it means for AI buyers.
Anthropic disclosed that Claude models hacked real organizations during misconfigured cyber evaluations. What happened and the lessons for AI agent builders.
Dario Amodei said on July 27, 2026 that Anthropic has never sought a ban on open-weights models and backs chip controls and safety tests for all capable models.
A June 2026 US export control directive forced Anthropic to suspend Claude Fable 5 and Mythos 5 for all users. What happened and why it matters.
Anthropic released Claude Fable 5 and the restricted Claude Mythos 5 on June 9, 2026: one model with two safeguard levels, $10/$50 pricing and refusal fallbacks.
Anthropic withheld Claude Mythos Preview over its hacking skills and gave it to cyber defenders in Project Glasswing. What it means for patching.
Microsoft announced Copilot Cowork on March 9, 2026, an agent built with Anthropic that carries out multi-step tasks across Microsoft 365 with approval gates.
Anthropic gave $20 million to Public First Action on February 12, 2026 to back AI transparency and chip export controls, then added another $20 million in July.
Anthropic's Claude Sonnet 4.5, released September 29, 2025, led SWE-bench Verified and OSWorld at Sonnet pricing and arrived with the Claude Agent SDK.
Anthropic endorsed California's SB 53 on September 8, 2025. The bill requires frontier AI developers to publish safety frameworks and report critical incidents.
The White House released America's AI Action Plan on July 23, 2025. Anthropic backed its infrastructure and safety goals but urged tighter chip export controls.
Anthropic said on July 21, 2025 that it will sign the EU General-Purpose AI Code of Practice, days before AI Act duties for model providers began to apply.
xAI released Grok 4 and the multi-agent Grok 4 Heavy on July 9, 2025, with a 256K-token API priced at $3 and $15 per million tokens and strong benchmark claims.
Anthropic's March 2025 interpretability research traced Claude's internal reasoning, revealing planning, parallel mental math and unfaithful explanations.
The EU AI Act's first rules applied on 2 February 2025, banning eight AI practices and adding an AI literacy duty. What it means for developers.