Claude Code Skills·Claude Skills·The open SKILL.md registry for Claude
ClaudSkillsEngineering › Testing › Page 7

Testing (Page 7 of 66)

3955 Claude Code skills in the Testing sub-category of Engineering.

3,955 skills · updated 2026-08-26 · showing 361–420 of 3,955 by quality score

For the full experience including quality scoring and one-click install features for each skill — upgrade to Pro.

Comprehensive UI QA and browser automation using the native monomind browse CDP client — full test lifecycle from discovery through performance profiling, with structured…
Use when evaluating, designing, or pressure-testing the business model of an AI agent product. Triggers on "agent business model", "agent economics", "agent canvas", "evaluate…
Comprehensive documentation of Claude's capabilities for visual regression testing, CI/CD integration, and quality assurance automation.
Use this agent when hands-on code implementation, debugging, testing, technical execution, or build repair is required.
Curated bundle for building, testing, and coordinating AI agent systems. Includes multi-agent orchestration, agent testing, cross-agent handoff, MCP server building, and prompt…
End-to-end testing specialist using Vercel Agent Browser (preferred) with Playwright fallback. Use PROACTIVELY for generating, maintaining, and running E2E tests.
Use when designing evaluations for AI agents, skills, routers, prompts, tool-use policies, or multi-step workflows: task sets, rubrics, graders, hard negatives, regression cases,…
Design and implement evaluation frameworks for AI agents. Use when testing agent reasoning quality, building graders, doing error analysis, or establishing regression pro — from…
Design and implement evaluation frameworks for AI agents. Use when testing agent reasoning quality, building graders, doing error analysis, or establishing regression pro — from…
Evaluate and improve Claude Code commands, skills, and agents. Use when testing prompt effectiveness, validating context engineering choices, or measuring improvement quality.
Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less — from…
Use when benchmarking, replaying, regression-testing, diagnosing, comparing, or release-gating an LLM agent harness across tasks, providers, policies, failures, restarts,…
Observability for Agentforce: adoption, deflection, latency, cost, quality. NOT for agent evaluation/testing (see agentforce-eval-harness) or raw platform-event monitoring.
Test strategy, execution, and coverage analysis. Use when designing tests, running test suites, or analyzing test results beyond baseline checks.
Test strategy, execution, and coverage analysis. Use when designing tests, running test suites, or analyzing test results beyond baseline checks.
Review pull request test coverage quality and completeness, with emphasis on behavioral coverage and real bug prevention.
End-to-end QA testing for Agent2 agents. Starts the agent service, sends real API requests from eval datasets, validates responses against the output schema, and generates a…
Expert QA engineer specializing in comprehensive quality assurance, test strategy, and quality metrics.
Agent davranis testi ve protokol uyumluluk dogrulamasi. Agent'larin tanimli rollerine uygun davranip davranmadigini assertion-based test'lerle olcer.
Generates adversarial test cases targeting safe-failure behavior — refusals, hedging, graceful degradation.
Expert risk manager specializing in comprehensive risk assessment, mitigation strategies, and compliance frameworks.
Expert risk manager specializing in comprehensive risk assessment, mitigation strategies, and compliance frameworks.
Design safe AI agent systems — capability restriction, sandboxed execution, human-in-the-loop gates, anomaly detection, rollback on unexpected behavior, blast radius limiting, and…
Expert Site Reliability Engineer balancing feature velocity with system stability through SLOs, automation, and operational excellence.
Test-Driven Development for the Reactive Agents codebase. Effect-TS aware TDD with mandatory timeout flags, Effect.flip error testing, Layer isolation, and dangling-server…
Test-Driven Development specialist enforcing write-tests-first methodology. Use PROACTIVELY when writing new features, fixing bugs, or refactoring code. Ensures 80%+ test coverage.
Run agent evaluation tests via `sf agent test run` against the target org's Testing Center. Produces severity-graded findings; CI mode emits SARIF for GitHub Code Scanning.
Expert test automation engineer specializing in building robust test frameworks, CI/CD integration, and comprehensive test coverage.
Test agent delegation patterns to verify hierarchy and escalation paths. Use after modifying agent structure.
Use an agent to write tests that check behaviour rather than restate the implementation, and verify they can actually fail. Use when adding coverage to existing code.
Test automation specialist — applies general testing practice and Playwright-specific patterns (TypeScript POM, fixtures, CI, optional Gherkin/BDD).
Use when testing, evaluating, or building regression suites for Agentforce agents: conversation testing in Agent Builder, topic coverage and utterance testing, Testing API and…
Framework de test complet pour agents IA — tests unitaires, d'intégration, de bout en bout et adversariaux.
Test AI agent systems including tool use, multi-turn conversations, error recovery, and non-deterministic outputs.
Drive terminal UI (TUI) applications programmatically for testing, automation, and inspection. Use when: automating CLI/TUI interactions, regression testing terminal apps,…
Optimize ElevenLabs conversational AI agents for real estate applications. Use when creating new agents, improving conversation quality, selecting voices, engineering system…
XXE specialist (H1 #63). Use for testing XML parsing endpoints, file upload processors, SOAP services, SVG handlers, and any feature accepting XML input.
Use when defining or refining the tone, voice, and behavioral personality of an Agentforce agent: system instruction encoding, brand voice alignment, adaptive response formats,…
Design Agentforce testing: topic coverage, action unit tests, deterministic golden sets, adversarial prompts, and regression harness.
Use when implementing code from a spec through a bounded build-test-fix loop with explicit acceptance criteria, automated checks, live verification when needed, and clear stop…
Read-only report of where the agent-generated tests (*.agentic.spec.ts, *AgenticTest.java, test_agentic_*.py, ...) stand — agent-tests-only coverage, goal MET/NOT MET from…
Testing guidelines for Jarvy CLI - unit testing patterns, integration tests with assert_cmd, test environment variables, platform-specific testing, and CI coverage strategies.
Verify that agent tests (*.agentic.spec.ts, *AgenticTest.java, test_agentic_*.py, ...) truly lock behavior by running mutation testing against them only, then strengthening agent…
5-phase AI-powered SDLC — Discovery → Planning → Implementation → Validation → Production. UI-first, test-plan-first, human checkpoints at every phase.
The premier Agent-ready food delivery skill. Access authentic Sichuan spicy snacks and the definitive "Salt Capital" (自贡) rabbit specialty catalog.
Remove every agent-generated test file (*.agentic.spec.*, *.agentic.test.*, *AgenticTest.java, test_agentic_*.py, *_agentic_test.c/.cpp) plus the agentic plan files, after an…
Reconcile agent tests after the user INTENTIONALLY changed main-code behavior — show each failing agent test as old-locked vs new-actual behavior, and update tests only with…
Generate agent-authored unit tests ("agent tests") to reach a user-defined coverage goal for one target file/class or the whole project, without ever modifying main code or…
Use early in an AI-agent project — before ship, before real traffic — to build a starter test suite for the agent and run it offline.
Unterstützt Agentursteuerung, Ausschreibung, Lastenheft und Abnahme barrierefreier Websites. Formuliert Anforderungen, Akzeptanzkriterien, Nachweis- und Re-Test-Pflichten,…
Use when apple-dev has finished feature code + Unit tests (Swift Testing) and is about to enter code-review.
Design, review, and validate TDD/BDD workflows for agile-run app and game stories. Use for test strategy, TDD flows, BDD scenarios, test-light eligibility, app UI/integration…
What counts as evidence in ai-agents and how to produce it. Covers the TESTING-RIGOR pos+neg+edge bar, test layout and collection reality, coverage proof commands,…
PBI INPUT PACKAGE から PlanGate の plan.md / todo.md / test-cases.md を B-1→B-2→B-3 フローで作成する。Use when: docs/working/TASK-XXXX/pbi-input.md を元に実行計画を作りたい時。
Operational prompt engineering for production LLM apps: structured outputs (JSON/schema), deterministic extractors, RAG grounding/citations, tool/agent workflows, prompt safety…
Comprehensive AI prompt engineering safety review and improvement prompt. Analyzes prompts for safety, bias, security vulnerabilities, and effectiveness while providing detailed…
中文优先:用于AI回归测试相关任务,帮助识别、设计、实现或验证对应工作流。English keywords: Regression testing strategies for AI-assisted development.
Use when building, integrating, debugging, or reviewing an AI-powered reporting, dashboard, or data-visualization application where AI generates report structure.
Use after QA strategy and test-case synthesis to build the requirements-to-test traceability matrix, identify missing coverage and test blockers, and score readiness for QA…
Use when QA scope and strategy are defined and you need to generate detailed, executable test cases plus smoke, regression, and user acceptance suites tied to requirements, roles,…
All Engineering skills →
More in EngineeringDevops (3,719) · Architecture (3,060) · Backend (2,477) · Frontend (1,674) · Languages (1,461) · Code Quality (1,434) · Cloud Platforms (1,292) · Databases (890) · Performance (843) · Mobile (630) · Observability (438) · Data Engineering (371) · Docs Engineering (319) · Workflow Orchestration (286) · ML AI Eng (280) · API Tooling (23)