Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Ab Test Framework

ASecurity

Compare models with A/B testing for selection

14 stars
0 votes
0 copies
6 views
Added 9/7/2026
testingjavascriptgojavatestingsecurityperformance

Security Analysis

A92/100
mediumInstalls packages at runtime which could introduce malicious dependencies

Pro scans all 4 files and shows the line behind each finding

Scanned 9/7/2026

$npx -y skills add modbender/skill-library-mcp --skill ab-test-framework --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Ab Test Framework?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Ab Test Framework
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/modbender-ab-test-framework/badge)](https://www.skillsdirectory.com/skills/modbender-ab-test-framework)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: A/B Testing Framework
version: 1.0.0
description: Compare models with A/B testing for selection
author: OpenClaw SysAdmin Community
tags: [20, sysadmin, manual]
trigger:
  type: manual
  schedule: ""
  enabled: true
input:
  model_a:
    type: string
    description: First model
    required: True
  model_b:
    type: string
    description: Second model
    required: True
  test_prompts:
    type: array
    description: Test prompts
    required: True
output:
  status: string
  details: object
  winner: string
  confidence: number
dependencies:
  - openclaw/llm
  - stats-library
security:
  - A/B test security per Category 8; prevent test manipulation
  - Validate all inputs per Category 8
  - Use least privilege principle
  - Log all operations for audit
---

# A/B Testing Framework

## Description

Compare models with A/B testing for selection

## Source Reference

This skill is derived from **20. Testing & Quality Assurance** of the OpenClaw Agent Mastery Index v4.1.

**Sub-heading**: A/B Testing Frameworks for Model Selection

**Complexity**: high

## Input Parameters

| Name | Type | Required | Description |
|------|------|----------|-------------|
| `model_a` | string | Yes | First model |
| `model_b` | string | Yes | Second model |
| `test_prompts` | array | Yes | Test prompts |

## Output Format

```json
{
  "status": <string>,
  "details": <object>,
  "winner": <string>,
  "confidence": <number>
}
```

## Usage Examples

### Example 1: Basic Usage

```javascript
const result = await openclaw.skill.run('ab-test-framework', {
  model_a: "value",
  model_b: "value",
  test_prompts: 123
});
```

### Example 2: With Optional Parameters

```javascript
const result = await openclaw.skill.run('ab-test-framework', {
  model_a: "value",
  model_b: "value",
  test_prompts: []
});
```

## Security Considerations

A/B test security per Category 8; prevent test manipulation

### Additional Security Measures

1. **Input Validation**: All inputs are validated before processing
2. **Least Privilege**: Operations run with minimal required permissions
3. **Audit Logging**: All actions are logged for security review
4. **Error Handling**: Errors are sanitized before returning to caller

## Troubleshooting

### Common Issues

| Issue | Cause | Solution |
|-------|-------|----------|
| Permission denied | Insufficient privileges | Check file/directory permissions |
| Invalid input | Malformed parameters | Validate input format |
| Dependency missing | Required module not installed | Run `npm install` |

### Debug Mode

Enable debug logging:
```javascript
openclaw.logger.setLevel('debug');
const result = await openclaw.skill.run('ab-test-framework', { ... });
```

## Related Skills

- `model-routing-manager`
- `performance-benchmarker`
 * @param {string} params.model_a - First model
 * @param {string} params.model_b - Second model
 * @param {Array} params.test_prompts - Test prompts

Attribution

modbendermodbender
View sourceSee grades on GitHubMore from modbender →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Screen Reader Testing

Practical guide to testing web applications with screen readers for comprehensive accessibility validation.

401991 votes

Tdd Workflow

在编写新功能、修复错误或重构代码时使用此技能。强制执行测试驱动开发,包含单元测试、集成测试和端到端测试,覆盖率超过80%。

2456590 votes

Eval Harness

克劳德代码会话的正式评估框架,实施评估驱动开发(EDD)原则

2456590 votes

Python Testing

使用pytest、TDD方法、夹具、模拟、参数化和覆盖率要求的Python测试策略。

2456590 votes

Django Tdd

Django测试策略,包括pytest-django、TDD方法论、factory_boy、模拟、覆盖率以及测试Django REST Framework API。

2456590 votes
View all in testing →