Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Yolo Object Detection

ASecurity

Deploy state-of-the-art YOLO models (YOLOv8, YOLOv10, YOLO11) for real-time object detection, instance segmentation, and pose estimation. Triggers when training custom YOLO models, exporting to TensorRT FP16/INT8, running ONNX Runtime inference, executing ByteTrack multi-object tracking, or building high-throughput FastAPI/Triton inference services.

8 stars
0 votes
0 copies
0 views
Added 9/29/2026
ai-agentspythongofastapiapiperformance

Works with

api

Security Analysis

A100/100

Scanned 9/29/2026

$npx -y skills add hamzabellouch/agent-skills --skill yolo-object-detection --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Yolo Object Detection?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Yolo Object Detection
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/hamzabellouch-yolo-object-detection/badge)](https://www.skillsdirectory.com/skills/hamzabellouch-yolo-object-detection)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: yolo-object-detection
metadata:
  category: Computer Vision and Spatial AI
description: >-
  Deploy state-of-the-art YOLO models (YOLOv8, YOLOv10, YOLO11) for real-time object detection, instance segmentation, and pose estimation.
  Triggers when training custom YOLO models, exporting to TensorRT FP16/INT8, running ONNX Runtime inference, executing ByteTrack multi-object tracking,
  or building high-throughput FastAPI/Triton inference services.
compatibility: Python (>= 3.9), Ultralytics (>= 8.1.0), PyTorch (>= 2.1), TensorRT (>= 8.6), ONNX Runtime GPU
---

# YOLO Object Detection & Tracking

End-to-end production pipelines for custom training, TensorRT quantization, multi-object tracking (ByteTrack), and real-time inference serving with YOLO.

---

## 1. Pipeline Architecture

```text
+---------------------+      +------------------------------+      +---------------------------+
| Custom Dataset      | ---> | YOLO Model Training          | ---> | Export to TensorRT        |
| (Roboflow / COCO)   |      | (PyTorch / Ultralytics)      |      | (FP16 / INT8 Calibration) |
+---------------------+      +------------------------------+      +---------------------------+
                                                                                 |
                                                                                 v
+---------------------+      +------------------------------+      +---------------------------+
| Stream Output       | <--- | Real-Time Multi-Object       | <--- | TensorRT Engine           |
| (Bounding Boxes/IDs)|      | Tracking (ByteTrack / BoT)   |      | High-Throughput Inference |
+---------------------+      +------------------------------+      +---------------------------+
```

---

## 2. Custom Dataset Definition & Model Training (`train_yolo.py`)

### Dataset YAML Config (`dataset.yaml`)

```yaml
path: /data/datasets/manufacturing_defects
train: images/train
val: images/val
test: images/test

names:
  0: scratch
  1: dent
  2: crack
```

### PyTorch Training Pipeline (`train_yolo.py`)

```python
from ultralytics import YOLO
import torch
import logging

logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)

def train_custom_yolo():
    device = "cuda:0" if torch.cuda.is_available() else "cpu"
    logger.info(f"Using device: {device}")

    # Load baseline pre-trained YOLO model (e.g. YOLOv8x or YOLO11x)
    model = YOLO("yolov8x.pt")

    # Execute Distributed Training
    results = model.train(
        data="dataset.yaml",
        epochs=100,
        imgsz=640,
        batch=32,
        device=device,
        workers=8,
        optimizer="AdamW",
        lr0=0.001,
        weight_decay=0.0005,
        val=True,
        save=True,
        project="yolo_defects_project",
        name="experiment_v1"
    )

    # Validate trained model
    metrics = model.val()
    logger.info(f"mAP50-95: {metrics.box.map}")
    logger.info(f"mAP50: {metrics.box.map50}")

if __name__ == "__main__":
    train_custom_yolo()
```

---

## 3. TensorRT Export & INT8 Quantization (`export_tensorrt.py`)

Convert PyTorch `.pt` weights into high-performance NVIDIA TensorRT `.engine` models.

```python
from ultralytics import YOLO

def export_to_tensorrt():
    model = YOLO("yolo_defects_project/experiment_v1/weights/best.pt")

    # Export to TensorRT FP16 for maximum GPU throughput
    model.export(
        format="engine",
        imgsz=640,
        half=True,        # Enable FP16 Precision
        dynamic=False,     # Static shape for max performance
        workspace=4,       # 4GB GPU Workspace memory for engine building
        device=0
    )
    print("Successfully exported model to TensorRT Engine format.")

if __name__ == "__main__":
    export_to_tensorrt()
```

---

## 4. Multi-Object Tracking Pipeline with ByteTrack (`track_video.py`)

Combine YOLO object detection with ByteTrack to assign consistent IDs across video frames.

```python
import cv2
from ultralytics import YOLO
import numpy as np

def run_realtime_tracking(video_path: str, engine_path: str):
    # Load TensorRT Engine Model
    model = YOLO(engine_path, task="detect")

    cap = cv2.VideoCapture(video_path)

    while cap.isOpened():
        success, frame = cap.read()
        if not success:
            break

        # Execute Detection & Tracking using ByteTrack
        results = model.track(
            source=frame,
            persist=True,
            tracker="bytetrack.yaml", # Built-in ByteTrack config
            conf=0.4,
            iou=0.5,
            verbose=False
        )

        # Render Bounding Boxes with Track IDs
        annotated_frame = results[0].plot()

        # Extract Tracking Bounding Boxes and Object IDs
        if results[0].boxes and results[0].boxes.id is not None:
            boxes = results[0].boxes.xyxy.cpu().numpy()
            track_ids = results[0].boxes.id.int().cpu().numpy()
            cls_ids = results[0].boxes.cls.int().cpu().numpy()

            for box, track_id, cls_id in zip(boxes, track_ids, cls_ids):
                x1, y1, x2, y2 = map(int, box)
                # Process individual tracked object (e.g. ROI cropping)

        cv2.imshow("Real-Time YOLO ByteTrack", annotated_frame)
        if cv2.waitKey(1) & 0xFF == ord("q"):
            break

    cap.release()
    cv2.destroyAllWindows()

if __name__ == "__main__":
    run_realtime_tracking("input_feed.mp4", "best.engine")
```

---

## 5. FastAPI High-Throughput Batch Serving (`serve_yolo.py`)

```python
from fastapi import FastAPI, UploadFile, File, HTTPException
from ultralytics import YOLO
import cv2
import numpy as np
import io
from PIL import Image

app = FastAPI(title="YOLO Real-Time Inference API", version="1.0.0")

# Load TensorRT Model Engine
MODEL = YOLO("best.engine", task="detect")

@app.post("/detect")
async def detect_objects(file: UploadFile = File(...), confidence: float = 0.5):
    if not file.content_type.startswith("image/"):
        raise HTTPException(status_code=400, detail="Invalid image file format")

    contents = await file.read()
    image = Image.open(io.BytesIO(contents)).convert("RGB")
    frame = np.array(image)

    # Run TensorRT Inference
    results = MODEL.predict(source=frame, conf=confidence, verbose=False)
    
    detections = []
    for box in results[0].boxes:
        coords = box.xyxy[0].tolist()
        conf = float(box.conf[0])
        cls_id = int(box.cls[0])
        label = MODEL.names[cls_id]

        detections.append({
            "label": label,
            "confidence": round(conf, 4),
            "bbox": [round(c, 2) for c in coords]
        })

    return {"count": len(detections), "detections": detections}
```

---

## 6. Optimization Checklist for Production

1. **Resolution Sizing**: Match export resolution (`imgsz=640`) strictly with inference stream dimensions to prevent CPU resize overhead.
2. **TensorRT Engine Reuse**: Cache build engine files (`.engine`); avoid re-building engine files on container restart.
3. **Batching**: Group incoming API image frames into batch sizes of 4, 8, or 16 (`MODEL.predict(source=[f1, f2, f3, f4])`) for higher GPU throughput.

Attribution

hamzabellouchhamzabellouch
View sourceSee grades on GitHubMore from hamzabellouch →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →