Skip to content
Back to skills

Webgpu Performance

ASecurity

Use when High-performance browser graphics and compute mastery. Transitioning from WebGL to WebGPU API, WGSL compute shaders, explicit GPU memory management, and browser-side tensor calculations.

  • 5 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 27, 2026
ai-agentstypescriptbashapiperformance

Works with

  • api

Security analysis

A100/100

Scanned September 27, 2026

npx -y skills add Harmitx7/tribunal-kit --skill webgpu-performance --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Webgpu Performance?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Webgpu Performance
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/harmitx7-webgpu-performance-tribunal-kit/badge)](https://www.skillsdirectory.com/skills/harmitx7-webgpu-performance-tribunal-kit)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: webgpu-performance
description: "Use when High-performance browser graphics and compute mastery. Transitioning from WebGL to WebGPU API, WGSL compute shaders, explicit GPU memory management, and browser-side tensor calculations."
version: 5.0.0
last-updated: 2026-09-13
skills:
  - 60fps-animation
  - fixing-motion-performance
  - motion-engineering
tools: Read, Grep, Glob, Bash, Edit, Write
scripts-binding:
  - .agent/scripts/lint_runner.js
  - .agent/scripts/verify_all.js
---

# WebGPU Performance Mastery

---

## 🛠️ Technical Architecture & Reference Recipes

---

## 1. Core Principles

- **Explicit > Implicit:** Unlike WebGL, WebGPU doesn't hide state. You must explicitly configure Pipelines, BindGroups, and CommandEncoders.
- **Compute First:** Leverage Compute Shaders (`@compute @workgroup_size(X, Y)`) for heavy array manipulation, physics, or ML tensor operations, keeping the CPU entirely free.
- **Buffer Alignment:** WGSL requires strict 4-byte or 16-byte alignment (`vec4<f32>`, `mat4x4<f32>`). Always pad structs exactly to prevent silent memory corruption.

## 2. WGSL Compute Shader Pattern

When performing parallel calculations (e.g., particle physics or ML matrix multiplication):

```wgsl
struct SystemData {
    deltaTime: f32,
    particleCount: u32,
}

@group(0) @binding(0) var<uniform> data: SystemData;
@group(0) @binding(1) var<storage, read_write> particles: array<vec4<f32>>;

@compute @workgroup_size(64)
fn main(@builtin(global_invocation_id) global_id: vec3<u32>) {
    let index = global_id.x;
    if (index >= data.particleCount) { return; }

    var pos = particles[index];
    pos.y -= 9.8 * data.deltaTime; // Gravity
    particles[index] = pos;
}
```

## 3. WebGPU Execution Pipeline

To run the above compute shader from TypeScript:

1. **Initialize:** `navigator.gpu.requestAdapter()` -> `requestDevice()`.
2. **Create Buffers:** `device.createBuffer({ size, usage: GPUBufferUsage.STORAGE | GPUBufferUsage.COPY_DST })`.
3. **Write Data:** `device.queue.writeBuffer(buffer, 0, float32Array)`.
4. **Bind Group:** Group buffers into a `GPUBindGroup`.
5. **Command Encoder:**
   ```typescript
   const encoder = device.createCommandEncoder();
   const pass = encoder.beginComputePass();
   pass.setPipeline(computePipeline);
   pass.setBindGroup(0, bindGroup);
   pass.dispatchWorkgroups(Math.ceil(count / 64));
   pass.end();
   device.queue.submit([encoder.finish()]);
   ```

## 4. LLM Traps & Pre-Flight Checks

- **TRAP:** Assuming WebGPU works everywhere.
- **FIX:** Always feature-detect with `if (!navigator.gpu) { fallbackToWebGL(); }`.
- **TRAP:** Struct alignment issues in WGSL.
- **FIX:** Never use `vec3<f32>` inside arrays without padding. It behaves as 16-bytes anyway. Use `vec4<f32>` to be perfectly aligned.
- **TRAP:** Reading buffers back to the CPU synchronoulsy.
- **FIX:** Use `mapAsync(GPUMapMode.READ)` and await it. Do not block the main thread.

## Verification Protocol

Before submitting code, ensure:

1. Devices and adapters are properly null-checked.
2. WGSL workgroup sizes align with the dispatch sizes dynamically.
3. GPUBuffers used for compute have `GPUBufferUsage.STORAGE` flags.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…