
Claude Skills by VectorSpaceLab
github.com/VectorSpaceLabCompute RLAIF AI-labeler alignment, pairwise win rate, harmless rate, and proxy target consistency metrics.
Build RLAIF pairwise AI-feedback labels from label-token logits, optional chain-of-thought rationales, and swapped-order position-bias mitigation.
Train and validate tiny soft-label pairwise reward models that distill RLAIF AI preference distributions with cross-entropy.
Apply ToxiGen ALICE classifier-in-the-loop weighted beam scoring for adversarial or detoxifying generation experiments.
Run a bounded ToxiGen-style classifier update and report validator-compatible before-after loss, AUC, and parameter evidence.
Compute ToxiGen-style human harm labels, classifier attack flags, and aggregate toxicity validation metrics.
Build and validate ToxiGen-style balanced identity-label demonstration prompts for implicit toxicity generation and recovery experiments.
Route latent samples through anchor image realism or patch realism for few-shot adaptation.
Compute cross-domain distance consistency softmax-KL losses for few-shot generator adaptation experiments.
Compute diversity and correspondence diagnostics for few-shot image generation recovery experiments.
Execute a bounded CDC-regularized few-shot adaptation optimizer step with route-aware realism terms.
Compute and inspect a DCL-style contrastive proxy loss with same-latent positives and real-target anti-collapse negatives.
Compute a deterministic intra-cluster diversity proxy mirroring intra-LPIPS evaluation for reduced few-shot generation recovery.
Validate DCL latent-paired source/target feature batches and separate real-target negatives for few-shot GAN adaptation recovery.
Execute a bounded finite-difference optimizer step for DCL proxy recovery and emit validator-compatible training traces.
Compute and test the DDPM epsilon-prediction mean squared error training objective.
Build DDPM beta schedules and closed-form forward noising coefficients for recovery experiments.
Run a bounded soft-mode proxy DDPM recovery experiment using generated module skills.
Compute DDPM predicted x0 and reverse posterior statistics from predicted epsilon.
Build shared noised DDPM adaptation batches and reconstruct clean-image predictions for source and adapted denoisers.
Run a bounded DDPM-PA proxy optimizer step and emit mechanism-faithful recovery traces and metrics.
Extract Haar high-frequency components and compute DDPM-PA high-frequency preservation and detail losses.
Compute DDPM-PA pairwise cosine-softmax distributions and adapted-to-source KL preservation loss.
Compute FID-style activation mean/covariance summaries for generated or real feature matrices.
Assemble a bounded soft-mode recovery protocol that exercises FID statistics, FID distance, and TTUR update skills.
Compute the Fréchet Gaussian distance used by FID with stable numerical handling.
Run deterministic two-time-scale update traces for GAN-style min-max proxy dynamics.
Validate paper-critical guided diffusion UNet architecture flags before claiming architecture-faithful recovery or sampling.
Apply noisy classifier log-probability gradients to reverse diffusion steps with explicit guidance-scale checks.
Build and validate Gaussian diffusion beta schedules and posterior coefficients for guided diffusion recovery experiments.
Preserve guided diffusion target metadata, proxy declarations, classifier scale, and evaluation protocol consistency.
Compute and validate LDM spatial latent-grid sampling plans for high-resolution and dense conditional tasks.
Validate and run deterministic LDM-style cross-attention conditioning between latent features and modality tokens.
Execute and validate a reduced LDM latent noising, epsilon prediction loss, and optimizer-step training primitive.
Validate LDM autoencoder latent compression contracts, spatial downsampling, regularization metadata, and reconstruction diagnostics.
Build executable soft-mode LDM proxy recovery artifacts from generated skills with mechanism checks and validator evidence.
Evaluate BAPPS-style two-alternative forced choice perceptual similarity triplets with auditable per-item predictions.
Fit non-negative layer calibration weights for LPIPS-style distances using bounded 2AFC ranking feedback.
Compute LPIPS-style perceptual patch distances from normalized multi-layer feature differences with optional non-negative calibration weights.
Run a bounded source-safe LPIPS-style recovery experiment and emit validation-compatible recovery artifacts.
Compute CIC normalized skill-transition contrastive loss, logits, margins, and tiny dependency-free training updates.
Estimate CIC intrinsic reward using k-nearest-neighbor particle entropy over learned transition embeddings.
Run a bounded soft-mode CIC proxy experiment that invokes generated batching, contrastive loss, and entropy reward skills.
Build and validate aligned CIC state-transition and continuous skill batches for contrastive intrinsic control recovery experiments.
Classify offline RL datasets by D4RL challenge properties and explain benchmark implications from metadata.
Compute and audit D4RL normalized returns from raw, random-baseline, and expert-reference scores.
Validate fixed offline reinforcement learning transition datasets and summarize D4RL-style provenance before recovery experiments.
Execute a bounded D4RL-style offline recovery proxy with fixed-dataset training and normalized-score evidence.
Compute DIAYN discriminator log-probabilities, cross-entropy diagnostics, and intrinsic rewards from skill-conditioned state predictions.
Execute a bounded maximum-entropy DIAYN policy surrogate update with before-after loss and parameter-change evidence.