
Claude Skills by VectorSpaceLab
github.com/VectorSpaceLabCompare robustness interventions across clean and shifted metrics using direction-aware evidence signs.
Assemble executable soft-mode proxy recovery results with metrics, traces, and mechanism checks.
Validate clean and shifted robustness benchmark protocols with class-overlap and metric-gap contracts.
Compute AUROC and AUPR detection metrics with paper-faithful score orientations for softmax confidence baselines.
Run a bounded mechanism-faithful proxy experiment for maximum softmax OOD detection recovery artifacts.
Compute maximum softmax probability detector scores from logits or probabilities for misclassification and OOD detection.
Preprocess English text with WordNet-style collocation matching, tokenization, and inflectional normalization.
Select or defer WordNet-style sense pointers using current and previous sentence context.
Build and validate compact WordNet-style synset and semantic-pointer taxonomies for recovery experiments.
Estimate WordNet-style semantic distance by traversing typed semantic-pointer graphs.
Run a reduced executable WordNet taxonomy recovery using generated preprocessing, tagging, and distance skills.
Compute Direct Preference Optimization logistic losses, implicit rewards, and log-ratio diagnostics from policy and reference log probabilities.
Normalize pairwise preference records into prompt, chosen, rejected, and metadata fields for Direct Preference Optimization training.
Run a bounded scalar DPO optimization that records loss, parameter changes, preference accuracy, and mechanism checks for reduced recovery.
Compute response-only sequence log probabilities for DPO by masking prompts and padding before summing token log probabilities.
Measure whether activated FFN neurons preferentially raise logits for their projected promoted tokens versus controls.
Group top promoted vocabulary tokens into human-readable concept labels with purity and unresolved-token evidence.
Run a bounded mechanism-faithful proxy experiment for FFN value-vector concept promotion using generated skills.
Extract and normalize feed-forward output value vectors from transformer-like model weights for concept-promotion analysis.
Project FFN value vectors through an LM head to rank promoted tokens for mechanistic concept analysis.
Compute PPLM-style bag-of-words or linear-classifier attribute losses and gradients for controlled generation.
Evaluate PPLM controlled-generation proxies with target-mass gain, KL fluency cost, and target consistency checks.
Fuse PPLM perturbed and unperturbed token distributions using geometric mixing for fluent controlled decoding.
Run PPLM-style iterative normalized gradient perturbations with KL regularization and optimizer trace evidence.
Create bounded multi-continuation records per prompt for toxic degeneration evaluation with explicit generator metadata.
Normalize RealToxicityPrompts-style prompt records for toxic and non-toxic prompted generation evaluation.
Run an executable bounded recovery harness that composes prompt normalization, generation, scoring, and aggregation with mechanism checks.
Compute expected maximum toxicity and toxicity probability for RealToxicityPrompts-style continuation sets.
Attach numeric toxicity scores to generated continuations using Perspective API or declared offline proxy scoring.
Compute frozen-attention direct, head, and virtual-head path contributions for attention-only circuits.
Build and validate bounded mechanism-faithful recovery artifacts for Transformer Circuits proxy experiments.
Detect induction-head copying behavior on repeated token sequences with explicit mechanism checks.
Expand attention-head QK and OV weights into token-level circuit matrices and copying diagnostics.
Apply logit-lens unembedding and additive residual contribution checks for mechanistic transformer circuit analysis.
Generate coefficient schedules for curriculum-regularized PINN recovery experiments.
Diagnose PINN failure-mode recovery runs with relative errors and optimizer progress checks.
Build deterministic periodic PDE benchmarks for PINN failure-mode recovery experiments.
Compute PINN data, boundary, and PDE residual losses for reduced failure-mode experiments.
Run bounded reduced PINN recovery experiments that exercise generated benchmark, objective, diagnostics, and curriculum skills.
Build smooth residual trial networks for Deep Ritz variational PDE objectives.
Generate stochastic interior and boundary quadrature samples for Deep Ritz variational PDE training.
Run bounded Deep Ritz optimization and emit auditable recovery artifacts for variational PDE experiments.
Compute Deep Ritz variational energy losses with gradient energy and boundary penalties.
Build a lightweight gated PINN-style scalar model that exposes trainable parameters and gradient paths for reduced recovery experiments.
Update PINN data-loss weights from residual and data-fit gradient statistics using the paper's moving-average annealing rule.
Construct deterministic reduced Helmholtz PINN benchmark data with analytic solution, boundary samples, collocation samples, forcing, and relative L2 scoring.
Compute separated Helmholtz PINN residual and boundary losses so gradient imbalance can be diagnosed and corrected.
Run bounded reduced Helmholtz PINN recovery with optimizer updates, lambda traces, relative L2 metrics, and validator-compatible evidence.
Maintain bounded positive-curvature L-BFGS correction memory for large-scale quasi-Newton optimization.
Build validator-compatible soft-mode recovery metrics and mechanism checks for L-BFGS proxy experiments.