Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,914
skills in category
997
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 15,457–15,480 of 23,914 skills

Benchmark Recovery HarnessA

Run an auditable reduced recovery experiment for the SBI benchmark by composing generated task, sampling, and C2ST skills.

researchgo
0
247
Reduced Recovery HarnessA

Execute a bounded soft-mode proxy experiment that validates P3O mechanisms with generated skills and numeric training evidence.

researchpythongo
0
247
Reprogramming EvaluationA

Validate mechanism evidence for adversarial reprogramming proxy or full recovery experiments.

researchpython
0
247
Safe Corpus ProtocolA

Build and validate harmless corpus splits that mirror visual jailbreak optimization and held-out evaluation protocols.

researchpythongo
0
247
Jailbreak Proxy EvaluatorA

Compute held-out benign-versus-adversarial obedience deltas and mechanism checks for safe jailbreak proxies.

researchpythongo
0
247
Tecoa Recovery HarnessA

Run an executable reduced TeCoA recovery experiment with generated skill invocations and mechanism checks.

researchpythongo
0
247
Openclip Scale ProtocolA

Build and validate OpenCLIP-style scale tables for CLIP scaling-law experiments, including total compute and error metric derivation.

researchpython
0
247
Visual Instruction Data BuilderA

Build LLaVA-style visual instruction records from captions and boxes for mechanism-faithful recovery experiments.

researchpython
0
247
Clip Proxy RecoveryA

Run a bounded soft-mode CLIP proxy recovery that logs contrastive training and zero-shot evaluation evidence.

researchperformance
0
247
Robust Optimization ObjectiveA

Specify and validate the robust min-max objective used for PGD adversarial training from Madry et al. 2017.

researchpythonbash
0
247
Sil Recovery EvaluationA

Validate Self-Imitation Learning recovery evidence for source boundaries, executable metrics, and mechanism-faithful proxy checks.

researchpythonbash
0
247
Sac Update StepA

Execute one deterministic reduced Soft Actor-Critic critic actor and target update for recovery evidence.

researchpython
0
247
Ppo Recovery Evaluation HarnessA

Validate reduced PPO recovery evidence with mechanism checks, source-boundary checks, and pass-rate metrics.

researchpython
0
247
Go Explore Robustification EvaluationA

Evaluate Go-Explore archived trajectories with deterministic replay and bounded perturbation checks for recovery evidence.

researchpythongo
0
247
Periodic Pde BenchmarkA

Build deterministic periodic PDE benchmarks for PINN failure-mode recovery experiments.

researchpythonreact
0
247
Induction Head DetectorA

Detect induction-head copying behavior on repeated token sequences with explicit mechanism checks.

researchgotesting
0
247
Ffn Proxy Recovery HarnessA

Run a bounded mechanism-faithful proxy experiment for FFN value-vector concept promotion using generated skills.

researchpythongit
0
247
Proxy Recovery HarnessA

Run a bounded mechanism-faithful proxy experiment for maximum softmax OOD detection recovery artifacts.

researchgit
0
247
Proxy Recovery EvaluationA

Decide whether a declared soft-mode proxy exercises the Accuracy on the Line mechanism.

researchpython
0
247
Reduced Recovery HarnessA

Run a bounded mechanism-faithful proxy experiment for probabilistic bilevel coreset selection with executable evidence.

researchgo
0
247
Generalization Gap BoundA

Validate paper-target metadata, aggregate influence tolerance, and observed pruning metric gaps for generalization-influence pruning.

researchpython
0
247
Early Example ScoringA

Compute EL2N and GraNd example-importance scores for supervised classification pruning experiments.

researchpythonbash
0
247
Nle Task RewardsA

Compute reduced NLE task rewards and clipping for symbolic recovery experiments.

researchpython
0
247
Opal Recovery EvaluationA

Validate OPAL recovery evidence, mechanism checks, source boundaries, and proxy metric comparison.

researchpythonbash
0
247