Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 15,457–15,480 of 23,914 skills
Run an auditable reduced recovery experiment for the SBI benchmark by composing generated task, sampling, and C2ST skills.
Execute a bounded soft-mode proxy experiment that validates P3O mechanisms with generated skills and numeric training evidence.
Validate mechanism evidence for adversarial reprogramming proxy or full recovery experiments.
Build and validate harmless corpus splits that mirror visual jailbreak optimization and held-out evaluation protocols.
Compute held-out benign-versus-adversarial obedience deltas and mechanism checks for safe jailbreak proxies.
Run an executable reduced TeCoA recovery experiment with generated skill invocations and mechanism checks.
Build and validate OpenCLIP-style scale tables for CLIP scaling-law experiments, including total compute and error metric derivation.
Build LLaVA-style visual instruction records from captions and boxes for mechanism-faithful recovery experiments.
Run a bounded soft-mode CLIP proxy recovery that logs contrastive training and zero-shot evaluation evidence.
Specify and validate the robust min-max objective used for PGD adversarial training from Madry et al. 2017.
Validate Self-Imitation Learning recovery evidence for source boundaries, executable metrics, and mechanism-faithful proxy checks.
Execute one deterministic reduced Soft Actor-Critic critic actor and target update for recovery evidence.
Validate reduced PPO recovery evidence with mechanism checks, source-boundary checks, and pass-rate metrics.
Evaluate Go-Explore archived trajectories with deterministic replay and bounded perturbation checks for recovery evidence.
Build deterministic periodic PDE benchmarks for PINN failure-mode recovery experiments.
Detect induction-head copying behavior on repeated token sequences with explicit mechanism checks.
Run a bounded mechanism-faithful proxy experiment for FFN value-vector concept promotion using generated skills.
Run a bounded mechanism-faithful proxy experiment for maximum softmax OOD detection recovery artifacts.
Decide whether a declared soft-mode proxy exercises the Accuracy on the Line mechanism.
Run a bounded mechanism-faithful proxy experiment for probabilistic bilevel coreset selection with executable evidence.
Validate paper-target metadata, aggregate influence tolerance, and observed pruning metric gaps for generalization-influence pruning.
Compute EL2N and GraNd example-importance scores for supervised classification pruning experiments.
Compute reduced NLE task rewards and clipping for symbolic recovery experiments.
Validate OPAL recovery evidence, mechanism checks, source boundaries, and proxy metric comparison.