Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 20,617–20,640 of 22,620 skills
Systematic quality assessment using BetterBench 46-criterion framework — 5 benchmarks, 30 papers, 40 web searches
Evaluation Methodology Archaeology Campaign — 5 strategies for systematic analysis of AI/ML benchmarks, metrics, and leaderboards. Reveals construct validity issues, saturation, data contamination, and evaluation protocol inconsistencies.
Produce final structured baseline report integrating all analysis results
Select appropriate baselines for experimental comparison
Quality control metrics for ChIP-seq experiments including FRiP, NSC/RSC, IDR, and library complexity measurements to assess enrichment quality and replicate reproducibility.
SOTA Performance Baseline Campaign — 5 strategies for systematically collecting, standardizing, and analyzing performance data across methods. Produces standardized comparison tables, progress curves, and headroom analysis.
Gather independent ranking ballots from multiple judges or perspectives for a given candidate set.
Identify and suspend fundamental assumptions via de Bono PO. Systematically negate axioms to reveal hidden solution spaces.
Generate specific attack strategies for a given threat surface, producing concrete probes that can be executed.
Compute overall resilience score (0.0-1.0) based on attack results, coverage, and vulnerability severity distribution.
Systematic stress testing of assumptions — surface, classify by vulnerability, attack, assess fragility. Combines assumption-surfacing (shared), abp-vulnerability-classification, and clr-validation SOPs.
Classic reductio ad absurdum: negate the core claim, derive logical consequences, seek contradiction or absurdity.
Systematic extraction, challenge, and sensitivity analysis of assumptions underlying a decision to identify load-bearing beliefs.
Assumption Destruction Campaign — open new solution spaces by negating, reversing, and challenging fundamental assumptions.
Measure how much conclusions change when each assumption is negated. Ranks assumptions by their impact on the final result.
Which assumptions are most fragile? — Vulnerability ranking + impact assessment of experiment assumptions
Challenge each assumption's validity — shared cross-repo SOP
Tactic: Surface assumptions, sort by dependency, attack root assumptions first, then trace cascade failures through the dependency graph.
Build assumption dependency graphs and trace cascade failures when root assumptions are invalidated.
Surface all assumptions, classify by vulnerability (load-bearing × likely-false), validate causal logic. Focus on dangerous assumptions — high load-bearing + non-explicit.
Standardize assignee names and identify corporate group affiliations across patent offices
Rate each identified obstacle's difficulty — overcomability, time cost, workaround existence. May optionally use search tools to validate assessments.
Present obstacles with their severity assessments and proposed mitigations to the user. Ask whether they can accept these obstacles. If unacceptable after 2 rounds, return to present-candidates.
Deep WHY probing inspired by i* Intentionality modeling. Understand the user's motivation, success definition, risk tolerance, innovation preference, independence preference, time urgency, and learning willingness. The most important SOP in actor-profiling — understanding WHY drives everything downstream.