Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,827
skills in category
868
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 6,865–6,888 of 20,827 skills

Iclr PointsA

Quantifies the average research effort required to produce one publication at top-tier conferences across 27 computer science subfields. It enables cross-area comparisons of faculty productivity and publication effort by normalizing faculty headcounts against publication counts. Use when the user has predictions and gold and needs to compute ICLR points.

researchpythongo
0
3
Ici Response Prediction EvalA

Evaluates the cross-cohort generalisability of transcriptomic models (bulk and single-cell RNA-seq) for predicting immune checkpoint inhibitor (ICI) response in cancer patients. Probes robustness to cohort-specific transcriptomic context, tumour type, immune composition, and class imbalance. Use when the user wants to benchmark on Cho et al., Ribas et al., Poddubskaya et al., Gondal et al., Franken et al., Luoma et al., Reinstein et al., or asks about evaluating this task. Reports macro F1 sc...

researchpythongo
0
3
Ice Guard EvalA

Measures intervention consistency in LLM decision-making by checking whether swapping irrelevant features (demographic names, authority credentials, or framing phrasing) causes the model to change its verdict. Probes susceptibility to spurious feature reliance and systematic bias across high-stakes domains. Use when the user wants to benchmark on ICE-Guard Benchmark, or asks about evaluating this task. Reports flip_rate.

researchpython
0
3
Ice Flare EvalA

Evaluates bilingual (Chinese and English) financial large language models across 14 NLP tasks, including sentiment analysis, classification, question answering, and information extraction. It probes cross-lingual adaptability, domain-specific reasoning, and instruction-following capabilities on financial text. Use when the user wants to benchmark on FE, StockB, CFPB, CFiQA-SA, FPB, FiQA-SA, Corpus, AFQMC, NL, NL2, NSP, FinevalF, StcokA, CACL18, CBigData18, CIKM18, ACL18, BigData18, RE, CHeadl...

researchpythongo
0
3
Ice Bench EvalA

Evaluates image generation and editing models across 31 fine-grained tasks spanning text-to-image creation, reference-guided creation, and various editing scenarios. It probes capabilities in aesthetic quality, imaging quality, prompt adherence, source/reference consistency, and controllability. Use when the user wants to benchmark on ICE-Bench, or asks about evaluating this task. Reports prompt following (PF).

researchpythonperformance
0
3
Icdar2019 Sroie EvalA

Evaluates end-to-end document understanding on low-quality scanned receipts, specifically testing text localization, character-level OCR, and structured key information extraction (e.g., company, cash, date, address). Use when the user wants to benchmark on ICDAR2019 SROIE, or asks about evaluating this task. Reports primary metric.

researchpythongo
0
3
Iccma17 EvalA

Evaluates computational argumentation solvers on their ability to correctly compute extensions (e.g., semi-stable, stage, ideal) across diverse argumentation frameworks ranging from random graphs to application-derived structures. Use when the user wants to benchmark on ICCMA'17 Benchmark Suite, or asks about evaluating this task. Reports exact-match accuracy.

researchpythongo
0
3
Ic IndexA

Evaluates whether machine learning models can correctly capture non-additive interaction effects in drug-target affinity prediction, rather than merely learning global means or individual drug/target main effects. It measures the proportion of correctly predicted interaction directions across test pairs. Use when the user has predictions and gold and needs to compute IC-index.

researchpythongo
0
3
Iao Prompting EvalA

Evaluates LLMs' ability to perform structured reasoning and knowledge application across arithmetic, logical, commonsense, and symbolic tasks using a template-based prompting framework. Use when the user wants to benchmark on GSM8K, AQuA, Date Understanding, Object Tracking, StrategyQA, CommonsenseQA, Last Letter, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
I2e Event Classification EvalA

Evaluates the classification performance of Spiking Neural Networks trained on synthetic event streams generated from static images, and tests the transferability of these models to real-world neuromorphic sensor data. Use when the user wants to benchmark on I2E-CIFAR10, I2E-CIFAR100, I2E-ImageNet, CIFAR10-DVS, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
I Star EvalA

This evaluation probes how anisotropic regularization (I-STAR) affects the downstream performance of fine-tuned language models across standard NLP benchmarks. It also measures the geometric properties of the resulting embedding spaces, specifically isotropy and intrinsic dimensionality, to correlate representation structure with task accuracy. Use when the user wants to benchmark on SST-2, QNLI, RTE, MRPC, QQP, COLA, STS-B, SST-5, SQUAD, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Hyrec Cmteb EvalA

Evaluates the retrieval capability of hybrid (dense + sparse/lexicon) models on Chinese text. It probes how well the model ranks relevant passages for a given query across diverse Chinese domains like medical, e-commerce, and general web. Use when the user wants to benchmark on C-MTEB, or asks about evaluating this task. Reports nDCG@10.

researchpythonperformance
0
3
Hypothesis Ranking EvalA

Measures an LLM's ability to correctly rank a groundtruth hypothesis against a set of negative hypotheses using pairwise comparisons. It evaluates discriminative judgment in scientific reasoning. Use when the user wants to benchmark on ResearchBench Hypothesis Ranking, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Hypothesis Composition EvalA

Assesses an LLM's capability to synthesize novel research hypotheses by combining a given research background with retrieved inspiration papers. It probes the model's ability to mutate and recombine scientific concepts into coherent, groundtruth-aligned proposals. Use when the user wants to benchmark on ResearchBench Hypothesis Composition, or asks about evaluating this task. Reports Normalized Composition Score.

researchpythongo
0
3
Hypocenter Inversion EvalA

Evaluates the accuracy and uncertainty quantification of a physics-informed neural network for locating earthquake hypocenters using synthetic seismic arrival times. Use when the user wants to benchmark on Synthetic Seismic Array, or asks about evaluating this task. Reports location uncertainty.

researchpythongo
0
3
Hyperjump EvalA

Evaluates the optimization quality and time efficiency of hyperparameter search algorithms by comparing the test error rate of recommended configurations against wall-clock time across neural architecture and traditional ML benchmarks. It measures how quickly each optimizer converges to near-optimal configurations under sequential and parallel deployment settings. Use when the user wants to benchmark on NATS-Bench, LIBSVM Covertype, or asks about evaluating this task. Reports test_error_rate.

researchpythongo
0
3
Hyperhelm EvalA

Evaluates the ability of mRNA language models to predict diverse biological properties (e.g., protein expression, degradation, thermostability) and annotate antibody sequence regions. It also probes model robustness to out-of-distribution sequence lengths and extreme GC content, testing generalization in hierarchical biological representation learning. Use when the user wants to benchmark on Ab1, Ab2, mRFP, COVID-19 Vaccine, Drosophila melanogaster, Saccharomyces cerevisiae, Pichia pastoris, ...

researchpythongo
0
3
Hyperglm EvalA

Evaluates a multimodal model's ability to generate and anticipate scene graphs from video frames, capturing spatial object relationships and causal temporal transitions. It also tests video question answering, captioning, and relation reasoning capabilities by leveraging hypergraph structures to model multi-way interactions. Use when the user wants to benchmark on VSGR, PVSG, Action Genome, or asks about evaluating this task. Reports Recall (R) / mean Recall (mR).

researchpythongo
0
3
Hyperfm250k EvalA

Evaluates hyperspectral foundation models and task-specific deep learning architectures on pixel-level regression tasks for retrieving cloud optical and microphysical properties (COT, CER, CWP, CTH) from NASA PACE-OCI imagery. Use when the user wants to benchmark on HyperFM250K, or asks about evaluating this task. Reports MSE.

researchpythongit
0
3
Hyface Vc EvalA

This protocol evaluates face-based voice conversion models by measuring how well synthesized audio matches the target speaker's identity and pitch characteristics using only facial images as input. It probes cross-modal alignment, speaker homogeneity, diversity, and explicit fundamental frequency (F0) estimation accuracy. Use when the user wants to benchmark on LRS3, or asks about evaluating this task. Reports Pitch deviation.

researchpythongo
0
3
Hydraulic Anomaly Detection EvalA

Evaluates semi-supervised and traditional machine learning models for detecting hydraulic system anomalies using only normal data for training. It probes the ability of models to generalize from normal-condition features and identify leakage faults under class-imbalanced testing conditions. Use when the user wants to benchmark on Unspecified hydraulic condition monitoring dataset, or asks about evaluating this task. Reports F1_Score.

researchpythontesting
0
3
Hybridrag Bench EvalA

Evaluates retrieval-augmented models' ability to perform multi-hop reasoning over hybrid knowledge (unstructured text and knowledge graphs) using time-framed, external scientific literature to prevent parametric memorization. Use when the user wants to benchmark on Arxiv-AI, Arxiv-CY, Arxiv-BIO, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Hybridqa EvalA

Multi-hop question answering that requires integrating information from both tabular and textual sources. It probes a model's ability to perform cross-modal reasoning and extract precise answers from heterogeneous data. Use when the user wants to benchmark on HybridQA, or asks about evaluating this task. Reports exact match (EM).

researchpythongo
0
3
Hybridna EvalA

Evaluates the capability of DNA foundation models to perform short-range and long-range genomic understanding tasks, as well as their ability to generate biologically plausible cis-regulatory elements. It probes sequence classification, variant effect prediction, and generative design across multiple species and cell types. Use when the user wants to benchmark on GUE, BEND, LRB, CRE (regLM), or asks about evaluating this task. Reports MCC.

researchpythongit
0
3