Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,835
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,033–7,056 of 20,835 skills

Health Indicator EvalA

Evaluates the ability of unsupervised contrastive learning models to extract robust, degradation-sensitive health indicators from sensor data. It probes how well the learned features correlate with actual wear or track operational degradation over time, while remaining invariant to noise and operating condition shifts. Use when the user wants to benchmark on Milling Machine Wear Dataset, Railway Wheel Dataset, or asks about evaluating this task. Reports correlation value to the wear.

researchpythongo
0
3
Headlinecause EvalA

Evaluates a model's ability to detect implicit causal relationships between two news headlines without relying on explicit causal linking words. It probes commonsense reasoning and world knowledge to distinguish between causal, refutational, same-event, and unrelated headline pairs. Use when the user wants to benchmark on HeadlineCause, or asks about evaluating this task. Reports causality ROC AUC.

researchpythongo
0
3
Head Qa EvalA

Evaluates complex reasoning and domain-specific knowledge integration in healthcare by testing models on multi-choice questions derived from real Spanish medical specialization exams. It probes the ability to handle long, context-rich questions requiring cross-domain inference and precise medical knowledge. Use when the user wants to benchmark on HEAD-QA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Head Ct Radiology Classification EvalA

This benchmark evaluates a model's ability to perform multi-label classification on radiological text reports, specifically identifying the presence of 13 clinical findings in head CT scans. It probes the model's capacity to handle significant class imbalance and generalize from general-domain pretraining to specialized medical NLP tasks. Use when the user wants to benchmark on Head CT Reports, or asks about evaluating this task. Reports Sample-weighted F1-score.

researchpython
0
3
Hdr Gopro EvalA

Evaluates the model's ability to reconstruct high-quality dynamic HDR radiance fields and synthesize novel views and time steps from alternating-exposure monocular videos. It probes exposure-invariant geometric reconstruction, temporal coherence, and radiometric accuracy under extreme exposure variations. Use when the user wants to benchmark on HDR-GoPro, or asks about evaluating this task. Reports PSNR.

researchpythongo
0
3
Hdp Real Vehicle EvalA

Evaluates end-to-end autonomous driving planning models in real-world closed-loop scenarios, measuring success rate, trajectory stability, and safety compliance during urban driving. Use when the user wants to benchmark on Real-world driving dataset, or asks about evaluating this task. Reports closed-loop success rate.

researchpythonperformance
0
3
Hd209458b Retrieval EvalA

Evaluates atmospheric retrieval models on exoplanet transmission spectra to constrain chemical abundances, temperature-pressure profiles, and cloud properties. It probes the model's ability to disentangle spectral features across multi-wavelength observations and quantify detection significance of trace gases. Use when the user wants to benchmark on JWST NIRCam transmission spectra, HST WFC3 transmission spectra, HST STIS transmission spectra, or asks about evaluating this task. Reports reduc...

researchpythongo
0
3
Hc3 Human EvalA

Evaluates the ability to distinguish AI-generated responses from human expert answers and assesses perceived helpfulness across multiple domains. It probes linguistic realism, factual reliability, and stylistic alignment with human communication. Use when the user wants to benchmark on HC3, or asks about evaluating this task. Reports detection accuracy.

researchpythongo
0
3
Hazard EvalA

Evaluates embodied agents' decision-making and planning capabilities in dynamically changing disaster environments (fire, flood, wind). It probes the ability to reason about evolving object states, environmental propagation dynamics, and spatial-temporal trade-offs to successfully rescue valuable items. Use when the user wants to benchmark on HAZARD, or asks about evaluating this task. Reports rescued value rate (Value).

researchpythongo
0
3
Hateful Memes EvalA

Evaluates multimodal models' ability to detect hateful or harmful memes by analyzing the alignment between image and text content. It probes robustness against visual and textual confounders that appear benign individually but become harmful when combined. Use when the user wants to benchmark on HatefulMemes, HarMeme, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Hateful Meme Detection EvalA

Evaluates multimodal models' ability to detect hateful or offensive memes across multiple domains and under low-resource, out-of-distribution conditions. It probes robustness to distribution shifts, adversarial image perturbations, and the effectiveness of retrieval-augmented inference versus standard fine-tuning or in-context learning. Use when the user wants to benchmark on HatefulMemes, HarMeme, MAMI, Harm-P, MultiOFF, PrideMM, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Hate Speech Ordinal EvalA

Evaluates deep learning models' ability to predict continuous, interval-scaled hate speech scores from raw text comments. It benchmarks against existing APIs and transformer baselines using cross-validated error and correlation metrics. Use when the user wants to benchmark on Custom hate speech corpus (YouTube, Reddit, Twitter), or asks about evaluating this task. Reports RMSE.

researchpythongo
0
3
Hate Speech Offensive Language EvalA

Evaluates a model's ability to distinguish between hate speech, offensive language, and neutral text in social media posts. It probes the classifier's sensitivity to contextual nuances, reclaimed slurs, and demographic-specific biases in labeling. Use when the user wants to benchmark on Hate Speech and Offensive Language Dataset, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Hate Speech Detection EvalA

Evaluates binary and multi-label hate speech detection models on Brazilian Portuguese text. It specifically probes a model's sensitivity to targeted minority groups and its ranking quality under severe class imbalance. Use when the user wants to benchmark on ToxiGen-PT, Portuguese Superset Benchmark, HateBR, OLID-BR, TuPy-E, ToLD-BR, or asks about evaluating this task. Reports Macro-Recall.

researchpython
0
3
Hasper EvalA

Evaluates the ability of computer vision models to classify hand shadow puppet silhouettes into one of 15 distinct categories. It probes feature extraction robustness, particularly for rotationally asymmetric and visually similar silhouettes under varying lighting and motion dynamics. Use when the user wants to benchmark on HaSPeR, or asks about evaluating this task. Reports top-1 accuracy.

researchpythongo
0
3
Harvard Eye Fairness EvalA

Evaluates deep learning models for eye disease screening (AMD, DR, glaucoma) on 2D fundus and 3D OCT images. It measures both overall diagnostic performance and demographic fairness across race, gender, and ethnicity to assess equitable model behavior. Use when the user wants to benchmark on Harvard-EF30k, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Harrison EvalA

This benchmark evaluates a model's ability to recommend relevant hashtags for real-world social media images using only visual input. It probes contextual image understanding and multi-label classification by measuring how well predicted hashtags align with actual user-generated tags. Use when the user wants to benchmark on HARRISON, or asks about evaluating this task. Reports Precision@1.

researchpythongo
0
3
Harood Ood EvalA

Evaluates out-of-distribution (OOD) generalization in sensor-based human activity recognition (HAR) across four realistic domain-shift scenarios: cross-person, cross-position, cross-device, and cross-time. It benchmarks how well 16 OOD methods with CNN or Transformer backbones maintain performance when trained on one domain and tested on unseen domains. Use when the user wants to benchmark on DSADS, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Harmonysset EvalA

Evaluates multimodal large language models' ability to align video and music across four dimensions: rhythmic synchronization, thematic coherence, emotional congruence, and cultural relevance. It tests both open-ended descriptive reasoning and multiple-choice selection to measure temporal and semantic alignment capabilities. Use when the user wants to benchmark on HarmonySet, or asks about evaluating this task. Reports HarmonySet-OE Score.

researchpythongo
0
3
Harmbench Asr EvalA

Evaluates the robustness of LLM safety defenses against multi-turn human and automated jailbreak attacks. It probes whether current refusal mechanisms and machine unlearning methods can withstand adversarial red teaming aimed at recovering harmful or dual-use knowledge. Use when the user wants to benchmark on HarmBench, WMDP-Bio, or asks about evaluating this task. Reports ASR.

researchpythonsecurity
0
3
HareA

Evaluates the clinical quality and diagnostic alignment of machine-generated histopathology reports by measuring semantic alignment of extracted pathological entities and their interrelations against ground truth reports. It probes a model's ability to accurately capture domain-specific terminology, diagnostic conclusions, and their contextual connections. Use when the user has predictions and gold and needs to compute HARE Score.

researchpythongo
0
3
Hardvs2.0 EvalA

Evaluates multi-modal human activity recognition capabilities by classifying 300 action categories from synchronized RGB frames and asynchronous event streams under challenging real-world conditions such as low light, occlusion, and dynamic backgrounds. Use when the user wants to benchmark on HARDVS 2.0, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Har Weighted F1 EvalA

Evaluates the ability of lightweight convolutional neural networks to accurately classify human activities from wearable sensor time-series data. It probes the trade-off between model compression (parameter count and FLOPs) and classification performance on highly imbalanced, multi-class activity recognition tasks. Use when the user wants to benchmark on UCI-HAR, OPPORTUNITY, PAMAP2, UNIMIB-SHAR, WISDM, or asks about evaluating this task. Reports weighted F1 score.

researchpythongo
0
3
Har Wearable Sensor EvalA

Evaluates deep learning architectures (CNNs, LSTMs, DNNs) for frame-by-frame human activity recognition using wearable sensor time-series data. It probes the models' capacity to capture temporal dependencies and generalize across diverse domains (kitchen gestures, lifestyle/exercise, and medical gait analysis) while handling severe class imbalance. Use when the user wants to benchmark on Opportunity, PAMAP2, Daphnet Gait, or asks about evaluating this task. Reports mean f1-score.

researchpythongo
0
3