Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,835
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,009–7,032 of 20,835 skills

Hest Benchmark EvalA

Evaluates the ability of histopathology foundation models to predict gene expression levels from H&E-stained whole-slide image patches. It probes the alignment between morphological features and transcriptomic profiles across diverse cancer types and organs. Use when the user wants to benchmark on HEST-Benchmark, or asks about evaluating this task. Reports Pearson correlation.

researchpythongo
0
3
Hescape EvalA

Evaluates cross-modal alignment between histology images and spatial gene expression profiles, and tests downstream capabilities in gene mutation classification and direct gene expression prediction from whole-slide images. Use when the user wants to benchmark on 5K, Multi, ImmOnc, Colon, Breast, Lung, or asks about evaluating this task. Reports Recall@5.

researchpythonexpress
0
3
Herobench EvalA

Evaluates long-horizon planning and structured reasoning in a grid-based RPG virtual environment. Agents must generate multi-step plans involving resource gathering, crafting, and combat, requiring integration of numerical calculations with action sequencing. Use when the user wants to benchmark on HeroBench, or asks about evaluating this task. Reports Success %.

researchpythongo
0
3
Heq EvalA

This benchmark evaluates extractive reading comprehension in Hebrew, a morphologically rich language. It probes a model's ability to accurately identify answer spans within a given context passage, handling challenges like affixation, spelling variations, and domain-specific vocabulary across news and encyclopedic text. Use when the user wants to benchmark on HeQ, or asks about evaluating this task. Reports TLNLS.

researchpythongo
0
3
Hephaestus Minicubes EvalA

Evaluates models on detecting volcanic ground deformation using multi-modal InSAR data. It probes the ability to classify deformation presence and segment deformation areas from spatiotemporal interferometric time-series, while handling atmospheric noise and class imbalance. Use when the user wants to benchmark on Hephaestus Minicubes, or asks about evaluating this task. Reports F1-score, IoU.

researchpythongo
0
3
Hepatobench EvalA

Evaluates pathology foundation models on fine-grained tissue classification of liver cancer patches and whole-slide tumor/non-tumor segmentation. It measures how well models can quantify tissue composition (e.g., fibrosis, necrosis, inflammation) within clinically defined regions to support reproducible digital pathology analysis. Use when the user wants to benchmark on HepatoBench, or asks about evaluating this task. Reports F1-score, Dice coefficient.

researchpythongo
0
3
Helpsteer2 EvalA

This dataset evaluates language model responses across five key dimensions: helpfulness, correctness, coherence, complexity, and verbosity. It probes an LLM's ability to follow instructions, maintain factual accuracy, produce logically consistent text, and adapt to varying levels of detail and difficulty. Use when the user wants to benchmark on HelpSteer2, or asks about evaluating this task. Reports helpfulness.

researchpython
0
3
Helo Apr EvalA

Evaluates a cross-lingual program repair framework on low-resource programming languages (Ruby and Rust) by measuring functional correctness against unit tests and syntactic validity via compilation or parsing rates. Use when the user wants to benchmark on xCodeEval (Compact Set for Ruby and Rust), Defects4Ruby, or asks about evaluating this task. Reports Pass@k.

researchpythonrust
0
3
Helms Exploration EvalA

Evaluates the exploration efficiency and adaptability of a robot planner in unknown environments. It measures how well the system covers free space while minimizing travel distance and adapting to natural language preferences. Use when the user wants to benchmark on Simulated dungeon environments, Indoor office environment (130m x 100m), or asks about evaluating this task. Reports Travel Distance.

researchpythonnode
0
3
Helmholtz Scattering EvalA

Evaluates the accuracy and computational efficiency of numerical PDE solvers (BEM vs. PINNs) for 2D acoustic wave scattering. Probes generalization capability beyond the training domain and physical fidelity in far-field regions. Use when the user wants to benchmark on 2D Helmholtz Scattering Benchmark, or asks about evaluating this task. Reports relative error.

researchpython
0
3
Helmet Long Context EvalA

Evaluates a model's ability to retain, process, and reason over extended contexts (8K to 128K tokens) across retrieval-augmented generation (RAG) and long-range question answering (LongQA) tasks. It probes robustness to noise, multi-hop reasoning, and memorization in long-context settings. Use when the user wants to benchmark on HELMET, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Helm Recommender EvalA

Evaluates LLM-powered recommender systems across five human-centered dimensions: intent alignment, explanation quality, interaction naturalness, trust & transparency, and fairness & diversity. It combines expert human ratings with automated accuracy and fairness metrics across multiple domains and interaction scenarios. Use when the user wants to benchmark on MovieLens-1M, Amazon Books, Yelp, or asks about evaluating this task. Reports NDCG@10.

researchpythonrust
0
3
Helm Lite EvalA

This evaluation probes the robustness of open benchmarks against test-set memorization and data leakage. It measures whether small language models can artificially inflate leaderboard scores by overfitting directly to public test sets, revealing flaws in current benchmarking practices. Use when the user wants to benchmark on HELM-lite, or asks about evaluating this task. Reports Exact Match.

researchpythongo
0
3
Helm EvalA

Evaluates long-horizon vision-language-action (VLA) manipulation capabilities, specifically testing a model's ability to maintain cross-phase context, predict action failures before execution, and recover from perturbations via rollback or replanning. Use when the user wants to benchmark on LIBERO-LONG, CALVIN ABC→D, LIBERO-Recovery, or asks about evaluating this task. Reports TSR.

researchpythongo
0
3
Hello Chat EvalA

Evaluates an end-to-end Large Audio Language Model's capabilities in audio understanding (ASR, QA, translation, reasoning, emotion/event recognition, instruction following) and text-to-speech synthesis (naturalness, intelligibility, speaker similarity). Use when the user wants to benchmark on AIShell, WeNet, LibriSpeech, AlpacaEval, LLaMA Questions, Web Questions, Synthetic multilingual (Claude-generated), MMAU-Mini, EmoBox, AudioSet, CochlScene, Seed-TTS-Eval (Chinese), or asks about evaluat...

researchpythongo
0
3
Helelena Ce EvalA

Evaluates deep learning architectures for pilot-based channel estimation in 5G-NR OFDM systems. It probes the model's ability to reconstruct full Channel State Information (CSI) from sparse pilot measurements across varying SNR levels, Doppler shifts, and 3GPP TDL propagation profiles. Use when the user wants to benchmark on 5G Deep Learning Data Synthesis (MATLAB), or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Held Out Test LossA

Evaluates language model generalization and overfitting by measuring cross-entropy loss on a held-out test set. It probes how well the model retains predictive performance when trained on repeated or constrained data subsets. Use when the user has predictions and gold and needs to compute held-out test loss.

researchpythongo
0
3
Hed Benchmark EvalA

This benchmark evaluates whether large language models and automated essay scoring systems can correctly distinguish between harmful essays (containing toxic or discriminatory content) and argumentative essays (which present controversial but non-harmful viewpoints). It also measures safety alignment by tracking refusal rates and the tendency to redirect harmful prompts into ethical, argumentative responses. Use when the user wants to benchmark on HED benchmark, or asks about evaluating this ...

researchpythongit
0
3
Hebrew Tts EvalA

Evaluates the quality of a diacritic-free Hebrew text-to-speech system by measuring content accuracy, speech naturalness, and speaker similarity against baseline models and different text tokenization strategies. Use when the user wants to benchmark on Hebrew TTS Test Set, or asks about evaluating this task. Reports Content Preservation.

researchpythonperformance
0
3
Hebrew G2p EvalA

This benchmark evaluates a model's ability to convert unvocalized Hebrew text into fully-specified IPA transcriptions, including accurate stress placement and shva realization. It also measures downstream text-to-speech quality and inference latency to assess real-time applicability. Use when the user wants to benchmark on ILSpeech, SASPEECH, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Heartbeat Memory Pollution EvalA

This evaluation probes how socially manipulated content encountered during an AI agent's background 'heartbeat' execution influences its downstream behavior. It measures both immediate same-session carry-over and cross-session long-term memory pollution under varying social credibility cues, agent personas, and realistic content dilution. Use when the user wants to benchmark on MissClaw (Custom Testbed), or asks about evaluating this task. Reports ASR.

researchpython
0
3
Hear EvalA

Evaluates the zero-shot generalization and transferability of pre-trained audio representations across 19 diverse downstream tasks spanning speech, environmental sounds, and music. The benchmark requires models to perform without fine-tuning, emphasizing robustness and cross-domain adaptability. Use when the user wants to benchmark on FSD50K, ESC-50, GTZAN, Vocal Imitations, LibriCount, CREMA-D, VoxLingua107, Speech Commands, DCASE 2016 Task 2, Gunshot Triangulation, Beijing Opera, Mridingham...

researchpythonperformance
0
3
Healthgpt Medical Vqa Generation EvalA

Evaluates a medical vision-language model's ability to perform visual question answering (comprehension) and medical image synthesis (generation) on heterogeneous datasets. It probes the model's capacity to unify multiple downstream tasks using parameter-efficient fine-tuning without task interference. Use when the user wants to benchmark on VL-Health, VQA-RAD, SLAKE, PathVQA, IXI, SynthRAD2023, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Health Orsc Bench EvalA

This benchmark evaluates large language models' tendency to over-refuse benign health-related queries and their ability to provide safe, helpful completions in medical contexts. It specifically probes the trade-off between safety alignment and utility by measuring refusal rates on carefully curated boundary prompts across varying difficulty levels. Use when the user wants to benchmark on Health-ORSC-Bench, or asks about evaluating this task. Reports Over-Refusal Rate (ORR).

researchpythongo
0
3