Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,574
skills in category
983
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 4,969–4,992 of 23,574 skills

Sccluebenc EvalA

Evaluates the clustering accuracy, label consistency, and biological interpretability of various single-cell RNA-seq analysis methods. It probes how well traditional, deep learning, graph-based, and foundation model algorithms recover known cell type annotations across diverse tissues and dataset sizes. Use when the user wants to benchmark on scCluBench (36 human & mouse scRNA-seq datasets), or asks about evaluating this task. Reports Normalized Mutual Information (NMI).

researchpythongo
0
3
Scb Mt En Th 2020 EvalA

This protocol evaluates neural machine translation quality between English and Thai. It measures translation accuracy on a newly curated 1M-parallel corpus (SCB_1M) and a filtered OPUS corpus, while also testing cross-domain generalization on the IWSLT 2015 Thai-English benchmark. Use when the user wants to benchmark on SCB_1M, MT_OPUS, IWSLT 2015 Thai-English, or asks about evaluating this task. Reports SacreBLEU.

researchpythontesting
0
3
Scatspotter EvalA

Evaluates object detection and instance segmentation capabilities on real-world images of dog feces. It specifically probes model robustness to camouflage, occlusion, varying lighting conditions, and small object detection in outdoor urban environments. Use when the user wants to benchmark on ScatSpotter, or asks about evaluating this task. Reports mAP.

researchpythongo
0
3
Scared EvalA

Evaluates stereo correspondence and depth estimation methods for endoscopic surgical scenes. It probes how accurately models can reconstruct quasi-dense depth maps from stereo image pairs captured with structured light on biological tissue. Use when the user wants to benchmark on SCARED, or asks about evaluating this task. Reports mean absolute error in mm.

researchpython
0
3
Scanreason EvalA

Evaluates a model's ability to perform 3D visual grounding by jointly reasoning about implicit human instructions and localizing target objects in 3D scenes. It probes spatial, functional, logical, emotional, and safety-related reasoning capabilities alongside precise 3D bounding box localization. Use when the user wants to benchmark on ScanReason, or asks about evaluating this task. Reports matching score.

researchpythongo
0
3
Scannet 3d Detection EvalA

Evaluates the ability of 3D object detectors to localize and classify indoor objects using variable-frame sparse RGB-D inputs. It probes generalization across different input modalities (reconstructed point clouds, multi-view RGB-D, monocular RGB-D) and varying numbers of input views. Use when the user wants to benchmark on ScanNet, or asks about evaluating this task. Reports mAP@0.25.

researchpythongo
0
3
Scannerf EvalA

Evaluates the rendering quality and novel-view synthesis capability of Neural Radiance Field (NeRF) methods on real-world inward-facing object scans. It probes how well models generalize to unseen camera poses when trained with varying image densities and localized acquisition patterns. Use when the user wants to benchmark on ScanNeRF, or asks about evaluating this task. Reports PSNR.

researchpythongo
0
3
Scanner Mner EvalA

Evaluates multi-modal named entity recognition (MNER) and visual grounding capabilities, specifically probing the model's ability to generalize to unseen entities by leveraging external knowledge (Wikipedia) and image-based features. Use when the user wants to benchmark on MNER, GMNER, or asks about evaluating this task. Reports F1 score.

researchpythonrust
0
3
Scandinavian Sentiment EvalA

Evaluates whether translating low-resource language data into English and applying large-scale English/multilingual models outperforms training native monolingual models for sentiment classification. It probes the efficiency and effectiveness of cross-lingual data reuse versus isolated language-specific pre-training. Use when the user wants to benchmark on Sentiment datasets (Swedish, Danish, Norwegian, Finnish, English), or asks about evaluating this task. Reports binary accuracy.

researchpythongo
0
3
Scandeval Benchmark EvalA

Evaluates the performance of monolingual and multilingual language models across five Scandinavian languages (Danish, Norwegian, Swedish, Icelandic, Faroese) on question answering, linguistic acceptability, and named entity recognition. It also probes cross-lingual transfer capabilities between these languages by measuring performance variance across language groups. Use when the user wants to benchmark on ScandiQA, ScaLA, MIM-GOLD-NER, WikiANN, or asks about evaluating this task. Reports acc...

researchpythongo
0
3
Scam Typographic Robustness EvalA

This benchmark evaluates the typographic robustness of vision-language models (VLMs) and large vision-language models (LVLMs) by measuring their susceptibility to adversarial handwritten or synthetic text inserted into images. It probes whether models can correctly identify the primary object in an image despite the presence of misleading attack words, revealing vulnerabilities in multimodal alignment and text-visual reasoning. Use when the user wants to benchmark on SCAM, or asks about evalu...

researchpythongo
0
3
Scalar EvalA

Evaluates long-context academic reasoning by testing whether LLMs can correctly identify masked citations within scientific papers. It probes the model's ability to understand semantic context, attributional claims, and descriptive references across varying context lengths and difficulty levels. Use when the user wants to benchmark on SCALAR, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Sc Heureka Bench EvalA

Evaluates AI co-scientist agents' ability to autonomously plan, execute, and interpret single-cell biology workflows to answer open-ended research questions (OEQs) and multiple-choice questions (MCQs). It probes hypothesis generation, code execution, data analysis, and scientific reasoning in a domain-specific setting. Use when the user wants to benchmark on sc-HeurekaBench-Lite, or asks about evaluating this task. Reports Correctness [1-5].

researchpythongo
0
3
Sbsc Math Olympiad EvalA

Evaluates LLMs' ability to solve complex, Olympiad-level mathematics problems across Algebra, Combinatorics, Number Theory, and Geometry. It specifically probes multi-turn code generation and execution feedback for iterative problem decomposition and constraint handling. Use when the user wants to benchmark on AIME, AMC-12, MathOdyssey, OlympiadBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Sbr Srgi EvalA

Evaluates session-based recommendation models by predicting the next item in a user's session sequence. It probes the model's ability to capture sequential item transitions and leverage global item-transition patterns across sessions to improve ranking accuracy. Use when the user wants to benchmark on Diginetica, Tmall, Nowplaying, or asks about evaluating this task. Reports P@20.

researchpythongo
0
3
Sbr Intent EvalA

Evaluates a session-based recommendation model's ability to predict the next item in a user session using validated and enriched LLM-generated intents. It probes the model's capacity to leverage semantic intent signals alongside sequential interaction patterns for accurate item ranking. Use when the user wants to benchmark on Beauty (Amazon), Yelp, Books (Amazon), or asks about evaluating this task. Reports Hit Rate@10.

researchpython
0
3
Saw Bench EvalA

Evaluates multimodal foundation models' ability to perform observer-centric spatial reasoning, path integration, and camera geometry inference using egocentric videos recorded from smart glasses. It probes sustained tracking of intermediate movements and the distinction between physical translation and camera rotation. Use when the user wants to benchmark on SAW-BENCH, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Save Video Text Retrieval EvalA

Evaluates a model's ability to retrieve relevant videos given a natural language query in an audio-visual setting. It specifically probes how well speech-aware representations and early vision-audio alignment improve cross-modal matching accuracy across diverse video-text benchmarks. Use when the user wants to benchmark on MSRVTT-9k, MSRVTT-7k, VATEX, Charades, LSMDC, or asks about evaluating this task. Reports SumR.

researchpythongo
0
3
Satellite Llp EvalA

Evaluates the ability of lightweight deep learning models to predict fine-grained class proportions (e.g., vegetation density, population) from satellite image chips. The protocol measures how well models trained on coarse administrative-level label proportions can recover fine-grained spatial distributions, using both proportion regression and pixel-level segmentation accuracy. Use when the user wants to benchmark on esaworldcover, humanpop, or asks about evaluating this task. Reports MAE.

researchpythongit
0
3
Sass EvalA

This benchmark probes the ability of toxicity detection models to identify nuanced, adversarially crafted harmful content (e.g., gaslighting, manipulation, sarcasm) that mainstream tools often miss due to reliance on normative annotations and profanity cues. Use when the user wants to benchmark on SASS, or asks about evaluating this task. Reports F1-Score.

researchpythongo
0
3
Sasrec Sequential Rec EvalA

Evaluates a model's ability to predict the next item in a user's interaction sequence based on historical behavior. It probes the model's capacity to capture long-range dependencies and adapt to varying data sparsity across different domains. Use when the user wants to benchmark on Amazon (Beauty), Amazon (Games), Steam, MovieLens-1M, or asks about evaluating this task. Reports Recall@K.

researchpythontesting
0
3
Sarena Icon EvalA

Evaluates a model's ability to generate scalable vector graphics (SVG) from text prompts and reference images, measuring visual fidelity, semantic alignment, structural success, and code efficiency. Use when the user wants to benchmark on SArena-Icon, or asks about evaluating this task. Reports SR.

researchpythongo
0
3
Sard Ocr EvalA

Evaluates the robustness and accuracy of OCR models on synthetic, book-style Arabic documents with high typographic diversity across 10 fonts. It measures character-level precision, word-level accuracy, and overall sequence fluency to benchmark vision-language and traditional OCR systems. Use when the user wants to benchmark on SARD, or asks about evaluating this task. Reports CER.

researchpythonperformance
0
3
Saplma Truthfulness EvalA

Evaluates whether an LLM's internal hidden layer activations can predict the veracity of a given statement. It probes the model's implicit knowledge of truthfulness by training a classifier on neural activations rather than relying on explicit prompting or output probabilities. Use when the user wants to benchmark on True-False Dataset, LLM-Generated Statements, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3