Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,504
skills in category
980
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 4,681–4,704 of 23,504 skills

Simbarca Traffic Forecasting EvalA

Evaluates the ability of deep learning models to forecast urban traffic speeds at both the individual road segment and regional levels. It probes spatio-temporal forecasting capabilities under varying congestion levels, testing how well models integrate multi-source sensor data (drone trajectories and loop detectors) to predict future traffic states. Use when the user wants to benchmark on SimBarca, or asks about evaluating this task. Reports MAE.

researchpythonnode
0
3
Simba Benchmark Analysis EvalA

Evaluates a framework for analyzing language model performance matrices by identifying dataset-model correlations, discovering minimal representative dataset subsets, and predicting held-out model performance while preserving model rankings. Use when the user wants to benchmark on HELM, MMLU, BigBenchLite, or asks about evaluating this task. Reports coverage ($\eta$).

researchpythonperformance
0
3
Sim3d EvalA

Evaluates 3D anomaly detection and segmentation capabilities in industrial settings using multiview and multimodal (image + depth) inputs. It probes a model's ability to identify and localize defects across multiple object categories under both in-domain (real-to-real) and out-of-domain (synthetic-to-real) conditions. Use when the user wants to benchmark on SiM3D, or asks about evaluating this task. Reports I-AUROC.

researchpythongo
0
3
Sim1 Tshirt Fold EvalA

Evaluates a robot policy's ability to perform structured deformable manipulation (t-shirt folding) in real-world settings after being trained exclusively on simulation data. It probes sim-to-real transfer, out-of-domain robustness to environmental shifts, and data scaling efficiency. Use when the user wants to benchmark on SIM1 T-shirt Folding, or asks about evaluating this task. Reports success.

researchpythongo
0
3
Silicone Prep Anomaly EvalA

Evaluates multimodal vision-language models' ability to detect context-dependent visual anomalies in robotic scientific laboratory workflows using first-person imagery and stage-specific textual prompts. Use when the user wants to benchmark on Silicone Preparation Workflow, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Silicone EvalA

Evaluates a model's ability to perform sequence labelling on spoken dialogues, specifically predicting dialog acts (DA) and emotion/sentiment (E/S) labels per utterance within multi-utterance conversations. Use when the user wants to benchmark on SILICONE, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Significance EvalA

Evaluates the ability of ML classifiers to distinguish signal from background events in high-energy physics simulations. It measures classification quality via AUC and quantifies discovery potential using a likelihood-ratio-based significance metric optimized over a probability threshold. Use when the user has predictions and gold and needs to compute significance.

researchpythongo
0
3
Signalmc Med EvalA

This benchmark evaluates biosignal foundation models on synchronized, long-duration single-lead ECG and PPG recordings from emergency department visits. It probes the models' ability to extract clinically meaningful representations for tasks such as age and sex prediction, emergency disposition, laboratory value regression, and ICD-10 diagnosis classification. Use when the user wants to benchmark on SignalMC-MED, or asks about evaluating this task. Reports AUROC.

researchpythongo
0
3
Signal Quality Auditing EvalA

Evaluates the effectiveness of various Signal Quality Indices (SQIs) and machine learning models for classifying noisy vs. clean ECG signals, detecting outliers, and denoising time-series data. Use when the user wants to benchmark on Physionet 2011 ECG Challenge (PICC) Set A, MIT-BIH Arrhythmia & NSTDB, or asks about evaluating this task. Reports AUC, Mean Squared Error (MSE).

researchpythontesting
0
3
Sign Recommender EvalA

This protocol evaluates a model's ability to detect beneficial feature interactions for recommendation and graph classification tasks. It measures prediction accuracy and ranking performance while assessing how well the model filters out irrelevant feature pairs to improve generalization. Use when the user wants to benchmark on Frappe, MovieLens-tag, Twitter, DBLP, or asks about evaluating this task. Reports accuracy (ACC).

researchpythongo
0
3
Sigmacollab EvalA

Probes an AI system's ability to understand and assist in physically situated, goal-directed collaboration tasks using multimodal egocentric sensing. It evaluates real-time scene understanding, interaction modeling, and proactive guidance in fluid, human-AI collaborative scenarios. Use when the user wants to benchmark on SigmaCollab, or asks about evaluating this task. Reports classification.

researchpythongo
0
3
Sifthinker EvalA

Evaluates a model's spatial reasoning, fine-grained visual perception, and 3D spatial grounding capabilities. It also measures self-correction ability through dynamic bounding box refinement and general visual-language understanding across multiple benchmarks. Use when the user wants to benchmark on SpatialBench, SAT-Static, CV-Bench, VisCoT_s, V*Bench, RefCOCO, RefCOCO+, RefCOCOg, OVDEval, MME, MMBench, SEED-Bench, VQAv2, POPE, or asks about evaluating this task. Reports Top-1 Accuracy@0.5.

researchpythongo
0
3
Sidon Speech Restoration EvalA

Evaluates multilingual speech restoration quality by measuring acoustic fidelity, speaker preservation, and transcription accuracy on noisy speech. It also assesses downstream utility by training TTS models on cleansed data and measuring synthetic speech quality, alongside inference speed benchmarks. Use when the user wants to benchmark on test-clean/test-other subsets (English), Multilingual test set, TED-LIUM Release 3, or asks about evaluating this task. Reports DNSMOS.

researchpythongo
0
3
Sid Generalization EvalA

Evaluates a model's ability to generalize synthetic image detection across diverse generative architectures (GANs, Diffusion Models, DiTs) and real-world image sources, focusing on robustness to unseen generators and varying image resolutions. Use when the user wants to benchmark on SID Generalization Benchmark (ForenSynths, Self-Synthesis, Ojha, GenImage, DiTFake), or asks about evaluating this task. Reports ACC.

researchpythontesting
0
3
Sib200 Xlt EvalA

This evaluation probes a model's ability to perform zero-shot and fully-supervised cross-lingual text classification across typologically diverse languages. It specifically measures how well parameter-efficient soft prompt tuning methods transfer knowledge from high-resource source languages to low-performing or unseen target languages without language-specific fine-tuning. Use when the user wants to benchmark on SIB-200, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Shrutilipi Asr EvalA

Evaluates the quality, diversity, and downstream effectiveness of the Shrutilipi audio-text dataset for low-resource Indian language ASR. It measures how adding mined data improves Word Error Rate (WER) on standard and noisy benchmarks compared to existing datasets. Use when the user wants to benchmark on Shrutilipi, MUCS, Kathbath, CommonVoice, or asks about evaluating this task. Reports WER.

researchpythonperformance
0
3
Shredbench EvalA

Evaluates multimodal LLMs' ability to reconstruct shredded documents from fragmented visual inputs. It probes cross-modal semantic reasoning, visual discontinuity alignment, and fine-grained positional continuity awareness across natural language, source code, and tabular data. Use when the user wants to benchmark on ShredBench, or asks about evaluating this task. Reports NED.

researchpythonjava
0
3
Shpi Recommendation EvalA

Evaluates offline reinforcement learning methods for session-based recommendation systems in optimizing long-term user retention versus short-term clicks. It tests the ability of algorithms to learn from fixed logging policies and generalize to online rollouts across synthetic, simulated, and real-world recommendation environments. Use when the user wants to benchmark on Synthetic recommendation problem, RecoGym, HIV treatment simulator, Private dataset X, or asks about evaluating this task. ...

researchpythongo
0
3
Showui Gui EvalA

Evaluates a vision-language-action model's ability to perform GUI visual grounding and task-oriented navigation across web, mobile, and online environments. It probes the model's capacity to interpret screenshots, locate interactive elements, and execute correct action sequences to complete user instructions. Use when the user wants to benchmark on Screenspot, Mind2Web, AITW, MiniWob, or asks about evaluating this task. Reports Zero-shot grounding accuracy.

researchpythonperformance
0
3
Showdown Clicks EvalA

Evaluates a vision-language model's ability to accurately predict the correct UI element to click based on a screenshot and a task instruction. It isolates low-level visual grounding and interaction skills in ambiguous or icon-heavy interfaces. Use when the user wants to benchmark on Showdown-Clicks, or asks about evaluating this task. Reports Top-1 Accuracy.

researchpythongo
0
3
Shotbench EvalA

Evaluates vision-language models on expert-level cinematic understanding by testing their ability to reason about shot-level visual properties. It probes fine-grained visual reasoning and spatial cognition across eight cinematography dimensions such as shot size, framing, camera angle, lens, lighting, composition, and movement. Use when the user wants to benchmark on ShotBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Shot2story EvalA

Evaluates multi-modal video understanding across three tasks: single-shot captioning, multi-shot summarization, and question answering. It probes a model's ability to process visual frames, optional ASR text, and shot-level structure to generate coherent descriptions or answer temporal, holistic, and audio-related questions. Use when the user wants to benchmark on Shot2Story, or asks about evaluating this task. Reports CIDEr.

researchpythongo
0
3
Shorter Splatting EvalA

Evaluates the training efficiency and reconstruction fidelity of a 3D Gaussian Splatting method that uses scale reset and entropy-constrained alpha blending to reduce Gaussian list lengths. Use when the user wants to benchmark on Mip-NeRF 360, Deep Blending, Tanks and Temples, or asks about evaluating this task. Reports PSNR.

researchpythonperformance
0
3
Shopping Queries EvalA

Evaluates e-commerce search models on query-product semantic matching, ranking, and multiclass relevance classification. It probes the system's ability to distinguish between Exact matches, Substitutes, Complements, and Irrelevant products across English, Spanish, and Japanese. Use when the user wants to benchmark on Shopping Queries Dataset, or asks about evaluating this task. Reports nDCG.

researchpythongo
0
3