Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,504
skills in category
980
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 4,609–4,632 of 23,504 skills

Socialnav Sub EvalA

Evaluates Vision-Language Models' ability to perform spatial, spatiotemporal, and social reasoning in dynamic, crowd-filled robot navigation scenarios. It probes scene understanding by asking VLMs to answer visual questions based on image sequences and Bird's Eye View (BEV) representations. Use when the user wants to benchmark on SocialNav-SUB, or asks about evaluating this task. Reports PA.

researchpythongo
0
3
Socialcounterfactuals EvalA

Probes intersectional social bias in Large Vision-Language Models by measuring how model outputs vary when only perceived race, gender, or physical attributes change in counterfactual images. It specifically evaluates toxicity, stereotypical language, and competency ratings across different demographic groups. Use when the user wants to benchmark on SocialCounterfactuals, or asks about evaluating this task. Reports MaxToxicity.

researchpythongit
0
3
Social Media Bias EvalA

This benchmark evaluates the ability of models to automatically detect multiple dimensions of media bias (e.g., hate speech, racial, gender, political, linguistic, and text-level context bias) in social media posts across different topic domains. It probes a model's robustness to domain shift and severe class imbalance in multi-label bias identification tasks. Use when the user wants to benchmark on Social Media Bias Dataset (YouTube & Reddit), or asks about evaluating this task. Reports weig...

researchpythongo
0
3
Social Chem 101 EvalA

Evaluates a model's ability to reason about and generate rules-of-thumb (RoTs) that capture social and moral norms across 12 distinct dimensions of judgment, such as cultural pressure, legality, and moral foundations. It probes whether neural models can produce attribute-aware, context-sensitive normative judgments for unseen social scenarios. Use when the user wants to benchmark on Social-Chem-101, or asks about evaluating this task. Reports micro-F1.

researchpythongo
0
3
Soccsci210 EvalA

This evaluation probes an LLM's ability to predict individual human responses in social science experiments and match the overall distribution of those responses. It measures both point-wise accuracy and distributional alignment across unseen studies, conditions, outcomes, and participant demographics. Use when the user wants to benchmark on SocSci210, or asks about evaluating this task. Reports Accuracy, Wasserstein distance.

researchpythongo
0
3
Soccerchat EvalA

Evaluates multimodal video-language models on soccer-specific tasks: referee decision validation via question-answering and multi-label action classification. It probes the model's ability to align visual, auditory, and textual cues with ground-truth soccer events and rules. Use when the user wants to benchmark on XFoul validation dataset, SoccerNet-v2, or asks about evaluating this task. Reports QwQ Scorer, F1 Score (wt).

researchpythongo
0
3
Soc Dgl Dti EvalA

This evaluation protocol assesses a model's capability to predict binary drug-target interactions (DTI) using graph-based representations. It specifically probes performance under both balanced and highly imbalanced data distributions, as well as generalization to unseen drugs or targets in cold-start scenarios. Use when the user wants to benchmark on KIBA, Davis, BindingDB, DrugBank, or asks about evaluating this task. Reports AUROC.

researchpythonperformance
0
3
Soberdse EvalA

Evaluates a learning-based algorithm selection framework for High-Level Synthesis Design Space Exploration (DSE). It measures how accurately the model recommends the best-performing DSE algorithm for a given benchmark, and assesses the resulting optimization performance (ADRS) and runtime compared to heuristic and reinforcement learning baselines. Use when the user wants to benchmark on MachSuite & Polyhedral Benchmarks, or asks about evaluating this task. Reports recommendation_accuracy.

researchpythongo
0
3
Soar Rna EvalA

Evaluates large language models on zero-shot and chain-of-thought cell type annotation tasks using single-cell RNA-seq gene expression profiles. It probes the models' ability to translate structured genomic data into textual descriptions and accurately predict cell type labels without fine-tuning. Use when the user wants to benchmark on SOAR-RNA, or asks about evaluating this task. Reports Average BLEU.

researchpythonexpress
0
3
Soap Note Hallucination EvalA

Evaluates the hallucination rate of LLM-generated medical SOAP notes against physician-patient transcripts. It compares a literal, inference-unaware evaluation framework against a clinically informed, inference-aware framework to measure how often valid clinical reasoning is incorrectly flagged as hallucination. Use when the user wants to benchmark on Physician-Patient Transcripts, or asks about evaluating this task. Reports Mean Hallucination Rate.

researchpython
0
3
Soap Note Generation EvalA

Evaluates the ability of audio and text models to generate clinically accurate, well-structured SOAP notes from long-form doctor-patient conversations. Probes long-context audio reasoning, fact-grounding, and clinical documentation quality. Use when the user wants to benchmark on Doctor-Patient SOAP Conversations, or asks about evaluating this task. Reports Faithfulness.

researchpythondocumentation
0
3
Snr Detection ThresholdA

Evaluates the detection capability of the Lunar Gravitational-Wave Antenna (LGWA) for massive binary black hole mergers by computing the signal-to-noise ratio (SNR) of observed and simulated events against fixed thresholds. Use when the user has predictions and gold and needs to compute SNR.

researchpythongo
0
3
Snr Bench EvalA

Evaluates the robustness of audio deepfake detection models under varying signal-to-noise ratios (SNRs) by testing binary (real vs. spoof) and four-class (real+clean, real+noisy, spoof+clean, spoof+noisy) classification tasks on ASVspoof 2021 utterances augmented with MS-SNSD ambient noise. Use when the user wants to benchmark on ASVspoof 2021 (DF), or asks about evaluating this task. Reports EER.

researchpythontesting
0
3
Snntop1 Accuracy EvalA

Evaluates the classification accuracy of directly-trained spiking neural networks (SNNs) on both static image recognition and neuromorphic event-based vision tasks. It probes the model's ability to maintain gradient stability and high predictive performance while operating with minimal simulation timesteps, highlighting efficiency gains over traditional ANN-SNN conversion methods. Use when the user wants to benchmark on CIFAR-10, ImageNet, DVS-Gesture, DVS-CIFAR10, or asks about evaluating th...

researchpythongo
0
3
Snn Dfe Optical EvalA

Evaluates the communication performance and hardware efficiency of Spiking Neural Network-based Decision-Feedback Equalizers (DFEs) for optical channels compared to traditional Artificial Neural Network baselines. It probes the trade-off between bit error rate, computational complexity, and energy efficiency under varying quantization levels and FPGA resource constraints. Use when the user wants to benchmark on Custom Optical Communication DFE Benchmark, or asks about evaluating this task. Re...

researchpythongo
0
3
Snips Slu EvalA

Evaluates end-to-end spoken language understanding by measuring how well an embedded system extracts intents and slots from spoken audio. It probes the pipeline's ability to generalize to unseen queries and handle real-world ASR errors under strict resource constraints. Use when the user wants to benchmark on SmartLights, Weather, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Snip Pruning EvalA

Evaluates a single-shot pruning method's ability to identify and remove unimportant network connections at initialization, preserving classification accuracy across varying sparsity levels on standard vision and sequence datasets. Use when the user wants to benchmark on MNIST, CIFAR-10, Tiny-ImageNet, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Smtm Mobile Cnn EvalA

Evaluates the latency reduction, accuracy loss, memory overhead, energy saving, and early exit performance of a semantic memory caching mechanism (SMTM) for accelerating CNN inference on mobile devices. Use when the user wants to benchmark on UCF101, CIFAR-100 (long-tail), or asks about evaluating this task. Reports latency reduction.

researchpythongo
0
3
Smtlib Qf Bv EvalA

Evaluates the effectiveness of different MCSAT-based bitvector solving strategies and conflict explainers against a standard SMT benchmark suite. It measures how well each solver variant handles fixed-size bitvector formulas under a strict time limit. Use when the user wants to benchmark on SMT-LIB QF_BV, or asks about evaluating this task. Reports solved_instances.

researchpythongo
0
3
Smtce EvalA

Evaluates Vietnamese social media text classification across four tasks: constructive speech detection, complaint detection, emotion recognition, and hate speech detection. It probes the ability of monolingual versus multilingual BERT-based models to handle low-resource, domain-specific Vietnamese text with varying preprocessing requirements. Use when the user wants to benchmark on VSMEC, ViCTSD, ViOCD, ViHSD, or asks about evaluating this task. Reports macro-average F1 score.

researchpythonperformance
0
3
Smplolympics EvalA

Evaluates the ability of physically simulated humanoid agents to perform complex, long-horizon Olympic sports tasks using different control policies and motion priors. It probes task completion accuracy, physical realism, and the effectiveness of adversarial vs. hierarchical reinforcement learning in sparse-reward simulation environments. Use when the user wants to benchmark on SMPLOlympics Sports Environments, or asks about evaluating this task. Reports Suc Rate.

researchpythongo
0
3
Smolvla Robotics EvalA

Evaluates a vision-language-action model's ability to perform robotic manipulation tasks in both simulated and real-world environments. It probes visuomotor policy generalization, fine-grained task decomposition handling, and the impact of pretraining and inference modes on success rates. Use when the user wants to benchmark on LIBERO, Meta-World, SO100 Real-World Tasks, SO101 Real-World Tasks, or asks about evaluating this task. Reports Success Rate (SR).

researchpython
0
3
Smoldocling Doc EvalA

This evaluation probes a vision-language model's ability to perform end-to-end document conversion, including text recognition, layout analysis, table and chart structure extraction, and code/formula parsing. It measures how accurately the model reconstructs document content and spatial structure from page images into standardized markup formats. Use when the user wants to benchmark on DocLayNet, SynthCodeNet, Im2Latex-230k, FinTabNet, PubTables-1M, or asks about evaluating this task. Reports...

researchpythongo
0
3
Smol Chrf EvalA

Evaluates machine translation quality for 115 under-represented languages using professionally translated parallel data. It measures the improvement in character-level n-gram F-score (ChrF) after fine-tuning a baseline model on the Smol dataset compared to the unfine-tuned baseline. Use when the user wants to benchmark on SmolSent, SmolDoc, or asks about evaluating this task. Reports ChrF.

researchpythonperformance
0
3