Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,552
skills in category
982
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 4,825–4,848 of 23,552 skills

SemascoreA

Evaluates automatic speech recognition (ASR) transcription quality by measuring segment-wise semantic similarity and error weighting. It specifically probes robustness on disordered, noisy, and accented speech, testing alignment with human judgments and downstream NLU task metrics. Use when the user has predictions and gold and needs to compute SeMaScore.

researchpythongo
0
3
Semanticagent EvalA

Evaluates text-to-SQL generation capabilities across cross-domain parsing, knowledge-intensive reasoning, and enterprise-level SQL workflows. It also assesses the semantic validity, execution correctness, and diversity of synthetically generated training data. Use when the user wants to benchmark on Spider, BIRD, Spider2.0, EHRSQL, ScienceBenchmark, Spider-Syn, Spider-Realistic, Spider-DK, or asks about evaluating this task. Reports test-suite accuracy (TS), execution accuracy (EX).

researchpythongo
0
3
Semantic Textual Similarity EvalA

Evaluates a model's ability to quantify the degree of semantic similarity between pairs of sentences, including multilingual and cross-lingual contexts. It probes fine-grained semantic matching and cross-lingual generalization rather than binary paraphrase detection. Use when the user wants to benchmark on SemEval-2017 STS, or asks about evaluating this task. Reports Pearson correlation.

researchpythongo
0
3
Semantic Syntactic Word EvalA

Evaluates whether trained word vector models can capture semantic and syntactic relationships between words through simple algebraic operations in vector space. It probes the model's ability to solve analogy-style questions by measuring how well the vector arithmetic preserves linguistic regularities. Use when the user wants to benchmark on Semantic-Syntactic Word Relationship test set, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Semantic Kg EvalA

Evaluates the ability of semantic similarity methods to correctly classify pairs of natural language statements as semantically similar (label 1) or dissimilar (label 0). It specifically probes how well models handle controlled semantic variations (node and edge perturbations) across general and domain-specific knowledge domains. Use when the user wants to benchmark on Semantic-KG Benchmark, or asks about evaluating this task. Reports F1-score.

researchpythonnode
0
3
Semantic Change Detection EvalA

Evaluates the ability of contextualized language models to detect diachronic semantic change in words across different time periods. It probes whether models can accurately rank words by their degree of meaning shift compared to human-annotated gold standards. Use when the user wants to benchmark on SemEval-2020 Task 1, GEMS, or asks about evaluating this task. Reports Spearman's ρ.

researchpythongo
0
3
Selqa EvalA

Evaluates a model's ability to retrieve relevant answer sentences from a document given a question (selection), and to determine whether a document section contains an answer at all (triggering). It probes open-domain QA robustness against paraphrasing, varying question types, and section lengths. Use when the user wants to benchmark on SelQA, or asks about evaluating this task. Reports MAP.

researchpythongo
0
3
Seller Outcome Fairness EvalA

Evaluates the trade-off between platform revenue (GMV) and seller-side exposure fairness in online marketplace recommendation systems using simulated online environments trained on historical interaction data. Use when the user wants to benchmark on Proprietary Dataset, Electronics Event History (EVS) Dataset, or asks about evaluating this task. Reports GMV relative change.

researchpythongo
0
3
Selfcheckgpt EvalA

Evaluates a model's ability to detect hallucinated versus factual content in generated text using zero-resource consistency metrics across stochastic samples. It probes whether factual knowledge yields coherent, consistent outputs while hallucinated content exhibits divergence across multiple generations. Use when the user wants to benchmark on SelfCheckGPT dataset, or asks about evaluating this task. Reports AUC-PR.

researchpythongo
0
3
Self Adaptive Curriculum Nlu EvalA

Evaluates whether self-adaptive curriculum learning strategies, which use pre-trained model confidence to estimate example difficulty, improve fine-tuning performance over random and length-based sampling baselines across multiple NLU tasks. Use when the user wants to benchmark on SST-2, SST-5, HSOL, XNLI, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Seist Earthquake Monitoring EvalA

Evaluates a deep learning model's capability to perform multiple earthquake monitoring tasks, including seismic phase picking, detection, polarity classification, and magnitude estimation. It specifically probes cross-regional out-of-distribution generalization by training on Chinese seismic network data and testing on geologically distinct Pacific Northwest data. Use when the user wants to benchmark on DiTing, PNW (ComCat event subset), or asks about evaluating this task. Reports F1-Score.

researchpythongo
0
3
Seismic Segmentation Al EvalA

This evaluation probes a model's ability to perform semantic segmentation on 3D seismic data under an active learning regime. It measures how effectively a model generalizes to unseen geological volumes when trained on a sequentially selected subset of annotated sections, rather than a fixed passive dataset. Use when the user wants to benchmark on F3 benchmark, Parihaka, or asks about evaluating this task. Reports mIoU.

researchpythongo
0
3
Seismic Response EvalA

Evaluates a neural operator's ability to map high-frequency seismic wave excitations to building displacement responses across multiple floors. It specifically probes the model's capacity to capture oscillatory function spaces and handle amplitude-frequency disparities between different structural floors. Use when the user wants to benchmark on Custom seismic building response dataset, or asks about evaluating this task. Reports mean relative L2 error.

researchpythontesting
0
3
Seismic Picker EvalA

Evaluates the performance of deep learning and classical seismic phase pickers across three tasks: event detection, phase identification, and onset time picking. It probes cross-domain transfer capabilities and robustness to varying signal-to-noise ratios and waveform characteristics. Use when the user wants to benchmark on LenDB, GEOFON, INSTANCE, SCEDC, STEAD, ETHZ, Iquique, NEIC, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Seismic Phase Association EvalA

Evaluates the accuracy and computational efficiency of seismic phase associators on synthetic crustal and subduction zone datasets under varying event densities and noise levels. It probes the models' ability to correctly group seismic picks into events and maintain performance under high-stress conditions. Use when the user wants to benchmark on Synthetic Seismic Scenarios (Crustal & Subduction), or asks about evaluating this task. Reports event-level F1 score.

researchpythonrust
0
3
Seismic Inversion EvalA

Evaluates a model's ability to perform semi-supervised seismic impedance inversion using ultra-sparse well-log labels. It probes voxel-level accuracy, patch-level structural similarity, and percentage error on both synthetic and real-world 3D seismic volumes. Use when the user wants to benchmark on SEAM Phase I, Netherlands F3, Delft, or asks about evaluating this task. Reports MAE.

researchpythonexpress
0
3
Seismic Event Classification EvalA

Evaluates a model's ability to discriminate between three types of seismic events (earthquakes, quarry blasts, and background noise) using waveform and spectral features. It probes the model's capacity to learn physically meaningful seismological signatures like P/S-wave onsets and spectral decay patterns. Use when the user wants to benchmark on Curated Seismic Dataset, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Segmentmeifyoucan EvalA

This benchmark evaluates a model's ability to segment anomalous or hazardous objects in driving scenes that were not seen during training. It probes out-of-distribution detection and pixel-wise localization of unknown road obstacles, emphasizing safety-critical detection regardless of object class. Use when the user wants to benchmark on RoadAnomaly21, RoadObstacle21, or asks about evaluating this task. Reports AuPRC.

researchpythongo
0
3
Segbook EvalA

Evaluates the transfer learning and fine-tuning capabilities of volumetric medical image segmentation models across diverse imaging modalities, anatomical targets, and dataset sizes. It probes how well pre-trained models generalize to downstream segmentation tasks in clinical imaging scenarios and reveals non-linear performance scaling with dataset scale. Use when the user wants to benchmark on SegBook, or asks about evaluating this task. Reports Dice Score (DSC).

researchpythongo
0
3
Segale Doc EvalA

Evaluates whether a document-level machine translation evaluation framework can robustly handle translation anomalies (over-translation, under-translation, boundary shifts) and effectively score long-form texts without predefined sentence boundaries. Use when the user wants to benchmark on SEGALE Test Set, or asks about evaluating this task. Reports correlation with human judgments.

researchpythongo
0
3
Sega Layout EvalA

Evaluates a model's ability to generate content-aware graphic layouts from background images and instructions. It probes spatial reasoning, adherence to design principles (alignment, overlap, occlusion), and aesthetic quality. Use when the user wants to benchmark on PKU, CGL, Crello, or asks about evaluating this task. Reports Ali.

researchpython
0
3
Seephys EvalA

This benchmark evaluates multimodal LLMs' ability to perform physics reasoning using visual diagrams, text, or both. It probes visual interpretation, diagram-to-reasoning mapping, and the model's reliance on textual shortcuts versus actual visual perception across varying knowledge levels and diagram types. Use when the user wants to benchmark on SeePhys, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Seegull EvalA

Probes a model's propensity to generate or recognize stereotypical associations across diverse global and state-level identity groups. Evaluates the prevalence and cultural specificity of biases in English NLP models, highlighting regional disparities in stereotype content and offensiveness. Use when the user wants to benchmark on SeeGULL, or asks about evaluating this task. Reports stereotype_prevalence.

researchpythongo
0
3
Seeds Superpixel EvalA

Evaluates the quality of superpixel segmentation algorithms by measuring how well superpixel boundaries align with ground-truth object boundaries and how accurately superpixels can be used as indivisible units for downstream segmentation tasks. Use when the user wants to benchmark on Berkeley Segmentation Dataset (BSD), or asks about evaluating this task. Reports under-segmentation error (UE).

researchpythongo
0
3