Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,565
skills in category
982
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 4,873–4,896 of 23,565 skills

Seamlessm4t Human EvalA

Probes the semantic preservation and audio naturalness of speech-to-text and speech-to-speech translation systems across 24+ languages. Uses human annotators to score translations on a 1-5 scale for meaning similarity (XSTS) and speech quality/naturalness (MOS). Use when the user wants to benchmark on FLEURS test partition, or asks about evaluating this task. Reports XSTS.

researchpython
0
3
Seam EvalA

Evaluates vision-language models on their ability to reason across semantically equivalent inputs presented in different modalities (vision vs. language) across domain-specific notation systems. It probes cross-modal consistency, modality-agnostic reasoning, and identifies domain-specific perception and tokenization failure modes. Use when the user wants to benchmark on SEAM, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Sealqa EvalA

Evaluates a model's ability to reason over noisy, conflicting, and ambiguous real-world search results. It probes complex skills like contradiction resolution, temporal tracking, false-premise detection, and multi-document needle-in-a-haystack retrieval. Use when the user wants to benchmark on SealQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Seaexam EvalA

Evaluates LLMs' ability to answer local, culturally grounded multiple-choice questions in Southeast Asian languages (Indonesian, Thai, Vietnamese). It probes regional knowledge, language comprehension, and alignment with actual local usage compared to translated benchmarks. Use when the user wants to benchmark on SeaExam, or asks about evaluating this task. Reports accuracy (%).

researchpythongo
0
3
Seacrowd Benchmark EvalA

Evaluates the zero-shot capability of LLMs, VLMs, and speech models across 13 NLU/NLG tasks, ASR, and image captioning for Southeast Asian languages. It probes multilingual understanding, generation, and cross-modal alignment in low-resource and indigenous language settings. Use when the user wants to benchmark on SEACrowd NLU, SEACrowd NLG, SEACrowd ASR, SEACrowd VL, or asks about evaluating this task. Reports weighted F1 score, WER.

researchpythongit
0
3
Seabench EvalA

Evaluates LLMs' ability to handle open-ended, daily interaction scenarios in Southeast Asian languages. It probes contextual adaptation, instruction following, and safety in real-world multilingual usage. Use when the user wants to benchmark on SeaBench, or asks about evaluating this task. Reports LLM-as-a-Judge Score.

researchpythongo
0
3
Sea Vision EvalA

Evaluates multimodal language models on document parsing and text-centric visual question answering across 11 Southeast Asian languages. Probes the models' ability to extract structured information from complex documents and answer questions based on visual-textual alignment in low-resource scripts. Use when the user wants to benchmark on SEA-Vision, or asks about evaluating this task. Reports answer accuracy.

researchpythongo
0
3
Sea Spoof EvalA

This benchmark evaluates audio deepfake detection models on their ability to distinguish real speech from synthetic speech across six South-East Asian languages. It specifically probes cross-lingual generalization and robustness against diverse open-source and commercial text-to-speech and voice conversion systems. Use when the user wants to benchmark on SEA-Spoof, or asks about evaluating this task. Reports EER (%).

researchpythonexpress
0
3
Sea Helm EvalA

Evaluates Thai language models across eight competencies including instruction following, multi-turn dialogue stability, natural language understanding, generation, reasoning, safety, and code-switching resistance. Use when the user wants to benchmark on SEA-HELM, or asks about evaluating this task. Reports SEA-HELM Average Score.

researchpythongo
0
3
Se Toxicity EvalA

Evaluates the ability of contemporary toxicity detection models to correctly identify toxic language in software engineering contexts, such as code reviews and developer chat logs. It probes whether general-purpose classifiers can handle domain-specific terminology and contextual nuances without significant performance degradation. Use when the user wants to benchmark on Jigsaw Sample, Code Review, Gitter Ethereum, or asks about evaluating this task. Reports F-Score.

researchpythongo
0
3
Se Asr Wer EvalA

Evaluates how speech enhancement (SE) artifacts and noise errors affect automatic speech recognition (ASR) performance. It measures Word Error Rate (WER) on enhanced speech signals derived from simulated and real-world reverberant noisy conditions to isolate the impact of artifact components. Use when the user wants to benchmark on Simulated WSJ0+CHiME-3, CHiME-3 et05_real, or asks about evaluating this task. Reports WER [%].

researchpythonperformance
0
3
Sdsko Pub Vdr EvalA

Evaluates a model's ability to retrieve relevant Korean public document pages given a text query, comparing text-only parsing against multimodal visual understanding. It probes cross-modal reasoning, layout awareness, and the capacity to interpret tables, charts, and complex visual structures in administrative documents. Use when the user wants to benchmark on SDS KoPub VDR, or asks about evaluating this task. Reports Recall@k.

researchpythongo
0
3
Sdmuse Music Editing EvalA

Evaluates the quality and controllability of a stochastic differential music generation and editing model. It probes the model's ability to generate pop piano music from scratch or conditioned on control signals, and perform fine-grained editing tasks like stroke-based generation, inpainting, and style transfer. Use when the user wants to benchmark on ailabs1k7, or asks about evaluating this task. Reports pitch distribution similarity (PD).

researchpythongo
0
3
Sdg Gan Oversampling EvalA

Evaluates whether a GAN-based oversampling technique (SDG-GAN) improves binary classification performance on imbalanced tabular data compared to traditional and GAN-based baselines. Use when the user wants to benchmark on Credit Card Fraud Dataset, Pima Diabetes Dataset, Breast Cancer Wisconsin (Diagnostic) Dataset, Gambling Fraud Dataset, or asks about evaluating this task. Reports algorithmic performance.

researchpythongo
0
3
Sddfcs Simulation EvalA

Evaluates a reinforcement learning policy for same-day delivery routing on synthetic geographic settings, measuring the trade-off between overall service utility and regional fairness (minimum service rate). Use when the user wants to benchmark on SDDFCS Simulation, or asks about evaluating this task. Reports r_total.

researchpythonperformance
0
3
Sd Mae Histopath EvalA

Evaluates the ability of self-distillation augmented masked autoencoders to learn robust visual representations from histopathological images for downstream tasks like classification, segmentation, and detection, particularly in low-class or cross-domain settings. Use when the user wants to benchmark on PatchCamelyon (PCam), NCT-CRC-HE (NCT), MSIIvsMSS, MoNuSeg, Glas, NuCLS, or asks about evaluating this task. Reports top-1 accuracy.

researchpythongo
0
3
Scssl Bench EvalA

Evaluates how well self-supervised learning models learn single-cell representations for three downstream tasks: batch correction, cell type annotation, and missing modality prediction. It probes the trade-off between preserving biological variance and removing technical batch effects, as well as the ability to generalize across uni- and multi-omics data modalities. Use when the user wants to benchmark on PBMC-M, BMMC, PBMC, Pancreas, Immune Cell Atlas, MCA, Lung, Tabula Sapiens, HIC, or asks...

researchpythonexpress
0
3
Scs Frequency EvalA

Evaluates machine learning models' ability to predict severe convective storm (SCS) frequency and occurrence in European Russia under climate change scenarios. It probes binary classification of SCS events against non-events, as well as regression accuracy on the normalized annual cycle of storm activity using physics-informed deep learning architectures. Use when the user wants to benchmark on CMIP5 RCP8.5 & Meteorological Observations, or asks about evaluating this task. Reports RMSEAC.

researchpythongo
0
3
Scrolls EvalA

Evaluates long-text understanding capabilities across summarization, question answering, and natural language inference tasks. It probes whether models can effectively process and extract information from documents exceeding standard context windows (up to 16K tokens) using chunked encoding and cross-chunk fusion. Use when the user wants to benchmark on SCROLLS, or asks about evaluating this task. Reports Avg SCROLLS score.

researchpythongo
0
3
Script Identification EvalA

Evaluates the accuracy of a script identification tool on multilingual web corpora by checking if the predicted writing system matches the admissible scripts for the corpus's assigned language. It also measures script representation coverage in multilingual LLM tokenizers. Use when the user wants to benchmark on Multilingual C4 (mC4), OSCAR 22.01, or asks about evaluating this task. Reports ACC.

researchpython
0
3
Screenspot Pro EvalA

This benchmark evaluates a model's ability to perform GUI grounding in professional, high-resolution desktop environments. It probes whether vision-language models can accurately locate specific UI elements (both text and icons) based on natural language instructions, highlighting challenges with small targets and complex interfaces. Use when the user wants to benchmark on ScreenSpot-Pro, or asks about evaluating this task. Reports accuracy (center-point).

researchpythongo
0
3
Screenspot EvalA

Evaluates a model's ability to locate specific UI elements on screenshots based on natural language instructions. It measures how accurately the model can predict click coordinates that align with ground-truth bounding boxes across mobile, desktop, and web platforms. Use when the user wants to benchmark on ScreenSpot, or asks about evaluating this task. Reports click accuracy.

researchpythongit
0
3
Screenpr EvalA

Evaluates a model's ability to read and describe the content and layout of a GUI screenshot at a specific pointed location. It probes layout-aware screen reading, spatial reasoning, and the capacity to generate focused descriptions for mobile agent navigation. Use when the user wants to benchmark on ScreenPR, or asks about evaluating this task. Reports Content Acc.

researchpythongo
0
3
Screendrag EvalA

Evaluates a model's ability to perform fine-grained text dragging interactions on GUI screenshots. It measures whether the model correctly triggers a drag action, accurately selects the target text span, and aligns its predicted coordinates with ground truth. Use when the user wants to benchmark on SCREENDRAG, or asks about evaluating this task. Reports DTR.

researchpythongo
0
3