Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,554
skills in category
982
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 4,849–4,872 of 23,554 skills

Seegull EvalA

Probes a model's propensity to generate or recognize stereotypical associations across diverse global and state-level identity groups. Evaluates the prevalence and cultural specificity of biases in English NLP models, highlighting regional disparities in stereotype content and offensiveness. Use when the user wants to benchmark on SeeGULL, or asks about evaluating this task. Reports stereotype_prevalence.

researchpythongo
0
3
Seeds Superpixel EvalA

Evaluates the quality of superpixel segmentation algorithms by measuring how well superpixel boundaries align with ground-truth object boundaries and how accurately superpixels can be used as indivisible units for downstream segmentation tasks. Use when the user wants to benchmark on Berkeley Segmentation Dataset (BSD), or asks about evaluating this task. Reports under-segmentation error (UE).

researchpythongo
0
3
Seed X EvalA

Evaluates a multimodal model's ability to understand images and text (comprehension) and generate images from text instructions (generation). It probes fine-grained visual perception, reasoning, and compositional image synthesis. Use when the user wants to benchmark on VQAv2, GQA, POPE, MME, SEED, MMB, MM-Vet, MMMU, GenEval, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Seed Tts EvalA

Evaluates zero-shot voice conversion systems on linguistic preservation, speaker identity retention, and audio naturalness across English, Chinese, and cross-lingual settings. It also measures computational efficiency and latency for both streaming and offline inference modes. Use when the user wants to benchmark on Seed-TTS-Eval, or asks about evaluating this task. Reports WER (%).

researchpythongo
0
3
Seed Emotion Recognition EvalA

Evaluates a model's ability to classify EEG signals into three affective states (negative, neutral, positive) using a semi-supervised learning framework. It tests representation learning and classification performance on high-dimensional, noisy time-series data with limited labeled sessions. Use when the user wants to benchmark on SEED, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Sede EvalA

Evaluates text-to-SQL models on naturally occurring, under-specified user queries from Stack Exchange. Probes the model's ability to handle real-world ambiguity, nested subqueries, parameterized queries, and domain-specific schema knowledge without relying on perfectly specified instructions. Use when the user wants to benchmark on SEDE, or asks about evaluating this task. Reports PCM-F1.

researchpythongo
0
3
Securerouter EvalA

This evaluation probes the accuracy and inference efficiency of an encrypted routing framework for secure Transformer inference. It measures how well a cost-aware router dynamically selects smaller MPC-optimized models from a pool to balance privacy-preserving computation costs with task-specific accuracy requirements. Use when the user wants to benchmark on GLUE, or asks about evaluating this task. Reports Inference Speed-up.

researchpythongo
0
3
Secure Inference LatencyA

Measures the online and offline computation latency and communication bandwidth for cryptographic primitives and neural network operations under secure two-party computation. It evaluates how efficiently packed homomorphic encryption and garbled circuits handle matrix-vector products, convolutions, and activation functions without revealing inputs or model parameters. Use when the user has predictions and gold and needs to compute t_online.

researchpythongo
0
3
Secretbench EvalA

Evaluates the capability of automated secret detection tools to accurately identify hardcoded secrets (e.g., API keys, passwords, private keys) in source code repositories. It probes the tools' ability to balance high recall for true secrets against low false positive rates to mitigate alert fatigue. Use when the user wants to benchmark on SecretBench, or asks about evaluating this task. Reports Precision.

researchpythonexpress
0
3
Secbench EvalA

Evaluates large language models' cybersecurity knowledge retention and logical reasoning capabilities across multiple subdomains, languages, and difficulty levels using multiple-choice and short-answer questions. Use when the user wants to benchmark on SecBench, or asks about evaluating this task. Reports correctness percentage.

researchpythongo
0
3
Sec Gfd EvalA

Evaluates graph neural networks for fraud detection on real-world transaction and review graphs. It specifically probes a model's robustness to severe class imbalance and heterophily, where connected nodes often belong to different classes. Use when the user wants to benchmark on Amazon, YelpChi, T-Finance, T-Social, or asks about evaluating this task. Reports F1-macro, AUC.

researchpythonnode
0
3
Seas Safety EvalA

This evaluation probes the safety alignment and refusal capabilities of LLMs when exposed to harmful or adversarial prompts. It measures the frequency of unsafe model outputs to quantify vulnerability, while simultaneously tracking general instruction-following scores to ensure that safety hardening does not degrade overall utility. Use when the user wants to benchmark on SEAS-Test, BeaverTrail, HH-RLHF, XSTest, or asks about evaluating this task. Reports Attack Success Rate (ASR).

researchpythongo
0
3
Search3d Lerf EvalA

Evaluates a model's ability to perform open-vocabulary segmentation and localization in 3D radiance fields, specifically testing its capacity to understand and process hierarchical queries (e.g., object parts relative to whole objects) versus simple object-level queries. Use when the user wants to benchmark on Search3D (adapted), LERF dataset, or asks about evaluating this task. Reports mIoU.

researchpythontesting
0
3
Seamlessm4t Human EvalA

Probes the semantic preservation and audio naturalness of speech-to-text and speech-to-speech translation systems across 24+ languages. Uses human annotators to score translations on a 1-5 scale for meaning similarity (XSTS) and speech quality/naturalness (MOS). Use when the user wants to benchmark on FLEURS test partition, or asks about evaluating this task. Reports XSTS.

researchpython
0
3
Seam EvalA

Evaluates vision-language models on their ability to reason across semantically equivalent inputs presented in different modalities (vision vs. language) across domain-specific notation systems. It probes cross-modal consistency, modality-agnostic reasoning, and identifies domain-specific perception and tokenization failure modes. Use when the user wants to benchmark on SEAM, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Sealqa EvalA

Evaluates a model's ability to reason over noisy, conflicting, and ambiguous real-world search results. It probes complex skills like contradiction resolution, temporal tracking, false-premise detection, and multi-document needle-in-a-haystack retrieval. Use when the user wants to benchmark on SealQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Seaexam EvalA

Evaluates LLMs' ability to answer local, culturally grounded multiple-choice questions in Southeast Asian languages (Indonesian, Thai, Vietnamese). It probes regional knowledge, language comprehension, and alignment with actual local usage compared to translated benchmarks. Use when the user wants to benchmark on SeaExam, or asks about evaluating this task. Reports accuracy (%).

researchpythongo
0
3
Seacrowd Benchmark EvalA

Evaluates the zero-shot capability of LLMs, VLMs, and speech models across 13 NLU/NLG tasks, ASR, and image captioning for Southeast Asian languages. It probes multilingual understanding, generation, and cross-modal alignment in low-resource and indigenous language settings. Use when the user wants to benchmark on SEACrowd NLU, SEACrowd NLG, SEACrowd ASR, SEACrowd VL, or asks about evaluating this task. Reports weighted F1 score, WER.

researchpythongit
0
3
Seabench EvalA

Evaluates LLMs' ability to handle open-ended, daily interaction scenarios in Southeast Asian languages. It probes contextual adaptation, instruction following, and safety in real-world multilingual usage. Use when the user wants to benchmark on SeaBench, or asks about evaluating this task. Reports LLM-as-a-Judge Score.

researchpythongo
0
3
Sea Vision EvalA

Evaluates multimodal language models on document parsing and text-centric visual question answering across 11 Southeast Asian languages. Probes the models' ability to extract structured information from complex documents and answer questions based on visual-textual alignment in low-resource scripts. Use when the user wants to benchmark on SEA-Vision, or asks about evaluating this task. Reports answer accuracy.

researchpythongo
0
3
Sea Spoof EvalA

This benchmark evaluates audio deepfake detection models on their ability to distinguish real speech from synthetic speech across six South-East Asian languages. It specifically probes cross-lingual generalization and robustness against diverse open-source and commercial text-to-speech and voice conversion systems. Use when the user wants to benchmark on SEA-Spoof, or asks about evaluating this task. Reports EER (%).

researchpythonexpress
0
3
Sea Helm EvalA

Evaluates Thai language models across eight competencies including instruction following, multi-turn dialogue stability, natural language understanding, generation, reasoning, safety, and code-switching resistance. Use when the user wants to benchmark on SEA-HELM, or asks about evaluating this task. Reports SEA-HELM Average Score.

researchpythongo
0
3
Se Toxicity EvalA

Evaluates the ability of contemporary toxicity detection models to correctly identify toxic language in software engineering contexts, such as code reviews and developer chat logs. It probes whether general-purpose classifiers can handle domain-specific terminology and contextual nuances without significant performance degradation. Use when the user wants to benchmark on Jigsaw Sample, Code Review, Gitter Ethereum, or asks about evaluating this task. Reports F-Score.

researchpythongo
0
3
Se Asr Wer EvalA

Evaluates how speech enhancement (SE) artifacts and noise errors affect automatic speech recognition (ASR) performance. It measures Word Error Rate (WER) on enhanced speech signals derived from simulated and real-world reverberant noisy conditions to isolate the impact of artifact components. Use when the user wants to benchmark on Simulated WSJ0+CHiME-3, CHiME-3 et05_real, or asks about evaluating this task. Reports WER [%].

researchpythonperformance
0
3