Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,552
skills in category
982
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 4,801–4,824 of 23,552 skills

Sentencebench EvalA

This benchmark evaluates Persian grapheme-to-phoneme (G2P) systems on sentence-level text, specifically probing their ability to correctly map characters to phonemes and disambiguate homographs using contextual information. It measures both phonetic accuracy and contextual word-sense resolution capabilities. Use when the user wants to benchmark on SentenceBench, or asks about evaluating this task. Reports Homograph Acc. (%).

researchpythongit
0
3
Sentence Summarization EvalA

Evaluates abstractive summarization models on condensing source sentences into title-like summaries. It measures summary quality via lexical/semantic overlap and human judgments, while explicitly quantifying the degree of verbatim copying from the source text. Use when the user wants to benchmark on Gigaword, Newsroom, or asks about evaluating this task. Reports ROUGE-2.

researchpythongo
0
3
Sentence Stress Detection EvalA

Probes a model's ability to detect sentence-level prosodic stress at the word/token level using only audio input. It evaluates zero-shot generalization and alignment-free stress identification across diverse speech styles and synthetic/human datasets. Use when the user wants to benchmark on TinyStress-15K, Aix-MARSEC, Expresso, EmphAssess, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Sentence Representation EvalA

Evaluates the quality of sentence-level representations learned by a transformer-based autoencoder across semantic similarity, single- and multi-sentence classification, and controlled text generation. It probes the model's ability to capture semantic meaning, classify sentiment/acceptability, and reconstruct or modify text via vector arithmetic. Use when the user wants to benchmark on Semantic Textual Similarity (STS), GLUE benchmark, Yelp reviews, or asks about evaluating this task. Reports...

researchpythongo
0
3
Sentence Level Cal EvalA

Evaluates the recall and efficiency of continuous active learning systems for information retrieval when using sentence-level versus document-level relevance feedback. It measures how quickly a simulated reviewer can identify all relevant documents under varying effort models that account for assessment count and sentence reading time. Use when the user wants to benchmark on TREC Total Recall 2015 Track, HARD 2004 Track, or asks about evaluating this task. Reports Recall@E.

researchpythongo
0
3
Sentarl Trading EvalA

Evaluates a sentiment-aware reinforcement learning agent's ability to generate profitable and stable trading strategies across diverse market conditions, transaction cost regimes, and varying levels of financial news coverage. It benchmarks performance against a sentiment-free RL ablation and a buy-and-hold strategy using standard financial return and risk metrics. Use when the user wants to benchmark on 20-Asset Financial Time Series & News Corpus (2018-2020), or asks about evaluating this t...

researchpythongit
0
3
Sensorium 2023 EvalA

Predicts single-neuron responses in mouse primary visual cortex from dynamic video stimuli and behavioral covariates, probing spatio-temporal neural decoding and out-of-distribution generalization. Use when the user wants to benchmark on SENSORIUM 2023, or asks about evaluating this task. Reports R^2.

researchpythongo
0
3
Sensor Invariant Tactile EvalA

Evaluates the zero-shot transferability of tactile representations across different physical sensors for shape reconstruction, object classification, and 3D pose estimation. It probes whether learned features generalize across varying optical designs and manufacturing differences without retraining. Use when the user wants to benchmark on Real-world tactile contact dataset, or asks about evaluating this task. Reports top-1 accuracy.

researchpythongo
0
3
Seniortalk EvalA

Evaluates speech processing models on authentic, real-world conversations among super-aged Chinese speakers (75+). It probes capabilities in speaker verification, diarization, automatic speech recognition, and speech editing under conditions of age-related vocal degradation, dialectal variation, and presbyphonia. Use when the user wants to benchmark on SeniorTalk, or asks about evaluating this task. Reports EER.

researchpythonperformance
0
3
Semignn Alipay EvalA

Evaluates a semi-supervised graph neural network's ability to predict user loan defaults and classify user occupations using multiview graph data (social ties, app usage, nicks, addresses) on a large-scale financial platform dataset. Use when the user wants to benchmark on Alipay, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Semigda Medical Seg EvalA

Evaluates semi-supervised medical image segmentation performance under limited labeled data ratios (10% and 30%). It measures segmentation accuracy and boundary precision across multiple medical domains including colonoscopy, dermoscopy, pathology, and ultrasound. Use when the user wants to benchmark on Colonoscopy (CVC-ClinicDB, Kvasir, CVC-300), ISIC-2018, BCSS, BUSI, or asks about evaluating this task. Reports Dice coefficient (Dice).

researchpythonperformance
0
3
Semi Supervised Novelty Detection EvalA

Evaluates a model's ability to distinguish in-distribution (ID) samples from out-of-distribution (OOD) samples in a semi-supervised setting, where only labeled ID data and unlabeled mixed data are available during training. Use when the user wants to benchmark on MNIST, FashionMNIST, SVHN, CIFAR10, CIFAR100, ImageNet, or asks about evaluating this task. Reports AUROC.

researchpythongo
0
3
Semi Dynamic Context Compression EvalA

Evaluates the ability of context compression methods to preserve information density for downstream reading comprehension tasks. It probes whether adaptive, density-aware compression can maintain answer accuracy while significantly reducing context length compared to static baselines. Use when the user wants to benchmark on HotpotQA, SQuAD, Natural Questions, AdversarialQA, or asks about evaluating this task. Reports substring accuracy.

researchpythongo
0
3
Semeval2023 Task12 EvalA

Multilingual sentiment classification across low-resource African languages, including zero-shot generalization to unseen languages. Use when the user wants to benchmark on SemEval-2023 Task 12, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Semeval2022 Task10 EvalA

This benchmark evaluates structured sentiment analysis by testing a model's ability to extract sentiment targets, opinions, and their relational dependencies from text. It probes cross-lingual generalization and the capacity to repurpose semantic dependency parsers for sentiment graph generation. Use when the user wants to benchmark on SemEval-2022 Task 10, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Semeval2021task6 St3A

Detects persuasion techniques in multimodal memes by jointly analyzing text and image content. It probes cross-modal alignment and interaction modeling for persuasive content detection. Use when the user wants to benchmark on SemEval-2021 Task 6 Subtask 3, or asks about evaluating this task. Reports F1-Micro.

researchpythongo
0
3
Semeval2021task6 St2A

Identifies and classifies spans of persuasion techniques within unimodal text using sequence tagging. It probes the model's ability to localize and categorize persuasive spans at the token level. Use when the user wants to benchmark on SemEval-2021 Task 6 Subtask 2, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Semeval2021task6 St1A

Detects persuasion techniques in unimodal text by classifying which techniques are present in a given text snippet. It probes the model's ability to perform multilabel classification on propaganda and persuasive content. Use when the user wants to benchmark on SemEval-2021 Task 6 Subtask 1, or asks about evaluating this task. Reports F1-Micro.

researchpythongo
0
3
Semeval2020 Semantic Change EvalA

Evaluates a model's ability to detect and rank lexical semantic change over time across multiple languages. It probes both binary classification of whether a word's meaning has changed and graded ranking of the magnitude of that change. Use when the user wants to benchmark on SemEval 2020 Unsupervised Lexical Semantic Change Detection, or asks about evaluating this task. Reports Spearman's rank correlation.

researchpythongo
0
3
Semeval2017task4 EvalA

Evaluates the ability of models to classify sentiment in social media posts (tweets) across different languages (English and Arabic) and granularities (overall polarity, topic-specific polarity, and ordinal scales). Use when the user wants to benchmark on SemEval-2017 Task 4, or asks about evaluating this task. Reports macro-average recall.

researchpython
0
3
Semeval Question Relevancy EvalA

Evaluates a model's ability to rank question-answer pairs by relevance. It probes the system's capacity to understand semantic similarity and perform information retrieval tasks using language-independent features and multi-task learning. Use when the user wants to benchmark on SemEval-2016 Task 3, or asks about evaluating this task. Reports MAP.

researchpythongo
0
3
Semeval 2023 Task12 EvalA

Sentiment classification across twelve low-resource African languages and Creoles, evaluating model robustness to code-switching and varying degrees of lexical similarity to pretraining data. Use when the user wants to benchmark on SemEval-2023 Task 12, or asks about evaluating this task. Reports macro-F1.

researchpythongit
0
3
Semeval 2022 Task2 EvalA

Evaluates language models' ability to detect whether a multi-word expression (MWE) in a sentence is used idiomatically or literally, and to model the semantic similarity of idiomatic expressions. It tests compositionality understanding and contextual semantic representation across English, Portuguese, and Galician. Use when the user wants to benchmark on SemEval-2022 Task 2, or asks about evaluating this task. Reports macro F1 score.

researchpythongo
0
3
Semeval 2014 Task9 Sentiment EvalA

Evaluates sentiment polarity classification across diverse informal text genres, including tweets, sarcasm-marked tweets, and LiveJournal posts. It distinguishes between phrase-level contextual polarity and message-level sentiment, testing robustness to informal language, sarcasm-induced polarity inversion, and cross-platform generalization. Use when the user wants to benchmark on SemEval-2014 Task 9 Test Sets, or asks about evaluating this task. Reports macro- and micro-averaged F1.

researchpythongo
0
3