Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

21,231
skills in category
885
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 8,953–8,976 of 21,231 skills

Cicids2017 Adversarial EvalA

Evaluates the adversarial robustness of tree ensemble models (RF, XGB, LGBM, EBM) on enterprise network intrusion detection using the CICIDS2017 dataset. It measures how well models maintain detection performance on benign and malicious traffic when subjected to constrained adversarial perturbations of time-series traffic features. Use when the user wants to benchmark on CICIDS2017, or asks about evaluating this task. Reports F1S.

researchpythonperformance
0
3
Cic Ids2017 Nids EvalA

Evaluates the classification accuracy and adversarial robustness of a Graph Neural Network-based Network Intrusion Detection System (NIDS) on distinguishing benign traffic from various attack types in network flow data. Use when the user wants to benchmark on CIC-IDS2017, or asks about evaluating this task. Reports weighted F1-score.

researchpythongo
0
3
Churro Ds EvalA

Evaluates the ability of vision-language models and OCR systems to accurately transcribe historical documents, including both printed and handwritten text across diverse languages and scripts. It probes robustness to long-term document degradation, variable layouts, and long-context inputs in zero-shot and fine-tuned settings. Use when the user wants to benchmark on Churro-DS, or asks about evaluating this task. Reports normalized Levenshtein similarity.

researchpythongo
0
3
Chumor 1.0 EvalA

Evaluates large language models' ability to understand and explain culturally nuanced Chinese internet humor. It measures how well models can generate human-preferred, two-sentence explanations for jokes derived from the Chinese platform Ruo Zhi Ba. Use when the user wants to benchmark on Chumor 1.0, or asks about evaluating this task. Reports winning rate.

researchpythongo
0
3
Chronos Forecasting EvalA

Evaluates time series forecasting models on in-domain and zero-shot benchmarks across diverse domains and frequencies. It probes a model's ability to generalize to unseen temporal patterns using both probabilistic and point forecast metrics. Use when the user wants to benchmark on Benchmark I, Benchmark II, or asks about evaluating this task. Reports WQL.

researchpythonperformance
0
3
Chronoroot2 Segmentation EvalA

Evaluates deep learning models for simultaneous multi-organ segmentation in 2D plant images. It probes the model's ability to accurately delineate root systems and aerial parts across different plant species and light conditions, while preserving structural fidelity for downstream phenotypic analysis. Use when the user wants to benchmark on Arabidopsis thaliana 2D phenotyping dataset, Tomato 2D phenotyping dataset, or asks about evaluating this task. Reports Dice coefficient.

researchpythongit
0
3
Chren Bleu EvalA

Evaluates machine translation quality between Cherokee and English, focusing on low-resource, morphologically complex translation. It probes both in-domain and out-of-domain generalization, as well as the reliability of automatic metrics versus human judgment for polysynthetic languages. Use when the user wants to benchmark on ChrEn, or asks about evaluating this task. Reports BLEU.

researchpythongit
0
3
Chirp EvalA

Evaluates the quality of open-ended, long-form responses from Vision-Language Models (VLMs) using pairwise preference ranking across five fine-grained criteria: overall preference, relevance, reasoning, hallucinations, and details. Use when the user wants to benchmark on CHIRP, or asks about evaluating this task. Reports pairwise_preference.

researchpythongo
0
3
Chiral Action Recognition EvalA

Evaluates a video representation's sensitivity to temporal direction by distinguishing between temporally opposite actions (e.g., opening vs. closing a door). Also tests general action recognition capability via linear probing on standard benchmarks. Use when the user wants to benchmark on Something-Something v2, EPIC-Kitchens, Charades, Kinetics-400, UCF-101, HMDB-51, or asks about evaluating this task. Reports Chiral Accuracy.

researchpythongo
0
3
Chinese Toxicity Detection EvalA

This protocol evaluates models on Chinese toxicity detection across two tasks: binary sentence-level classification and fine-grained toxic span extraction. It measures classification accuracy and precision/recall, while also assessing the model's ability to extract contiguous, human-readable toxic spans and the faithfulness of those explanations via confidence masking. Use when the user wants to benchmark on COLD, ToxiCN, CNTP, or asks about evaluating this task. Reports F1, Overlap F1.

researchpythongo
0
3
Chinese Tibetan Medicine Qa EvalA

This benchmark evaluates a traceable cross-source retrieval-augmented generation framework for Chinese Tibetan medicine QA. It probes the model's ability to route queries across heterogeneous knowledge bases, fuse cross-source evidence, and generate faithful answers with correct citations. Use when the user wants to benchmark on Chinese Tibetan-medicine QA dataset, or asks about evaluating this task. Reports CrossEv@5.

researchpythongo
0
3
Chinese Llm Benchmarks EvalA

Evaluates the knowledge, reasoning, instruction-following, and safety alignment capabilities of Chinese instruction-tuned LLMs across academic, professional, open-ended, and safety-critical domains. Use when the user wants to benchmark on C-Eval, CMMLU, BELLE-EVAL, SafetyBench, or asks about evaluating this task. Reports log-likelihood.

researchpythongo
0
3
Chinese Llm Bench EvalA

Evaluates Chinese large language models' world knowledge, academic understanding, and multi-dimensional alignment after pretraining or instruction fine-tuning. It probes the model's ability to follow instructions, reason across domains, and maintain safety and helpfulness standards in Chinese. Use when the user wants to benchmark on C-Eval, CMMLU, Alignbench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Chinesafe EvalA

This benchmark evaluates the safety of large language models in Chinese by testing their ability to correctly classify text as safe or unsafe across multiple sensitive categories. It probes whether models can reliably detect harmful, policy-violating, or sensitive content in a Chinese-language context using both generation-based and perplexity-based strategies. Use when the user wants to benchmark on ChineseSafe, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Chimera Extraction EvalA

Evaluates a model's ability to identify scientific idea recombinations in research abstracts, extract the involved scientific concepts (entities), and classify the type of recombination relation (inspiration or blend) between them. Use when the user wants to benchmark on CHIMERA, or asks about evaluating this task. Reports Precision, Recall, F1.

researchpythongo
0
3
Chime4 EvalA

Evaluates end-to-end speech recognition and speech enhancement performance in noisy, reverberant multi-channel conditions. It probes the model's ability to jointly dereverberate, denoise, and transcribe speech using self-supervised learning representations. Use when the user wants to benchmark on CHiME-4, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Chime4 Asr Robustness EvalA

Evaluates the robustness of automatic speech recognition systems in noisy environments by measuring word error rates on enhanced speech from speaker extraction models across matched and mismatched noisy conditions. Use when the user wants to benchmark on CHiME-4, VoiceBank-DEMAND, WHAM!, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Chime4 Ami Asr EvalA

Evaluates automatic speech recognition systems on real-world meeting and close-talking microphone speech. It measures how well acoustic models trained on raw waveforms generalize to challenging multi-microphone environments compared to traditional feature-based baselines. Use when the user wants to benchmark on CHiME4, AMI, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Chime4 Aishell4 Asr EvalA

Evaluates multi-channel end-to-end speech recognition systems in noisy, real-world environments using neural beamforming front-ends combined with CTC-CRF acoustic models. It probes the model's ability to leverage single-channel data via pre-training, data scheduling, or simulation to improve robustness and accuracy. Use when the user wants to benchmark on CHiME4, AISHELL-4, or asks about evaluating this task. Reports WER.

researchpythonshell
0
3
Chime4 Adapter EvalA

Evaluates the robustness of Automatic Speech Recognition (ASR) models under various real-world noise conditions (bus, cafe, pedestrian, street junction) using adapter-based fine-tuning and speech enhancement front-ends. Use when the user wants to benchmark on CHiME-4, or asks about evaluating this task. Reports WER.

researchpythontesting
0
3
Chime3 Se EvalA

This benchmark evaluates speech enhancement models by measuring how well they clean noisy speech features before they are processed by a downstream automatic speech recognition (ASR) system. It probes the model's ability to preserve speech structure and reduce noise in challenging real-world far-field conditions. Use when the user wants to benchmark on CHiME-3, or asks about evaluating this task. Reports WER.

researchpythontesting
0
3
Chime2 Robust Asr EvalA

Evaluates the effectiveness of a time-domain speech enhancement frontend in improving automatic speech recognition performance on noisy and reverberant speech. It probes the model's ability to enhance speech without introducing distortion that degrades downstream ASR accuracy. Use when the user wants to benchmark on CHiME-2, or asks about evaluating this task. Reports WER.

researchpythonfrontend
0
3
Chikha Po EvalA

Evaluates fundamental lexical comprehension and generation capabilities of multilingual LLMs across thousands of languages. It probes word-level translation, context-aware translation, translation-conditioned language modeling, and bag-of-words machine translation tasks to measure basic linguistic competence beyond high-resource languages. Use when the user wants to benchmark on ChiKhaPo, or asks about evaluating this task. Reports language score.

researchpythongo
0
3
Chickenpox Hungary Forecasting EvalA

Evaluates the ability of recurrent graph neural networks to forecast spatiotemporal epidemiological time series. It probes how well models capture spatial dependencies between adjacent regions and temporal dynamics like seasonality and zero-inflation over multiple forecasting horizons. Use when the user wants to benchmark on Chickenpox Cases in Hungary, or asks about evaluating this task. Reports mean squared error.

researchpythonnode
0
3