Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,870
skills in category
870
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 8,257–8,280 of 20,870 skills

Credit Card Fraud Detection EvalA

Evaluates a model's ability to detect fraudulent credit card transactions in a streaming context by learning topological and sequential patterns from transaction graphs without manual feature engineering. Use when the user wants to benchmark on Credit Card Transaction Dataset (Feb-Sep), or asks about evaluating this task. Reports AP.

researchpythonnode
0
3
Creation Mmbench EvalA

Evaluates context-aware creative intelligence in multimodal and text-only models by assessing their ability to generate creative, contextually relevant content while maintaining visual factuality across diverse functional and creative writing tasks. Use when the user wants to benchmark on Creation-MMBench, or asks about evaluating this task. Reports VFS, Reward.

researchpythongit
0
3
Crbench EvalA

Evaluates whether multimodal large language models can perform genuine visual reasoning on charts by inferring values from axes and scales, rather than relying on OCR or pre-existing annotations. It probes the model's ability to interpret complex visual structures and perform multi-step estimation on both synthetic and real-world charts. Use when the user wants to benchmark on CRBench, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Craw4llm EvalA

Evaluates the efficiency and data quality of web crawling strategies for LLM pretraining by measuring downstream model performance after training on crawled or selected documents. It compares graph-connectivity-based, random, and pretraining-influence-based URL scoring methods against an oracle baseline. Use when the user wants to benchmark on ClueWeb22-A (English subset), or asks about evaluating this task. Reports Average performance on 22 core tasks (DCLM evaluation recipe).

researchpythongo
0
3
Craibench EvalA

CrAIBench probes the robustness of Web3 AI agents against context manipulation attacks, specifically memory injection and prompt injection. It evaluates whether agents can maintain user intent and resist adversarial goals when malicious instructions are embedded in historical memory or active prompts. Use when the user wants to benchmark on CrAIBench, or asks about evaluating this task. Reports Targeted Attack Success Rate (ASR).

researchpythongo
0
3
Crag EvalA

Evaluates the factual reliability and hallucination resistance of Retrieval-Augmented Generation (RAG) systems on realistic, dynamic, and long-tail questions. It measures how well models avoid generating incorrect information and appropriately abstain when knowledge is missing. Use when the user wants to benchmark on CRAG, or asks about evaluating this task. Reports truthfulness.

researchpythongo
0
3
Craft EvalA

Evaluates the ability of instruction-tuned LLMs to perform domain-specific multiple-choice question answering and text generation tasks. It measures how well models fine-tuned on synthetic, corpus-retrieved data generalize to held-out human-annotated benchmarks in biology, medicine, commonsense, recipe generation, and summarization. Use when the user wants to benchmark on ScienceQA (BioQA), MedMCQA (MedQA), CommonsenseQA 2.0 (CSQA), RecipeNLG (RecipeGen), CNN-DailyMail (Summarization), or ask...

researchpythongo
0
3
Cracknex EvalA

Evaluates few-shot crack segmentation performance under low-light conditions using illumination-invariant features. It tests the model's ability to generalize from well-illuminated support images to unseen low-light query images in both synthetic and real-world scenarios. Use when the user wants to benchmark on ll_CrackSeg9k, LCSD, or asks about evaluating this task. Reports mIOU.

researchpythonperformance
0
3
Cqa EvalA

Evaluates a model's ability to answer multiple-choice commonsense reasoning questions. It probes whether providing natural language explanations (human or model-generated) alongside questions improves reasoning performance compared to a baseline without explanations. Use when the user wants to benchmark on CQA, or asks about evaluating this task. Reports Accuracy (%).

researchpythongo
0
3
Cpt Merging EvalA

Evaluates the effectiveness of merging Continual Pretraining (CPT) models to build domain-specialized financial LLMs. It probes the model's ability to recover lost general knowledge, exhibit cross-domain complementarity, and demonstrate emergent reasoning capabilities through weight-space integration. Use when the user wants to benchmark on Financial Benchmark, or asks about evaluating this task. Reports Macro-Gain.

researchpythonperformance
0
3
Cpsc2018 Ecg Classification EvalA

Evaluates deep learning models on multi-lead ECG signal classification for arrhythmia detection under artificially balanced conditions. It probes the model's ability to extract discriminative temporal-spatial features from raw 12-lead cardiac signals and maintain robustness against various types of physiological noise. Use when the user wants to benchmark on CPSC2018, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Cps3d Seg EvalA

Evaluates 3D point cloud segmentation models for detecting surface defects on integrated circuit package substrates. It probes the model's ability to accurately classify high-density point clouds into normal and defect categories under industrial inspection conditions. Use when the user wants to benchmark on CPS3D-Seg, or asks about evaluating this task. Reports mIoU.

researchpythongo
0
3
Cppe 5 EvalA

Evaluates object detection models on fine-grained medical personal protective equipment (PPE) in complex, real-world scenes. It probes a model's ability to localize and classify coveralls, face shields, gloves, masks, and goggles from non-canonical perspectives, measuring detection accuracy across multiple IoU thresholds and object scales. Use when the user wants to benchmark on CPPE-5, or asks about evaluating this task. Reports AP (mean Average Precision).

researchpythongo
0
3
Cp Bench EvalA

This benchmark evaluates speech-LLMs on contextual and paralinguistic reasoning tasks. It probes the models' ability to integrate linguistic content with emotional, prosodic, and social cues from in-the-wild speech data to answer specific question types. Use when the user wants to benchmark on CP-Bench, or asks about evaluating this task. Reports LLaMA-3-70B judge score.

researchpythongo
0
3
Cow Bench EvalA

Evaluates modal, spatial, and temporal consistency in general world models through 18 sub-tasks across six task categories. It uses human-designed checklists to verify fine-grained physical laws, causal reasoning, and cross-modal alignment, moving beyond perceptual metrics to hard verification. Use when the user wants to benchmark on CoW-Bench, or asks about evaluating this task. Reports checklist_score.

researchpythongo
0
3
Covost2 St Mt EvalA

Evaluates multilingual speech-to-text translation and automatic speech recognition across 22 languages. It probes the model's ability to transcribe spoken audio and translate it into English (or from English) under monolingual, bilingual, and multilingual training regimes. Use when the user wants to benchmark on CoVoST 2, or asks about evaluating this task. Reports BLEU.

researchpythongo
0
3
CovochevalA

Evaluates zero-shot conversational voice cloning systems on their ability to generate natural, expressive speech that matches a target speaker's timbre and spontaneous style without prior training on the target speaker. It measures pronunciation accuracy, speaker similarity, and subjective qualities like naturalness, quality, and spontaneous style. Use when the user wants to benchmark on HQ-Conversations / CoVoC Test Prompts, or asks about evaluating this task. Reports FS.

researchpythongo
0
3
Covidx Classification EvalA

Evaluates deep learning models' ability to classify chest X-ray images into three diagnostic categories: normal, non-COVID-19 pneumonia, and COVID-19. It probes medical image classification performance under realistic class imbalance and tests whether models rely on clinically relevant lung regions or artifacts. Use when the user wants to benchmark on COVIDx, or asks about evaluating this task. Reports F-score.

researchpythongo
0
3
Covid19 Xray Classification EvalA

This evaluation probes a model's ability to classify chest X-rays as COVID-19 positive or negative using a cross-modal distillation setup where CT images are only used during training. It specifically tests the robustness of transfer learning under extremely small, patient-level paired cohorts and prevalence-heavy validation splits. Use when the user wants to benchmark on COVID-19 Image Data Collection, or asks about evaluating this task. Reports Accuracy.

researchpythonperformance
0
3
Covid19 Cxr Detection EvalA

Evaluates a deep CNN's ability to classify chest X-ray images into COVID-19 positive and negative/healthy categories using region-edge and channel-boosted features. Use when the user wants to benchmark on Three datasets (names not specified in section), or asks about evaluating this task. Reports 95% Confidence Interval (CI).

researchpythongo
0
3
Covid19 Cxr Ct Detection EvalA

Evaluates a lightweight CNN's ability to classify chest radiology images (CXR and CT scans) as positive or negative for COVID-19, and in a three-class setting. It also tests cross-modality generalization by training on CT and testing on CXR data. Use when the user wants to benchmark on CXR/CT Chest Radiology Dataset, or asks about evaluating this task. Reports accuracy.

researchpythontesting
0
3
Covid Net Cxr 2 EvalA

Binary classification of chest X-ray images to detect SARS-CoV-2 infection. It probes a model's ability to distinguish COVID-19 positive cases from negative cases (including no pneumonia and non-SARS-CoV-2 pneumonia) using a large, multinational dataset. Use when the user wants to benchmark on COVID-Net CXR-2 benchmark dataset, or asks about evaluating this task. Reports Sensitivity.

researchpythongo
0
3
Covid Drug Design EvalA

This benchmark evaluates deep graph generative models (JT-VAE and DQN) for their ability to design novel molecular structures optimized for high predicted potency against the SARS-CoV-2 3CL-protease, while balancing drug-likeness, lipophilicity, and synthesizability. It also assesses structural novelty relative to known antivirals and predicted binding affinity using in silico classifiers. Use when the user wants to benchmark on ChEMBL/BindingDB/ToxCat pharmacology dataset, or asks about eval...

researchpythongit
0
3
Covid Chestxray Enhancement EvalA

Evaluates the impact of five image enhancement techniques (histogram equalization, CLAHE, complement, gamma correction, BCET) on six CNN architectures for three-class classification (COVID-19, lung opacity, normal) using chest X-ray images. It also assesses whether lung segmentation improves classification accuracy and model interpretability. Use when the user wants to benchmark on COVQU-20, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3