Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

21,231
skills in category
885
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 8,929–8,952 of 21,231 skills

Clamp2 Music Retrieval EvalA

Evaluates a model's ability to classify symbolic music into genres, emotions, or composer styles, and to perform cross-modal semantic search between music scores (ABC/MIDI) and textual descriptions. It also probes multilingual retrieval capabilities by testing performance across machine-translated text queries. Use when the user wants to benchmark on WikiMT, VGMIDI, Pianist8, MidiCaps, or asks about evaluating this task. Reports Accuracy, MRR.

researchpythongo
0
3
Clam Adversarial EvalA

Evaluates the robustness of classical neural networks and quantum neural networks against adversarial attacks by measuring performance degradation on a malware classification task after injecting random noise into input features. Use when the user wants to benchmark on ClaMP_Integrated, or asks about evaluating this task. Reports accuracy.

researchpythongit
0
3
Cl Ifeval EvalA

Evaluates large language models' ability to follow complex, variable-driven instructions across multiple languages. It measures strict compliance with prompt constraints to reveal cross-lingual robustness disparities. The benchmark highlights how functional tasks expose performance gaps that static benchmarks often miss. Use when the user wants to benchmark on CL-IFEval, or asks about evaluating this task. Reports Strict Prompt Accuracy.

researchpythonperformance
0
3
Cl Gsmsym EvalA

Assesses mathematical reasoning and symbolic computation capabilities of LLMs across multiple languages. It uses dynamic, variable-driven templates to generate verifiable ground truths for each instance. The evaluation probes model resilience to linguistic variations and template-specific weaknesses. Use when the user wants to benchmark on CL-GSMSym, or asks about evaluating this task. Reports accuracy.

researchpythonperformance
0
3
Cl Engineering Regression EvalA

Evaluates continual learning strategies for 3D engineering regression tasks, measuring their ability to learn from sequential data streams while mitigating catastrophic forgetting and maintaining predictive accuracy across parametric and point cloud modalities. Use when the user wants to benchmark on SplitSHIPD-Par, SplitSHIPD-PC, SplitSHAPENET, SplitRAADL, SplitDRIVAERNET, SplitDRIVAERNET++-Par, SplitDRIVAERNET++-PC, or asks about evaluating this task. Reports MPE.

researchpythonperformance
0
3
Cl Drive Cognitive Load EvalA

Evaluates a model's ability to classify cognitive load levels from raw, multimodal physiological signals (EEG, ECG, EDA) collected during driving scenarios. It probes the model's capacity to learn temporal and cross-modal patterns without hand-crafted features. Use when the user wants to benchmark on CL-Drive, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Civil Comments Toxicity EvalA

Evaluates a RoBERTa-based classifier's ability to detect toxic or harmful content in online comments. It probes the model's sensitivity to explicit lexical cues versus implicit, context-dependent toxicity, highlighting failure modes that aggregate accuracy metrics miss. Use when the user wants to benchmark on Civil Comments, or asks about evaluating this task. Reports Accuracy.

researchpythonexpress
0
3
Cityscapes EvalA

Evaluates semantic scene understanding models on complex urban street scenes by measuring pixel-level classification accuracy and instance-level segmentation quality. It probes the model's ability to handle high-resolution imagery, diverse weather/lighting conditions, and fine-grained class distinctions in autonomous driving contexts. Use when the user wants to benchmark on Cityscapes, or asks about evaluating this task. Reports IoU.

researchpythonperformance
0
3
Cityintrusion EvalA

Evaluates real-time dynamic pedestrian intrusion detection from moving camera views, jointly performing area-of-interest segmentation and pedestrian detection to classify whether a pedestrian has intruded into a dynamic zone. It measures classification accuracy, segmentation quality, and detection precision while tracking computational efficiency. Use when the user wants to benchmark on Cityintrusion, Cityperson, Cityscape, or asks about evaluating this task. Reports PID_Acc.

researchpython
0
3
Citesum EvalA

This benchmark evaluates a model's ability to generate concise, single-sentence summaries of scientific papers using citation sentences as ground truth. It probes extreme summarization capabilities and domain adaptation across academic disciplines. Use when the user wants to benchmark on CiteSum, or asks about evaluating this task. Reports ROUGE-1, ROUGE-2, ROUGE-L.

researchpythongit
0
3
Citebench EvalA

Evaluates the capability of citation recommendation models to identify relevant academic references given local citation contexts. It probes robustness across varying contextual features, including context length, reference position, academic field, publication year, citation count, and part-of-speech tags. Use when the user wants to benchmark on S2ORC/S2AG Diagnostic Datasets, or asks about evaluating this task. Reports MRR.

researchpythongo
0
3
Citation Rec EvalA

Evaluates the ability of citation recommendation systems to retrieve and rank relevant academic papers given a query context. It probes ranking quality, recall of relevant candidates, and normalized discounted cumulative gain across multiple academic datasets. Use when the user wants to benchmark on ACL-200, FullTextPeerRead, Refseer, arXiv, ArSyTa, or asks about evaluating this task. Reports MRR.

researchpythongo
0
3
Cirthan EvalA

Evaluates composed image retrieval capabilities on culturally specific Thangka imagery. It tests the model's ability to align fine-grained sketch+text queries with target images across varying levels of textual semantic granularity, highlighting the domain gap between generic pre-training and specialized cultural retrieval. Use when the user wants to benchmark on CIRThan, or asks about evaluating this task. Reports Recall@K (R@K).

researchpythongo
0
3
Circuit EvalA

Evaluates large language models' ability to interpret analog circuit diagrams and netlists, and perform multi-level reasoning to calculate correct numerical values for circuit parameters. Use when the user wants to benchmark on CIRCUIT, or asks about evaluating this task. Reports accuracy.

researchpythonaws
0
3
Ciqi Bench EvalA

Evaluates a multimodal agent's ability to perform fine-grained visual classification and cultural reasoning on antique Chinese porcelain. It probes seven specific connoisseurship attributes (dynasty, reign period, kiln site, glaze color, decorative motif, vessel shape, and overall naming) through both multiple-choice and free-form generation tasks. Use when the user wants to benchmark on CiQi-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Cinic 10 EvalA

Evaluates image classification performance under significant domain shift between synthetic (CIFAR-10) and real-world/downsampled (ImageNet) sources. Probes model robustness to distributional bias and class-level statistical divergence across training and test domains. Use when the user wants to benchmark on CINIC-10, or asks about evaluating this task. Reports Test Error.

researchpythongo
0
3
Cine Tech Bench EvalA

Evaluates multimodal large language models and video generation models on fine-grained cinematographic understanding (shot scale, angle, composition, camera movement, lighting, color, focal length) and camera movement generation from video clips. Use when the user wants to benchmark on CineTechBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Cimi4d Annotation EvalA

Probes the accuracy of automatically generated 3D human pose and translation annotations for rock climbing motions. It evaluates how well a LiDAR-IMU fusion and blending optimization pipeline reconstructs off-ground climbing poses compared to manual ground truth. Use when the user wants to benchmark on CIMI4D, or asks about evaluating this task. Reports PMPJPE.

researchpythonrust
0
3
Cifar100 Miniimagenet EvalA

This evaluation protocol probes a model's ability to perform standard supervised image classification and adaptive few-shot episodic learning. It measures how well the model generalizes to unseen classes under limited supervision by averaging accuracy over multiple sampled episodes. Use when the user wants to benchmark on CIFAR-100, Mini-ImageNet, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Cifar10 Histopathology EvalA

Evaluates the generalization and uncertainty quantification of Bayesian Neural Networks trained with novel Jensen-Shannon divergence loss functions compared to standard KL divergence, specifically under noisy and class-biased data conditions. Use when the user wants to benchmark on CIFAR-10, Breast Histopathology Dataset, or asks about evaluating this task. Reports validation accuracy.

researchpythongo
0
3
Cifar Robustness EvalA

Evaluates the robustness of adversarially trained neural networks against Projected Gradient Descent (PGD) attacks on CIFAR-10 and CIFAR-100. It measures both clean (natural) classification accuracy and robust accuracy under varying attack strengths (PGD-20 and PGD-100). Use when the user wants to benchmark on CIFAR-10, CIFAR-100, or asks about evaluating this task. Reports PGD-20 accuracy.

researchpythongo
0
3
Cifake EvalA

Binary classification of real versus AI-generated synthetic images. It probes a model's ability to detect subtle background imperfections and artifacts introduced by latent diffusion models rather than semantic object content. Use when the user wants to benchmark on CIFAKE, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Ciciomt2024 Ifl EvalA

Evaluates incremental federated learning models for intrusion detection in IoT networks under evolving threat distributions. Probes the model's ability to adapt to concept drift over time while mitigating catastrophic forgetting in a federated setting. Use when the user wants to benchmark on CICIoMT2024, or asks about evaluating this task. Reports Accuracy (Acc).

researchpythongo
0
3
Cicids2017 EvalA

Evaluates network intrusion detection capability by classifying network traffic flows as benign or malicious (or specific attack types) using graph-structured representations of network connections. It probes the model's ability to learn from adaptive graph construction and contrastive learning under resource-constrained conditions. Use when the user wants to benchmark on CICIDS2017, or asks about evaluating this task. Reports accuracy.

researchpythonnode
0
3