Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,265–6,288 of 20,818 skills
Evaluates a deep learning model's ability to classify malware images across multiple tasks, including binary classification, malware family classification, and detection of obfuscation techniques across Windows, Android, macOS, and Linux platforms. Use when the user wants to benchmark on Maling benchmark dataset, or asks about evaluating this task. Reports accuracy.
Evaluates static, dynamic, and hybrid analysis pipelines for malware detection. Models are trained on opcode and API call sequences to distinguish malware families from benign Windows executables. Use when the user wants to benchmark on Malware Detection Dataset, or asks about evaluating this task. Reports Area under the ROC curve.
Evaluates the capability of anti-malware detection engines and behavior profiling methods to correctly group malware variants into their respective families based on runtime Windows API call sequences and parameters. Use when the user wants to benchmark on 40Bot, 419Mal, or asks about evaluating this task. Reports Pairwise Classification Score (PCS).
Evaluates the ability of LLMs and their ensembles to correctly classify malware samples into one of ten canonical families based on their behavior or code semantics. It probes robustness to class imbalance and the effectiveness of hierarchical decision-making under obfuscation. Use when the user wants to benchmark on Gold-standard malware family dataset, or asks about evaluating this task. Reports Macro F1-score.
Binary classification of software binaries as benign or malicious based on their control flow graphs. It probes the model's ability to learn graph-structured representations and route them through specialized experts for accurate detection. Use when the user wants to benchmark on BODMAS, DikeDataset, PMML, or asks about evaluating this task. Reports Accuracy.
This benchmark evaluates the effectiveness of unsupervised clustering algorithms on large-scale malware binary datasets. It probes how well different feature representations and clustering methods can group malware samples into coherent families while handling real-world noise and benign samples. Use when the user wants to benchmark on Bodmas, Ember, Security, or asks about evaluating this task. Reports Homogeneity.
This evaluation probes a model's ability to classify long-sequence binary and executable files into malware families or benign/malicious categories. It tests robustness to varying sequence lengths, compression formats, and real-world malware distribution characteristics compared to synthetic long-range benchmarks. Use when the user wants to benchmark on Kaggle (BIG 2015), Drebin, EMBER, LRA, or asks about evaluating this task. Reports accuracy.
Evaluates host-based intrusion detection models on their ability to classify multi-label malware behaviors from truncated Windows API call sequences. Probes how well different neural architectures handle sequential behavioral data and feature selection strategies for detecting overlapping malicious activities. Use when the user wants to benchmark on Behavioural Reports of Multi-Stage Malware, or asks about evaluating this task. Reports F1-score.
Evaluates a CNN's ability to classify Windows PE files as malware or benign by learning spatial features from grayscale images generated from dynamic API call argument sequences. It probes the model's resilience to obfuscation by leveraging behavioral temporal patterns converted into visual representations. Use when the user wants to benchmark on Windows PE Malware Dataset, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates an AI system's ability to analyze low-level process execution logs and identify malicious signals from malware detonations. It probes structured data parsing, security event correlation, and malware family classification capabilities. Use when the user wants to benchmark on CyberSOCEval Malware Analysis, or asks about evaluating this task. Reports accuracy.
Evaluates graph neural networks for Android malware family classification under intra-family and cross-family distribution shifts. It probes how semantic feature enrichment (function metadata and LLM embeddings) and test-time/domain adaptation methods mitigate performance degradation when models encounter unseen malware families. Use when the user wants to benchmark on MalNet-Tiny, MalNet-Tiny-Common, or asks about evaluating this task. Reports accuracy.
Evaluates memory-aware long sequence compression techniques in large-scale sequential recommendation. It probes how well methods balance memory overhead, computational cost, and ranking accuracy when processing long user interaction histories to predict future clicks. Use when the user wants to benchmark on Amazon-Electronic, MicroVideo1.7M, KuaiVideo, or asks about evaluating this task. Reports AUC.
Evaluates the ability of deep learning models to classify malware families from their visual representations (malware images). It probes feature extraction robustness and classification accuracy on an imbalanced dataset. Use when the user wants to benchmark on MalImg, or asks about evaluating this task. Reports Accuracy.
This benchmark evaluates the ability of machine learning models to classify malware binaries by converting them into grayscale images and predicting their specific family among 25 categories. It probes classification accuracy and computational efficiency across different neural network architectures. Use when the user wants to benchmark on Malimg, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to classify malware samples into one of seven predefined families based on their opcode sequences. It probes the model's capacity to capture structural and sequential patterns in low-level code representations. Use when the user wants to benchmark on Malicia, or asks about evaluating this task. Reports accuracy.
Evaluates radiative transfer models' ability to simulate exoplanet transit and direct-imaging spectra under controlled atmospheric conditions. Probes how atmospheric discretization, opacity treatments, and spectroscopic databases impact spectral predictions. Use when the user wants to benchmark on MALBEC Test Suite, or asks about evaluating this task. Reports ppm.
Evaluates the robustness and consistency of a model's mathematical reasoning by aggregating predictions across multiple inference runs per problem. It measures whether the most frequent prediction among k samples matches the ground truth. Use when the user has predictions and gold and needs to compute Majority@k.
Evaluates retrieval models' ability to follow complex, task-specific instructions across diverse domains and long-tail tasks. It measures how instruction tuning impacts generalization and performance on heterogeneous query-document relevance tasks compared to non-instruction-tuned baselines. Use when the user wants to benchmark on MAIR, or asks about evaluating this task. Reports nDCG@10.
Evaluates the instruction-following capability, output preference quality, and general reasoning performance of fine-tuned LLMs. It probes how well models adhere to explicit constraints, generate preferred responses relative to a baseline, and solve standard academic benchmarks. Use when the user wants to benchmark on AlpacaEval, IFEval, ARC, HellaSwag, Winogrande, MMLU, TruthfulQA, or asks about evaluating this task. Reports AlpacaEval.
Evaluates the accuracy and inference speed of monocular depth estimation models on resource-constrained mobile devices. It measures depth prediction quality using invariant standard root mean squared error and records inference time on a Raspberry Pi 4 to assess real-time capability. Use when the user wants to benchmark on MAI&AIM2022 challenge dataset, or asks about evaluating this task. Reports si-RMSE.
Evaluates a model-free runtime system for dynamically scaling uncore frequencies in heterogeneous CPU-GPU architectures. It probes the system's ability to balance energy efficiency and performance across diverse HPC, molecular dynamics, and deep learning workloads. Use when the user wants to benchmark on Altis, ECP proxy applications, AI-enabled applications, MLPerf benchmarks, Altis-SYCL, or asks about evaluating this task. Reports Energy Delay Product (EDP).
Evaluates a lightweight vision-language model's performance on reasoning, OCR, and real-world understanding benchmarks, alongside deployment efficiency metrics like inference latency and throughput on mobile hardware. Use when the user wants to benchmark on HallusionBench, MMBench, RealworldQA, MMStar, OCRBench, AI2D, TextVQA, CRPE, MME Realworld, DocVQA, or asks about evaluating this task. Reports accuracy.
Evaluates text-to-image generation models on their ability to produce images free of fine-grained artifacts, specifically probing subject anatomy, attributes, and interactions. It measures detection accuracy using a hierarchical taxonomy of artifact types to benchmark model robustness against visual inconsistencies. Use when the user wants to benchmark on MagicData340K, or asks about evaluating this task. Reports F1-Score.
Evaluates instruction-based image editing models on their ability to modify source images according to text instructions while preserving original content and style. It tests both single-turn and multi-turn editing capabilities against ground truth edited images. Use when the user wants to benchmark on MagicBrush, or asks about evaluating this task. Reports CLIP image similarity.